Clear all

Jobiba hiring network

Software Engineer Snowcommand Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software engineer snowcommand jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
1mo ago

About the Team GTM Growth Engineering builds AI-native products and systems that help OpenAI's go-to-market and B2B marketing organizations operate with greater speed, focus, and leverage. Our mandate is revenue leverage: products tied directly to pipeline quality, customer engagement, seller and marketer productivity, and the speed at which OpenAI can bring its technology to customers. We build the infrastructure and user experiences behind high-impact GTM workflows, including customer context, prioritization, routing, campaign execution, review surfaces, feedback loops, and measurement. Our work combines product craft, applied AI, reliable systems, and thoughtful operational design. About the Role We’re looking for a product-minded Software Engineer to build AI-powered products and full-stack experiences for GTM Growth Engineering. You will own meaningful product slices end to end, from user experience and frontend implementation to backend APIs, integrations, data models, instrumentation, and launch readiness. This is a role for engineers who want to build products that do real work in production. You will partner with Product, Design, Data Science, Sales, B2B Marketing, and operations teams to understand high-value workflows and ship systems that improve customer engagement, pipeline, conversion, and team productivity. The role is ideal for a strong product engineer who can move between product craft, systems engineering, applied AI, and measurable business outcomes. You should be excited to build from ambiguous problem statements, ship quickly, and improve products based on real user feedback. What You'll Do Build AI-powered products and workflows that help sales and B2B marketing teams identify opportunities, coordinate work, and engage customers more effectively. Own full-stack product experiences from prototype through launch, instrumentation, iteration, and production hardening. Design intuitive user journeys that combine polished interfaces, reliable servi

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team With Codex we’re building an AI software engineer. One that you can pair with, delegate to, or even ask to take on future tasks proactively. Our team is a fast-moving group within OpenAI, bringing together research, engineering, design, and product. We iteratively build the Codex agent harness and product to get the most out of the model, and we iteratively train the model to be great at complex software engineering tasks. The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. We operate across research, engineering, product, and infrastructure; owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. Codex Enterprise builds the ecosystem building blocks, discovery surfaces, and enterprise capabilities that help Codex spread across developers, teams, and organizations worldwide. It is a cross-cutting team that works across the stack to build both delightful product experiences and fundamental platform capabilities. Its customers range from individual developers and small teams to large enterprises, and our mission is critical to achieving the vision of Codex as a proactive teammate. About the Role As we grow, we’re focused on turning Codex from a powerful individual tool into a production-grade teammate for entire organizations. You will work across internal OpenAI teams and external customers, from fast-moving startups to large enterprises, to make it possible to deploy, operate, and trust Codex in increasingly demanding real-world environments. As Codex’s consumer adoption accelerates, enterprise demand is growing just as quickly, and there is also increasing opportunity to expand Codex through ecosystem capabilities that unlock new workflows, integrations, and discovery. This team helps turn messy, real-world team requirements into robust, repeatable, and scalable product and

pythonawsci/cd
View job →
O
1mo ago

About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun

awsazuregcp
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing

pythonawskubernetes
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. A majority of our users interact with our products in languages other than English, and our products must work seamlessly across languages, regions, and cultures. The Internationalization team builds the infrastructure that enables OpenAI products to ship globally by default. We develop the systems that power localization, international product launches, and high-quality global user experiences across all OpenAI products. About the Role As a Senior Software Engineer on the Internationalization team, you will build the systems that power localization and international product launches at OpenAI. You’ll work on the platform that manages product content, translation workflows, and localization infrastructure across our products. This role sits at the intersection of AI systems, developer platforms, and product infrastructure. In this role, you will Build and scale OpenAI’s localization, content, and experimentation platform used across OpenAI product teams, including open-source components: Develop AI-powered translation pipelines combined with human-in-the-loop review workflows. Design systems that reliably deliver localized product content across web and mobile apps. Build tools that enable linguists and localization teams to review and improve translations. Develop developer tooling that simplifies localization and internationalization workflows. Build and maintain internationalization libraries used across OpenAI products: Design systems that correctly handle numbers, currencies, dates, and pluralization across locales. Improve support for multilingual interfaces and right-to-left languages. Partner with product teams to improve the international readiness of new features. You might thrive in this role if you Have strong software engineering experience building backend or full-stack systems. Have familiarity with Java, React, MySQL, and cloud infrastructure p

javareactsql
View job →

About the Team The Platform Analytics team builds the systems OpenAI researchers use to understand the quality and behavior of the models we train including what models are doing, why they behave in a particular way, and how that behavior changes across experiments. Neptune is a core part of this work. It ingests, stores, queries, and visualizes large volumes of metrics from pretraining, post-training, and reinforcement learning. Hundreds of researchers depend on these systems in their daily work to compare experiments, debug unexpected behavior, and decide what to try next. Our scope is broader than metrics. We also build platforms that help researchers analyze samples, traces, evaluation results, and other structured or unstructured data through dashboards, APIs, and increasingly agent-driven workflows. These systems need to remain fast, reliable, and understandable as the scale and complexity of research change quickly. We are not trying to become a consulting team that builds a separate solution for every research project. We work directly with researchers to understand recurring problems, then turn them into reusable infrastructure and platform capabilities that many teams can build on. About the Role We’re looking for a hands-on experienced software engineer who can take ownership of a critical system and drive it from problem definition through production adoption. This person should be able to own a platform such as CacheHouse end to end: define its technical direction, design its data model and storage architecture, integrate it with several research dashboards and workflows, guide one or two engineers, and ensure the system works reliably for its users. The right candidate should already bring the technical judgment, ownership, and execution expected at this level. The primary learning curve should be OpenAI’s stack and research problem space, not learning how to lead a complex engineering effort or deliver a production system. You will work directly with

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users. That means working across the systems that make Codex actually function as an agent in the real world: the production performance envelope around tokens, latency, reliability, cost, and capacity; the core execution loop and interfaces that turn models into useful behavior; the shared infrastructure that enables other teams to build on Codex; and the feedback loops that turn real-world usage into better models and better agent behavior over time. About the Role We’re looking for engineers to build the infrastructure that powers Codex agents in production. This role focuses on the systems that let models safely execute code, interact with tools, complete long-running tasks, and operate reliably and efficiently at scale. You’ll design and operate the infrastructure behind sandboxed execution, orchestration, stateful workflows, app-server and SDK boundaries, and model rollouts. You’ll work at the intersection of distributed systems, developer tooling, and AI, building primitives that make Codex faster, safer, more reliable, and easier for the rest of the organization to build on. What You’ll Do Design and build execution environments for AI agents, including sandboxing, isolation, and reproducibility. Develop systems for agent orchestration across multi-step, tool-using workflows. Build infrastructure for running, testing, and debugging code generated by models. Create state and memory systems that allow agents to persist context across long-running tasks. Optimize tokens, latency, reliability, and cost across Codex’s production fleet. Support model rollouts, capacity planning, and the core tradeoffs between quality, speed, and economics to manage a fleet of frontier agents at scale. Build shared platform capabilities that unblock product teams, partner teams, and open source Codex. Yo

awsci/cdrest
View job →
O
1mo ago

About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t

awskubernetesrest
View job →

About the Team The Statsig team within OpenAI owns the experimentation, rollout, dynamic configuration, and analytics infrastructure that sits on the launch path for OpenAI products. Our systems help teams ship safely, evaluate product and model changes in production, and make high-confidence decisions from real-world usage. Statsig began as an independent company built around experimentation, feature management, and product analytics at scale. After Statsig joined OpenAI, the team began the next chapter: bringing that platform expertise and infrastructure into OpenAI as the experimentation and rollout foundation for every product we ship. This is infrastructure with a very direct product consequence. Teams working on ChatGPT, Codex, model measurement, consumer experiences including ads, business subscriptions, developer products, and shared platform systems depend on Statsig to evaluate configurations, move traffic safely, ingest experiment data, serve analytics, and roll changes forward or back when production reality demands it. We are at a critical point in the platform journey. Adoption is accelerating quickly across OpenAI, and the systems that were already important are becoming load-bearing for how the company launches. The infrastructure needs to stay fast under sharply increasing evaluation volume, reliable when more services depend on it, observable enough to debug quickly, and efficient enough to support OpenAI-wide scale. Recent SDK and server-side infrastructure work has already produced measurable wins in latency, reliability, memory usage, and compute efficiency across important services. The next phase is to make those gains systematic: a platform that can absorb rapidly growing product velocity while preserving low latency, data quality, operational safety, and developer trust. Based out of OpenAI's Bellevue office, we are a close-knit team that values in-person collaboration, technical depth, operational ownership, and building infrastructure that

vueawsrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The GPT Infrastructure team builds systems that turn advances in model inference and optimization into reliable production capabilities. We enable OpenAI workloads to be qualified and optimized across new accelerator platforms without requiring a one-off port and tuning effort for every hardware target. Our work spans distributed systems, model execution, compilers and runtimes, performance engineering, secure partner integrations, evaluation systems, and developer tooling. We build the infrastructure that makes optimization workflows automated, reproducible, and trustworthy. About the Role We are seeking a software engineer to help build the platform that qualifies and optimizes inference workloads across heterogeneous compute environments. You will develop both OpenAI-hosted services and secure partner-side software for running long-lived optimization workflows. These workflows generate candidate kernels, runtime configurations, and serving-stack changes; compile and execute them on target hardware; verify their correctness; measure their performance; and use the results to guide further optimization. You will work across model architecture, distributed execution, compilers, runtimes, networking, and accelerator systems. A central part of the role is turning research prototypes and one-off hardware bring-up efforts into reliable, reusable infrastructure with clear contracts, reproducible results, strong observability, and well-defined security boundaries. Key Responsibilities Design, build, and operate APIs and control-plane services for long-running workload qualification and optimization campaigns, including scheduling, retries, checkpointing, resource budgets, and observability. Build secure partner-side execution and evaluation software that can compile, run, verify, profile, and benchmark candidate artifacts on accelerator hardware. Integrate model workloads, hardware profiles, compiler toolchains, runtimes, serving engines, and distributed-exe

pythonawslinux
View job →
O
1mo ago

About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling, and infrastructure that help improve model performance across OpenAI’s training and inference workloads at frontier scale. About the Role We’re looking for an autonomous, high-ownership developer productivity engineer who cares deeply about helping other engineers move faster, safer, and with more confidence. This role will sit within OpenAI’s Model Performance organization, contributing to developer infrastructure, CI systems, testing workflows, tooling, and broader performance infrastructure efforts. There is also a strong opportunity to contribute to the Triton project and help improve the systems that support performance-critical engineering work across OpenAI. In this role you will: Improve development workflows for engineers working on model performance infrastructure Design and improve CI/CD, release, validation, and testing pipelines Build and maintain tools that improve reliability, iteration speed, and engineering confidence Partner closely with engineers to identify friction in testing, debugging, deployment, and development workflows Contribute to infrastructure efforts that support performance-critical training and inference systems Help improve developer experience across Python-heavy codebases and performance-oriented infrastructure Work in a high-context, ambiguous environment where ownership and good judgment matter You might thrive in this role if: You are motivated by enabling the people around you and helping engineers do their best work You have strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build/release workflows You are highly collaborative, empathetic, and comfortable partnering deeply with technical teams You are strong in Python and enjoy building reliable, scalable developer tools and infrastructure You have experience improving large-scale engineering work

pythonawsci/cd
View job →

About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches. About the Role In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs. In this role, you will Build and refine performance models that translate microbenchmark results into cost-to-serve estimates. Analyze inference workloads end to end across applications, models, and fleet infrastructure. Enhance tooling to identify bottlenecks across layers for latency and throughput. Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference. You might thrive in this role if you: Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency. Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling. Have deep expertise with performance profiling, benchmarking, analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve o

awsrestai
View job →

About the Team OpenAI’s Platform and Infrastructure Engineering organization advances the mission of deploying artificial general intelligence (AGI) for the benefit of all by delivering secure, scalable, and resilient technology solutions. Our team builds and maintains robust infrastructure that safeguards OpenAI’s data and systems while ensuring employees are well-equipped and seamlessly connected. By prioritizing security, reliability, and user-centric solutions, we empower OpenAI employees to drive impactful AI research, corporate operations, and product innovation. About the Role As a Software Engineer: Internal Applications, Enterprise, you will build internal products that make technology support and administration safer, faster, and less dependent on manual intervention. You will help reduce reliance on broadly privileged human actions, turn recurring technology problems into paved paths, and build agentic systems that can help resolve tickets end to end. A core part of the role is building the interfaces that bring employees, AI agents, and human responders together in a shared ITSM experience, with the right context, controls, and handoffs at each step. We are seeking engineers who enjoy working across frontend and backend layers on ambiguous, high-leverage enterprise problems. You should bring strong product judgment, solid backend engineering fundamentals, and an interest in building software that changes how technology support, system administration, and agent-assisted operations are delivered. The best fit will care as much about the quality of the operator and employee experience as the correctness of the backend systems behind it. In this role, you will: Build frontend experiences that let employees request help, let agents gather context and take safe actions, and let human responders review, approve, or take over without losing the thread. Reduce reliance on broadly privileged manual actions by replacing them with narrow, auditable, policy-aware aut

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time• From $230K/yr
1mo ago

About the Role The Engineering Acceleration Delivery / Continuous Deployment team builds and operates the systems that safely ship OpenAI’s infrastructure and product code to production. We own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across OpenAI to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale infrastructure orchestration. In This Role, You Will Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks. Build systems that support agent-assisted or autonomous deployment workflows using modern AI tooling. Technologies commonly used in this environment include: Kubernetes for large-scale container orchestration and runtime infrastructure Python and FastAPI for internal services Terraform for infrastructure as code GitOps-based deployment workflows (e.g., ArgoCD, Flux, or similar systems) Buildkite for CI orchestration You may be a strong fit if you: Have worked with Kubernetes-based deployment systems at scale Have experience building or operating continuous deployment platforms Are familiar with GitOps tooling such as

pythonawskubernetes
View job →
O
1mo ago

About the Team We’re hiring software engineers to make OpenAI’s networking teams more productive. These teams build and operate the high-performance networking systems that support OpenAI’s training and inference infrastructure at frontier scale. About the Role We’re looking for someone who cares deeply about the developer experience of engineers working on complex infrastructure systems — especially around build systems, test architecture, release pipelines, and reliable development workflows. This role will be embedded with OpenAI’s networking team: making it faster, safer, and easier for engineers to build, test, validate, and ship changes across multi-server, networked, and hardware-adjacent environments. In this role you will: Improve development workflows for engineers building and operating OpenAI’s networking systems Design and improve continuous deployment, release, and validation pipelines Build and maintain test harnesses for multi-server, networked, and hardware-backed environments Improve iteration speed across C++, Python, and build-system-heavy codebases Partner with engineers to identify friction in CI, testing, debugging, and deployment workflows Drive testing and reliability strategy for infrastructure components that support large-scale training and inference workloads Work closely with centralized developer experience teams while staying deeply embedded with the networking engineers closest to the systems You might thrive in this role if: You are motivated by helping other engineers move faster and with more confidence You have experience with CI/CD, release pipelines, testing infrastructure, or build systems You are comfortable moving between C++, Python, and build systems such as CMake, Bazel, or Blaze You enjoy building test harnesses, automation, and workflow improvements for complex systems You do not need to be a networking expert, but you are excited to learn enough about the domain to make the team meaningfully more effective When you see

pythonawsci/cd
View job →
🔔

Get new software engineer snowcommand jobs by email

Daily job updates · Unsubscribe anytime