NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally. What You Will Be Doing: Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software. Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can b
Jobs in United States
System Engineer in United States
5,037 active opportunities · Updated October 2026
Showing
15 jobs
Explore current system engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Expert Structural Analysis Engineer – Systems Stress — USA - North Charleston, Seychelles. Apply via Workday.
About the Role The Engineering Acceleration team builds and operates the foundational systems that engineers use to build, test, and ship ChatGPT, the API, and OpenAI's infrastructure. We are looking for an engineer to help evolve OpenAI's build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, and software quality. You will work on the systems that determine how quickly and confidently engineers can move: Bazel-based builds, Buildkite pipelines, test selection, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to make OpenAI one of the most productive engineering organizations in the world while preserving a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping useful systems instead of fighting infrastructure. In This Role, You Will Own and evolve Bazel-based build and test workflows across a large, polyglot monorepo. Design and maintain Starlark rules, macros, toolchains, and integrations that make builds reproducible, hermetic, and easy for product teams to adopt. Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, test sharding, retry behavior, and flake isolation. Build systems that reduce unnecessary CI work through affected-target detection, dependency graph analysis, test selection, caching, batching, and smarter scheduling. Improve local development workflows so engineers can reproduce CI behavior, debug build failures, and iterate quickly without learning every detail of the build stack. Operate and optimize build infrastructure across Docker/OCI images, Kubernetes-based runners, cloud resources, and remote cache/exec
About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p
From $272K/yr
About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Opportunity: Datadog’s Senior Staff Engineers are technical leaders operating at the forefront of large-scale systems design, building the infrastructure that will support our next five years of growth and beyond. They do this in three major ways: As individual contributors, they bring world-class technical depth to build industry-leading systems in areas such as observability data platforms, distributed query engines, and real-time event streaming at global scale. As technical leaders, they apply broad architectural perspective and deep systems thinking to align design decisions across teams and domains. They work across complex, multi-team problem spaces to define long-term technical direction, drive large-scale initiatives forward, and ensure consistent execution. As engineering stewards, they play a key role in evolving our systems and engineering culture. They actively participate in Datadog’s senior technical community, bringing external insights and internal experience to elevate engineering standards and mentor the next generation of technical leaders. Examples of projects a Senior Staff Engineer may lead include designing and launching a new distributed data storage engine capable of handling hundreds of millions of records per second, building the real-time infrastructure behind a new observability product, or re-architecting a core service to support exponential growth in throughput and complexity. What You’ll Do: Be the technical owner of multiple critical systems or architecture areas, often spanning several t
About the Team The Applied AI team safely brings OpenAI's technology to the world. We released ChatGPT, Plugins, DALL·E, and the APIs for GPT-4, GPT-3, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. We serve end-users directly through ChatGPT, and serve developers through our APIs, which power product features that were never before possible. About the Role The Engineering Acceleration team designs, builds and maintains the foundational systems that engineers use to build ChatGPT and the API. This is a fast-growing team and you will get a chance to own and define the strategy, vision, and plan for how to increase developer productivity. In this role, you will: Drive the design, development, and implementation of tools, systems, and processes that accelerate engineering velocity, reduce manual effort, and increase the quality of output. Use our latest AI tools to re-think how we can be the most productive team in the industry. Work closely with various teams within OpenAI to understand their workflows, challenges, and needs, and ensure the tools and systems built by the Engineering Acceleration team address these requirements. Bring new features and research capabilities to the world by partnering with product engineers to lay the necessary technical foundations. Guide and advise product engineering teams on best practices for ensuring observable, scalable systems. Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years of experience in engineering, including 3+ years of experience in infrastructure building tooling for developers. Have experi
About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About The Role The Security Team is responsible for securing all things Sentry: our customers, our code, and everything in between. We are a small but growing team with broad scope, high trust, and the autonomy to tackle hard security problems with creativity and an engineering mindset. We work at a company with a strong developer culture, building a product that millions of developers genuinely love and rely on. That context shapes everything about how we operate. We take a pragmatic approach to preventing and responding to security risks. In this role not only will you build and contribute to systems which detect malicious activity, you will have the unique opportunity to implement new controls to prevent future incidents. You will work across detection and response and corporate security domains. You'll contribute to practices that keep Sentry secure as we grow: alert triage for corporate and production, detection engineering, deploying preventative controls, identity and access management, investigations and incident response, and more. You'll partner with teams across the company to prevent and respond to security incidents. You will work as a technical collaborator who prioritizes preventative controls, defense in depth, and high signal alerting practices. As Sentry expands our agentic product capabilities and development practices, you'll also find yourself at the frontier of a new set of security approaches and challenges. In this role, you will Maintain, improve, and own detection engineering systems. We own and operate our own detection stack and are building agentic triage with thoughtful security response and orc
From $230K/yr
We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview As the Engineering Manager for the Service Platform team, you will lead the group responsible for Instacart’s core deployment and release engineering systems, including our canary infrastructure and end‑to‑end deployment tooling. This platform is foundational to the reliability of every service at Instacart, and your work will directly shape how developers safely, confidently, and efficiently ship code. You will partner closely with product, infrastructure, and service teams to define and evolve a platform that enables fast iteration without compromising reliability. We are looking for a product-focused engineering manager with a strong interest in developer experience and internal tools, who can guide the team in building intuitive, scalable, and highly resilient systems. You will set the vision, drive execution, and ensure that Instacart engineers
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We’re at the forefront of the data revolution, committed to building the world’s greatest data and applications platform. Snowflake started with a clear vision: develop a cloud data platform that is effective, affordable, and accessible to all data users. Data Cloud allows sharing live data in governed and secure ways so our customers can solve their business problems. Data Cloud allows sharing live data in governed and secure ways so our customers can solve their business problems. Come join our world-class team as a Software Engineer to build the backend services for data cloud to help our customers to share data internally within their company or outside with external organizations. We are investing in initiatives across multiple engineering areas that include: Data Engineering & Open Lakehouse, Database Engineering, Engineering Systems, Product Experiences, Apps & Collaboration,Platform Engineering, and Public Sector, Security and Governance AS A SOFTWARE ENGINEER, YOU WILL: Design and build features, and/or distributed platforms at scale. Drive impactful initiatives for the globally distributed infrastructure Collaborate with product managers, architects, other engineering teams, and business groups, to drive end-to-end solutions. Contribute to improving our en
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Engineering Manager, Cloud Efficiency Snowflake runs large scale cloud infrastructure to deliver its own service — production and internal deployments, Kubernetes fleets, CI/CD, etc. Our cloud spend is in billions of dollars per year. We are looking for an experienced Engineering Manager to lead the Cloud Efficiency engineering team. In this role, you will own the technical vision and execution for building a unified, self-serve cloud efficiency platform along with AI skills and agents that makes resource usage and spend attributable and governable while driving insights and optimization of our cloud spend. AS AN ENGINEERING MANAGER IN CLOUD EFFICIENCY, YOU WILL: Lead and grow our talented team of software engineers, fostering a culture of technical excellence, ownership, and continuous learning. Drive the roadmap for Cloud Efficiency — translating company-level spend objectives into engineering systems: authoritative cost data, resource ownership registry, attribution pipelines, cost and unit economics modeling, observability, governance policies, and optimization workflows — in partnership with Product, Engineering, Finance, and Data Science. Set technical strategy for backend systems, data pipelines, and APIs that measure, attribute and surface cost and usage insights at
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Software Engineer Intern at Roblox, your 12-week journey is designed around accelerated learning and tangible global impact. With the support of a dedicated mentor, you’ll apply your knowledge to production-scale reality as you take end-to-end ownership of a project tackling some of the hardest technical challenges, including distributed systems, real-time communication, 3D co-experience, data processing, rendering, and more. You Will: Join our supportive community of engineers, receiving dedicated mentorship as you deliver live production projects. Investigate and experiment with cutting-edge technologies, such as machine learning frameworks, agentic coding tools, and large language models (LLMs), to solve complex technical challenges and improve our engineering systems. Own a project from beginning to end, from coding and testing to deploying it to production, and presenting your work to peers and leaders across the company. Partner closely with cross-functional teams, including Design, Product, Data, QA, and DevOps, to deliver cohesive products and features. Build a close community with fellow interns and full-time builders alike, while gaining a broad perspective on the compa
From $192K/yr
As the Senior Product Manager for the Actions & Automations team, you will own the ecosystem that enables customers, partners, and Datadog teams to build, deploy, and operate AI agents on Datadog. You will drive the strategy and execution for the platform capabilities, developer experience, integrations, and extensibility model that make Datadog the best place to build agents that understand and act on production systems. Modern engineering organizations are entering a new era where software is not only monitored and operated by humans, but increasingly by AI-powered agents. As agentic workflows reshape how teams build, operate, secure, and troubleshoot systems, customers need a platform for creating specialized agents, connecting them to business and engineering systems, governing their behavior, and extending them to solve unique organizational problems. You will define and build the ecosystem that makes this possible. Agent Builder sits at the intersection of Datadog's products, AI capabilities, and ecosystem strategy. You will have the opportunity to work across the breadth of the Datadog platform, partner with teams throughout the company, and help establish Datadog as the foundation for operational AI. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Define the vision, strategy, and roadmap for Datadog's Agent Builder platform and ecosystem. Own the core platform capabilities that enable customers and partners to create, customize, deploy, and manage AI agents. Drive the extensibility model for agents, including integrations, tools, actions, context sources, APIs, SDKs, and developer workflows. Shape how agents perform actions across Datadog products and third-party systems. Partner closely with AI, platform, infrastructure, and product tea
Datadog's Software Engineers with Systems depth leverage their experience with systems and tooling to build software that ensures Datadog remains reliable, performant, and secure. For this track, their Software Engineering experience may resemble the Distributed Systems track, but is typically applied in combination with their systems experience to build and run internal platforms and tools that our products are built on. These people typically have deep experience building and managing large cloud infrastructure deployments, or leading reliability efforts for orgs similar to ours, or building release machinery to allow hundreds or thousands of devs to do their jobs without stepping on each others' toes. The systems and tooling where they may have experience depth may include (but not limited to): bazel, build tooling, cassandra, CDN, chef, configuration management, container orchestration, consul, docker, elasticsearch envoy, haproxy, kafka, kubernetes, load balancing, network architecture, postgres, redis, release management, RPC frameworks, service discovery, spinnaker, terraform, zookeeper. Bonus: You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. #LI-KM5 Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan. Th
From $399.4K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Why Roblox Studio? Roblox Studio is the creation engine behind millions of immersive 3D experiences built by creators around the world. It serves everyone from first-time developers to large professional studios, combining a powerful real-time engine with a deeply expressive toolchain. As creation on Roblox grows more sophisticated, Studio must continue evolving as a world-class development environment for interactive 3D experiences. That requires deep investment in the technical systems that underpin the IDE experience from scripting, debugging, architecture, and core IDE capabilities including performance. We are looking for a leader who can raise the bar on these foundational systems and help shape the future of Studio as a high-quality, extensible, and performant creation environment. Familiarity with creator tools is important, and experience with AI-assisted creation is a plus. The Role We are seeking a Director of Engineering for Studio Systems to lead a critical area of Roblox Studio’s technical foundation. In this role, you will be accountable for the architecture, execution, and long-term evolution of core systems that power Studio’s IDE and developer workflows. This includes area
Other cities to consider
More places hiring for this role
Get new system engineer jobs in United States by email
Daily job updates · Unsubscribe anytime