Jobs in United States

Performance And Systems Engineer in United States

2,914 active opportunities · Updated October 2026

Explore current performance and systems engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Workload team is responsible for designing and running OpenAI’s LLM training and inference infrastructure that powers frontier models at massive scale. Our systems unify how researchers train and serve models, abstracting away the complexity of performance, parallelism, and execution across vast GPU/accelerator fleets. By providing this foundation, the Workload team ensures that researchers can focus on advancing model capabilities while we handle the scale, efficiency, and reliability required to bring those models to life. About the Role We are looking for an engineer to design and implement the dataset infrastructure that powers OpenAI’s next-generation training stack. You will be responsible for building standardized dataset interfaces, scaling pipelines across thousands of GPUs, and proactively testing performance bottlenecks. In this role, you will collaborate closely with the multimodal researchers, and other infra groups to ensure datasets are unified, efficient, and easy to consume. In this role, you will: Design and maintain standardized dataset APIs, including for multimodal (MM) data that cannot fit in memory. Build proactive testing and scale validation pipelines for dataset loading at GPU scale. Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines, ensuring smooth adoption and a great user experience. Document and maintain dataset interfaces so they are discoverable, consistent, and easy for other teams to adopt. Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged once standardized. Debug and resolve performance bottlenecks in distributed dataset loading (e.g., straggler systems slowing global training). Provide visualization and inspection tools to surface errors, bugs, or bottlenecks in datasets. You might thrive in this role if you: Have strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.

AWSRestAIRust
H
📍 Colorado, United States of America, United States
✓ High-confidence listingCompany trend +103.7%

$59.4K – $89.6K/yr

Quick readStrong listing-quality and freshness signals

Software Quality Engineer Description - This role is responsible for maintaining the quality, reliability, and performance of software applications throughout the development lifecycle. The role identifies and rectifies defects, ensures adherence to established quality standards, and contributes to the overall improvement of the software development process. The role involves various activities aimed at preventing and detecting issues, thereby enhancing the end user experience. The role creates and executes comprehensive test plans, test cases, and test scripts based on project specifications. *Onsite in Ft. Collins 5-days a week Responsibilities • Executes established test plans and protocols for assigned portions of code for end-user applications, systems software, and firmware running on hardware, local, networked, and Internet- based platforms; identifies, logs, and debugs assigned issues. • Perform Functional and Solution Testing of Video/Collaboration Software • Additionally, codes and programs test scripts, automation, and integration activities based on specific test requirements. • Conducts functional, integration, regression, and performance testing to validate software functionality. • Automates testing processes using appropriate tools and frameworks to improve efficiency and repeatability. • Monitors and enforces adherence to established coding standards, design guidelines, and best practices. • Monitors software performance and conducts load and stress testing to identify bottlenecks and performance issues. • Prepares and maintains QA-related documentation, including test plans, test matrices, and testing reports. • Develops understanding of and relationship with internal and outsourced development partners on software applications design and development. • Participates as a member of project

PythonAIJenkins
H
📍 Colorado, United States of America, United States
✓ High-confidence listingCompany trend +103.7%

$59.4K – $89.6K/yr

Quick readStrong listing-quality and freshness signals

Software Quality Engineer Description - This role is responsible for maintaining the quality, reliability, and performance of software applications throughout the development lifecycle. The role identifies and rectifies defects, ensures adherence to established quality standards, and contributes to the overall improvement of the software development process. The role involves various activities aimed at preventing and detecting issues, thereby enhancing the end user experience. The role creates and executes comprehensive test plans, test cases, and test scripts based on project specifications. *Onside in Ft Collins 5-days a week Responsibilities • Executes established test plans and protocols for assigned portions of code for end-user applications, systems software, and firmware running on hardware, local, networked, and Internet- based platforms; identifies, logs, and debugs assigned issues. • Perform Functional and Solution Testing of Video/Collaboration Software • Additionally, codes and programs test scripts, automation, and integration activities based on specific test requirements. • Conducts functional, integration, regression, and performance testing to validate software functionality. • Automates testing processes using appropriate tools and frameworks to improve efficiency and repeatability. • Monitors and enforces adherence to established coding standards, design guidelines, and best practices. • Monitors software performance and conducts load and stress testing to identify bottlenecks and performance issues. • Prepares and maintains QA-related documentation, including test plans, test matrices, and testing reports. • Develops understanding of and relationship with internal and outsourced development partners on software applications design and development. • Participates as a member of project t

PythonAIJenkins
M
📍 California, United States of America, United States
✓ High-confidence listingCompany trend +1850%
Quick readStrong listing-quality and freshness signals

We anticipate the application window for this opening will close on - 10 Oct 2026 Careers that change lives start here. Medtronic is a global leader in healthcare technology with a Mission to alleviate pain, restore health, and extend life. Our 95,000 employees work across more than 150 countries to put patients first — developing innovative medical technologies that improve the lives of 72+ million patients each year. Your unique talents will help shape the future of healthcare while building a career grounded in purpose, growth, and impact. A Day in the Life The Neuromodulation Research & Development organization develops therapies and technologies that address chronic pain and other neurological conditions. Within the Interventional Pain portfolio, teams focus on minimally invasive therapies used to treat musculoskeletal pathologies, including vertebral compression fractures, nerve pain, metastatic bone tumors, and benign bone tumors. This role supports the development of system-level verification and validation strategies for new products and enhancements across mechanical, electrical, and software components. This position is based in Santa Clara, California, and follows an onsite work model. No travel is required for this role. As a Research & Development Test Engineer, you will lead and execute verification and validation activities that support the development of interventional pain therapies and associated medical device systems. You will collaborate with cross-functional partners to define test strategies, evaluate product performance, manage technical risks, and generate evidence needed to support product development and regulatory requirements. Primary Responsibiliti

RecruitmentHR
M
📍 Colorado, United States of America, United States
✓ High-confidence listingCompany trend +1850%
Quick readStrong listing-quality and freshness signals

We anticipate the application window for this opening will close on - 30 Sep 2026 Careers that change lives start here. Medtronic is a global leader in healthcare technology with a Mission to alleviate pain, restore health, and extend life. Our 95,000 employees work across more than 150 countries to put patients first — developing innovative medical technologies that improve the lives of 72&#43; million patients each year. Your unique talents will help shape the future of healthcare while building a career grounded in purpose, growth, and impact. A Day in the Life The Acute Care & Monitoring (ACM) business develops technologies that help clinicians monitor, assess, and manage patients across the continuum of care. Within ACM, the Test Engineering function supports product development, verification, reliability, and manufacturing readiness through the design and execution of test methods, fixtures, and equipment used to evaluate product performance and system functionality. This team partners closely with systems, mechanical, electrical, software, quality, and manufacturing engineering teams to support the development and sustainment of medical technologies. This position is based in Lafayette, Colorado and is an onsite role. Travel up to 10% may be required to support testing activities, cross-functional collaboration, and business needs. As a Mechanical R&D Engineer II, you will support system-level, reliability, mechanical, and electrical testing activities for Acute Care & Monitoring products. This role combines hands-on laboratory testing, fixture development, data analysis, and cross-functional collaboration to evaluate product performance, investigate issues, and support verification activities throughout the product lifecycle. Primary Responsibilities <li

PythonRecruitmentHR
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%
Quick readStrong listing-quality and freshness signals

As a Software Engineer - Frontend on a feature team, you'll be responsible for building an intuitive, responsive product. Your work will help thousands of customers to monitor the health and performance of their systems, no matter the scale. In this capacity, you’ll use tools like React and Typescript to help streamline the flow of data from platform to user, enable seamless pivots from one view to the next, and build a powerful yet easy-to-use platform. You will work on complex frontend challenges for thousands of customers with huge sets of data and work on exciting scalability and performance challenges. Our users are also developers so you will feel close to the product and have an impact on the development world. You will be part of a front-end community of 200+ passionate frontend engineers and you will be surrounded by experts. Join us to build the next generation of high-scale, data-powered features for our customers. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Work closely with backend engineers, product managers and designers to inform/refine/validate design concepts Break down complex features into distributable engineering tasks and rollout plans Design and deliver features that have direct impact on thousands of users Experiment with and advocate for new systems, design patterns, and tooling Participate in hackathons, sync with other Front end engineers at the monthly demos meeting and yearly Front End summit Use and contribute to a best-in-class in-house design system Own meaningful parts of our service, improve performance and address scalability limits Work in a fast-paced, high-growth environment that values diversity of talent, excellence of product, and exciting engineering challen

JavaScriptTypeScriptJavaReact
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are looking for a highly experienced RTL engineer to own critical on- and off-chip interconnect components for our custom AI accelerator platform. You will drive the microarchitecture and RTL implementation of scalable on-chip communication fabrics connecting high-bandwidth compute, memory, and I/O subsystems as well as purpose-built off-chip interfaces and protocols needed to enable custom computing at scale. This is a senior, hands-on engineering role with broad technical ownership. You will drive design from requirements through the full silicon lifecycle, from architecture definition and performance analysis through RTL implementation, verification closure, physical design convergence, bring-up, and production readiness. You will plan and oversee the work of junior engineers and help drive and develop productive engineering relationships with external partners and help manage partner execution. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the microarchitecture, RTL design, and delivery of major SoC interconnect components, including network-on-chip fabrics, switches, routers, bridges, protocol adapters, arbiters, and traffic-management logic as well as off-chip protocol bridges and interfaces. Drive third party engagements to develop novel networking and interface protocols and silicon IP while ensuring high quality and de

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role On the Accelerators team, you will help OpenAI evaluate and bring up new compute platforms that can support large-scale AI training and inference. Your work will range from prototyping system software on new accelerators to enabling performance optimizations across our AI workloads. You’ll work across the stack, collaborating with both hardware and software aspects - working on kernels, sharding strategies, scaling across distributed systems, and performance modeling. You'll help adapt OpenAI's software stack to non-traditional hardware and drive efficiency improvements in core AI workloads. This is not a compiler-focused role, rather bridging ML algorithms with system performance - especially at scale. In this role, you will: Prototype and enable OpenAI's AI software stack on new, exploratory accelerator platforms. Optimize large-scale model performance (LLMs, recommender systems, distributed AI workloads) for diverse hardware environments. Develop kernels, sharding mechanisms, and system scaling strategies tailored to emerging accelerators. Collaborate on optimizations at the model code level (e.g. PyTorch) and below to enhance performance on non-traditional hardware. Perform system-level performance modeling, debug bottlenecks, and drive end-to-end optimization. Work with hardware teams and vendors to evaluate alternatives to existing platforms and adapt the software stack to their architectures. Contribute to runtime improvements, compute/communication over

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a software engineer on the Scaling team, you’ll help build and optimize the low-level stack that orchestrates computation and data movement across OpenAI’s supercomputing clusters. Your work will involve designing high-performance runtimes, building custom kernels, contributing to compiler infrastructure, and developing scalable simulation systems to validate and optimize distributed training workloads. You will work at the intersection of systems programming, ML infrastructure, and high-performance computing, helping to create both ergonomic developer APIs and highly efficient runtime systems. This means balancing ease of use and introspection with the need for stability and performance on our evolving hardware fleet. This role is based in San Francisco, CA, with a hybrid work model (3 days/week in-office). Relocation assistance is available. In this role, you will: Design and build APIs and runtime components to orchestrate computation and data movement across heterogeneous ML workloads. Contribute to compiler infrastructure, including the development of optimizations and compiler passes to support evolving hardware. Engineer and optimize compute and data kernels, ensuring correctness, high performance, and portability across simulation and production environments. Profile and optimize system bottlenecks, especially around I/O, memory hierarchy, and interconnects, at both local and distributed scales. Develop simulation infrastructure to validate runtime b

PythonAWSRestAI
O
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Astral team builds high-performance developer tools to power the future of programming, at OpenAI and beyond, including Ruff, uv, and ty. The Astral toolchain sees hundreds of millions of installs per month and powers hundreds of millions of package downloads per day for the Python ecosystem. As a team, we are building on those foundations to continue solving impactful tooling problems as programming evolves. About the Role We are looking for an experienced software engineer to build next-generation programming language tooling. If you like writing high-performance Rust, it could be a good fit; if you like thinking about the future of programming, it could also be a good fit. Strong candidates tend to have deep experience with Rust, Python, open source, compilers, or developer tools — but few candidates are deep in all of these areas, and we've hired candidates without prior Rust or Python experience. In this role, you will: Design and implement features in Astral’s existing open source projects (Ruff, uv, ty, and python-build-standalone, and more). Support Astral’s open source projects as a maintainer, triaging user issues, reviewing pull requests, and participating in community discussions. Evolve the Astral toolchain to accelerate development velocity at OpenAI. Build entirely new tools, in entirely different programming ecosystems, to power the future of agentic software development. Your background might look something like: 5+ years of professional engineering experience, excluding internships, in relevant engineering roles. High agency and comfort operating in a fast-moving environment, with strong ownership of security, reliability, and operational excellence. Strong developer empathy and communication skills, including experience maintaining open source projects. Exceptional systems engineering fundamentals and a track record of leading complex projects from ambiguous problem statements through to user impact. Proficiency in one or more s

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Infrastructure Quality team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, general contractors, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. Our work spans from vendor qualification through commissioning, ensuring operational readiness across our global portfolio. About the Role We are seeking an experienced Manufacturing Quality Engineer (MQE) to establish, implement, and manage a manufacturing-focused quality program for datacenter infrastructure. This role will be responsible for vendor oversight, quality assurance, process improvement, and issue resolution for all critical systems. You will lead vendor audits, monitor performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced risks, and operational reliability. By partnering with vendors, construction teams, and internal stakeholders, you will help ensure OpenAI’s datacenters are delivered on time and built to the highest operational standards. Travel Domestic and international travel as needed (estimated 40–60%) to manufacturing sites, datacenter locations, and partner facilities. Key Responsibilities Vendor Oversight & Performance Management Conduct manufacturing evaluation, audits, and improve vendor performance across production, inspection, testing, and delivery phases. Develop and track quality metrics to assess manufacturing performance and identify trends. Partner with vendors to refine processes, training, and quality controls to mitigate risks before shipment. Program Development & Execution Develop and maintain a datacenter-focused m

AWSRestAIGo
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $192K/yr

Quick readStrong listing-quality and freshness signals

About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. You will: Solve a scaling bottleneck in a critical service Deploy a new feature to production, progressively rolling it out with feature flags Investigate and fix a production issue from a service your team owns Design a way to scale up a service for more traffic With your team, plan the most important projects to work on next Who You Are: You have significant experience in one or more languages You value code simplicity and performance You can design architecture to solve problems at high scale You have a BS/MS/PhD in a scientific field or equivalent experience You want to work in a fast, high-growth startup environment that respects its engineers and customers You have demonstrated ability to use AI coding tools in day-to-day workflows and build, validate, and refine AI-generated output in products You can design AI Backend systems, with awareness of quality, cost, and latency tradeoffs 6+ years of experience Bonus points: You've worked at high scale with systems like Redis, Cassandra, Kafka You wrote your own data pipelines once or twice before You have a strong background in statistics You have significant experience with Go, C, or Python You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how You’re motivated to push the boundaries of how AI can improve software engineering best practices and contribute to building AI-enabled products Datadog values people from all walks of life. We understand not everyone will meet all the above qualificat

PythonRedisAIGo
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $234K/yr

Quick readStrong listing-quality and freshness signals

The ML Observability team builds cutting-edge tools to monitor, explain, and improve AI systems in production, particularly those leveraging Large Language Models (LLMs) and generative AI. We provide robust, scalable observability for AI workloads, including drift detection and model evaluation, and behavior tracing, enabling customers to ship AI with confidence. As a Staff Engineer, you’ll lead the development of new features and foundational capabilities within Datadog’s LLM Observability product. You will shape product direction, drive experimentation, and apply your deep understanding of both AI systems and software engineering to solve open-ended problems in the fast-moving AI landscape. Your work will directly impact how our customers monitor, troubleshoot, and optimize LLM-based applications in production. Join us in building the foundational tools that make AI systems observable, understandable, and reliable in the real world. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Drive design and implementation of LLM observability features. Ideate, prototype, and scale new product features to provide insights and drive improvements for generative AI systems Work cross-functionally with other eng teams, product, UX, and applied science to iterate fast and find product-market fit Develop and extend tools for tracing, evaluating, and debugging LLMs Influence architecture decisions and mentor engineers to build resilient, high-performance systems Stay close to customer pain points and use those insights to guide product and engineering priorities Stay current with industry trends and advancements in machine learning and observability, driving innovation within the team Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or r

Machine LearningAIGoRust
O
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team OpenAI’s Forward Deployed Engineering team partners with customers to turn research breakthroughs into production systems. We operate at the intersection of customer delivery and core platform development. About the role As a foundational FDE manager, you’ll lead FDE through high-stakes, ambiguous customer deployments and own technical and business value outcomes end to end. You’ll grow a team that can operate under pressure and help OpenAI learn from the field. You’ll partner closely with Product, Research, Sales, and GTM to ensure fieldwork informs roadmap priorities, drives new exploration, and supports safe deployment at scale. Your decisions will influence how OpenAI is trusted by the customers closest to our deployment work. Your success will be measured by how consistently your team ships, how clearly you deliver signal to Research and Product, and how durable your team and delivery model prove to be. This role is based in New York City We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. This role also will require travel up to 25%. In this role you will Lead and grow a team of FDE delivering production systems with frontier models Own end-to-end delivery outcomes through clarity, speed, tight coordination, and technical quality Codify what works into tools, playbooks, and roadmap inputs that create leverage for both OpenAI and our wider developer community Notice early indicators and raise them with urgency, whether in product behavior, customer environments, or delivery practices Use judgement to distinguish what requires action and what does not Set a high bar for FDE performance and support each person’s growth through direct, actionable feedback Define how we staff and support field teams that can scale without added complexity You might thrive in this role if you Bring 8+ years of engineering or technical delivery experience, including 2+ years managing high-performing FDE or custo

JavaScriptPythonJavaAWS
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We’re looking for a GPU Inference Engineer to contribute to improvements in model serving efficiency for our Robotics research. This is a high-impact role where you’ll drive initiatives to optimize inference performance and scalability. You’ll also be engaged in model design, to help assist our researchers in developing inference-friendly models. This role is critical to scaling the team’s broader goals - it will directly enable leadership to focus on higher-leverage initiatives by building a stronger technical foundation. In this role you will: Perform engineering efforts focused on improving model serving, inference performance, and system efficiency Drive optimizations from a kernel and data movement perspective to improve system throughput and reliability Partner closely with research and product teams to ensure our models perform effectively at scale Design, build, and improve critical serving infrastructure to support Robotics growth and reliability needs You might thrive in this role if you: Have deep expertise in model performance optimization, particularly at the inference layer Have a strong background in kernel-level systems, data movement, and low-level performance tuning Are excited about scaling high-performing AI systems that serve real-world, multimodal workloads Can navigate ambiguity, set technical direction, and drive complex initiatives to completion This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. About OpenAI OpenAI i

AWSRestAIGo
🔔

Get new performance and systems engineer jobs in United States by email

Daily job updates · Unsubscribe anytime