Jobs in United States

Workload Porting And Performance Engineer in United States

382 active opportunities · Updated October 2026

Explore current workload porting and performance engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team The Workload team is responsible for designing and running OpenAI’s LLM training and inference infrastructure that powers frontier models at massive scale. Our systems unify how researchers train and serve models, abstracting away the complexity of performance, parallelism, and execution across vast GPU/accelerator fleets. By providing this foundation, the Workload team ensures that researchers can focus on advancing model capabilities while we handle the scale, efficiency, and reliability required to bring those models to life. About the Role We are looking for an engineer to design and implement the dataset infrastructure that powers OpenAI’s next-generation training stack. You will be responsible for building standardized dataset interfaces, scaling pipelines across thousands of GPUs, and proactively testing performance bottlenecks. In this role, you will collaborate closely with the multimodal researchers, and other infra groups to ensure datasets are unified, efficient, and easy to consume. In this role, you will: Design and maintain standardized dataset APIs, including for multimodal (MM) data that cannot fit in memory. Build proactive testing and scale validation pipelines for dataset loading at GPU scale. Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines, ensuring smooth adoption and a great user experience. Document and maintain dataset interfaces so they are discoverable, consistent, and easy for other teams to adopt. Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged once standardized. Debug and resolve performance bottlenecks in distributed dataset loading (e.g., straggler systems slowing global training). Provide visualization and inspection tools to surface errors, bugs, or bottlenecks in datasets. You might thrive in this role if you: Have strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.

AWSRestAIRust
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $326.1K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Security Software Engineer on the Production IAM team, you will set the technical direction for how identity and access work across Roblox's production infrastructure, from the mTLS-based identity that services use to authenticate to one another, to the privileged access controls that govern how engineers reach production. The team is accountable for Roblox's machine and workload identity platform, its centralized authorization engine, its production access management platform, production PKI and certificate lifecycle, and just-in-time privileged access for engineers. As an individual contributor in Production IAM, you will define multi-year strategy, drive alignment across Roblox Platform, mentor senior and staff engineers, and personally build the hardest parts of these systems. As AI agents become first-class actors in production, you will also help pioneer how they get identity, prove who they are, and receive safely-scoped access. You will Lead the architecture for production identity and access. Define and evolve the end-to-end design for machine, workload, human, and AI-agent identity across our hybrid on-prem and cloud fleet, making secure access invisible when

PythonJavaAWSGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team We’re hiring software engineers to make the Workload team more productive. The Workload team maintains the core components of OpenAI’s training and inference frameworks and helps execute frontier experiments. About the Role We’re looking for someone who cares about the developer experience of working in and around OpenAI’s core training and inference frameworks. In this role you will: Be responsible for optimizing the development workflows of the engineers around you Work within various Workload teams to address their specific needs, but collaborate with the centralized teams that own various aspects of development experience Optimize iteration speed, both broadly, and in particular by optimizing specific teams’ CI Improve reliability, for instance, by driving testing strategy for particular components Work through the long tail of things that it takes to build libraries and systems that will delight researchers You might thrive in this role if: You are motivated by helping people. You believe a thing that separates great teams from good teams are the players willing to do whatever work it takes, without ego. You believe in the power of developer experience. Something magical happens when people can quickly and confidently iterate on a simple codebase, but this magic is fragile and must be fought for. When you see someone trip over something, no matter how small, your first instinct is asking yourself what it would take for that to not happen again. Your second instinct is clicking merge on the PR you’ve already written to make it so. You are pragmatic. You have the ability to see the world through a perfectionist’s eyes, but are not yourself a perfectionist. You know which problems to pick and when to switch to making progress on a different problem. You like going end-to-end on things. You love co-design — that feeling when you were only able to find the right solution because you both deeply understand the users that interact with a system and the

PythonAWSRestAgile
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA silicon runs the world's AI infrastructure. The frequency it delivers across every voltage, process corner, and workload is not assumed. It is measured, correlated, and validated. This role does that work. The Silicon Co-Design Group is where architecture intent becomes silicon reality. We own the boundary between what was designed and what was built, and we are the team that knows the difference. When a program ships at frequency and at quality, this team is a reason why. You will be the person who follows through between simulation and silicon. When the model is wrong, a frequency corner that doesn't hold, a Vmin that walks, a critical path that timing analysis missed, you find out why, and your data is what the rest of the program acts on. Architecture, design, and product teams do not guess. They use your numbers. The engineers who do this well are rare. They think like circuit designers, work like experimentalists, and reason like data scientists. If that is you, read on. What you'll be doing: Own silicon speed characterization from first power-on through production sign-off, covering frequency, Vmin, Vmax, and timing margins across the full PVT space. Close the correlation gap. Tie pre-silicon timing analysis and critical path predictions to measured silicon, quantify where the model diverges from reality, and produce analysis that architecture and design can act on with confidence. Trace failures to their source, whether a microarchitectural bottleneck, a critical path that doesn't close under voltage, a clocking issue, or a process corner the model didn't anticipate, and drive the resolution. Own and build AI agents that work at your direction: automated test orchestration, intelligent data pipelines, and analysis flows that expand coverage and compress cycle time without sacrificing difficulty. Know where AI accelerates real work and whe

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -83.5%

From $113K/yr

Quick readStrong listing-quality and freshness signals

As a Security Sales Specialist, you'll partner with Enterprise Account Executives to drive adoption of Datadog’s Security platform across key accounts. This is a high-impact role focused on positioning our Security solutions (Cloud SIEM, Cloud Workload Security, CSPM, and more) into new and existing customers—expanding our footprint and helping customers modernize their security stack in the cloud. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as the subject matter expert (SME) for Datadog Security products across a targeted account patch Collaborate closely with Enterprise AEs to support net new logo acquisition and expansion in strategic accounts Own and drive the security sales cycle from discovery to technical close, working closely with Sales Engineers Evangelize Datadog’s security story to security leaders (CISO, Security Architects, SecOps) Work cross-functionally with Datadog's partner, channel, and alliance teams to drive joint go-to-market motions Co-sell effectively with AEs and partners, contributing to deal strategy, solution alignment, and stakeholder engagement Stay informed on security trends and competitive offerings to differentiate Datadog Who You Are: Proven success selling into security buyers (CISO, SecOps, GRC, etc.) Experience co-selling in a matrixed environment, supporting or partnering with AEs and cross-functional teams Strong understanding of the partner/channel sales model, including how to navigate and influence joint selling motions Familiarity with modern security solutions such as SIEM, CSPM, CWPP, container/Kubernetes security Ability to build strong relationships with internal stakeholders, partners, and customer technical teams Datadog values people from all walks of life. We understand not everyone will meet al

KubernetesAIGoRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team The GPT Infrastructure team builds systems that turn advances in model inference and optimization into reliable production capabilities. We enable OpenAI workloads to be qualified and optimized across new accelerator platforms without requiring a one-off port and tuning effort for every hardware target. Our work spans distributed systems, model execution, compilers and runtimes, performance engineering, secure partner integrations, evaluation systems, and developer tooling. We build the infrastructure that makes optimization workflows automated, reproducible, and trustworthy. About the Role We are seeking a software engineer to help build the platform that qualifies and optimizes inference workloads across heterogeneous compute environments. You will develop both OpenAI-hosted services and secure partner-side software for running long-lived optimization workflows. These workflows generate candidate kernels, runtime configurations, and serving-stack changes; compile and execute them on target hardware; verify their correctness; measure their performance; and use the results to guide further optimization. You will work across model architecture, distributed execution, compilers, runtimes, networking, and accelerator systems. A central part of the role is turning research prototypes and one-off hardware bring-up efforts into reliable, reusable infrastructure with clear contracts, reproducible results, strong observability, and well-defined security boundaries. Key Responsibilities Design, build, and operate APIs and control-plane services for long-running workload qualification and optimization campaigns, including scheduling, retries, checkpointing, resource budgets, and observability. Build secure partner-side execution and evaluation software that can compile, run, verify, profile, and benchmark candidate artifacts on accelerator hardware. Integrate model workloads, hardware profiles, compiler toolchains, runtimes, serving engines, and distributed-exe

PythonAWSLinuxRest
C
📍 East Lansing, United States
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary The Front End Operations Supervisor is responsible for leading a team of approximately 20 employees and overseeing the day-to-day operations of mail intake, document processing, and workflow management. This role provides direct supervision, coaching, and performance support through regular one-on-one meetings, team communications, and operational oversight to ensure service levels and business objectives are met. Key responsibilities include managing the daily workload associated with in-house printing, returned mail processing, document scanning, check handling, and mail indexing activities. The Supervisor is responsible for monitoring work queues, balancing resources, addressing operational issues, and ensuring timely and accurate processing of all incoming and outgoing work. This person serves as a primary liaison with the mailroom vendor, coordinating the receipt, processing, and distribution of incoming mail. The Supervisor partners with internal and external stakeholders to identify process improvement opportunities, resolve service issues, and ensure operational efficiency across all front-end mail operations. Additionally, this role supports departmental initiatives and special projects, assists with process documentation and workflow enhancements, and contributes to the implementation of operational improvements designed to increase productiv

ExcelCustomer Service
H
📍 Dallas, TX, United States
✓ High-confidence listingCompany trend +310%
Quick readStrong listing-quality and freshness signals

Become a part of our caring community We are looking for a highly motivated Senior Technology Leadership professional to join our IT Operations team. You will support the Lead of the Application Operations Center and Enterprise Post Production Validation teams, with a focus on process automation, innovation, and continuous improvement. You will work with both onshore and offshore resources, ensuring in daily tasks, enhancing application support, and driving improvements in monitoring and validation processes. Main Responsibilities: Collaborate with the Lead to provide operational and strategic support for Application Operations Center and Post Production Validation teams. Identify, evaluate, and implement automation opportunities to increase efficiency and reduce manual workload. Drive process innovation by recommending and deploying advanced tools and methodologies for application operations and validation activities. Analyze existing workflows and develop documentation for standard operating procedures and best practices. Oversee daily application support and monitoring activities to ensure system stability, performance, and reliability. Partner with onshore and offshore teams to coordinate task execution and promote consistent adoption of new processes and technologies. Develop and maintain dashboards and reports to track key performance indicators and present findings to leadership. Ensure automation and process improvements comply with organizational standards and regulatory requirements. Facilitate knowledge sharing, training sessions, and change management activities to support team development and successful project implementation. Engage with stakeholders to gather requirements, understand challenges, and communicate progress on automation initiatives. Ability to create

PythonJavaAzureAI
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. NVIDIA GH200 superchip provides performance and productivity required for strong scaling for HPC and generative AI workload. Scale out is inherent to design of this massive superchip. We are looking for expert engineers to come and help design rack level solutions for next generation scaling AI supercomputing platforms. We are looking for a strong technical architect to own end to end manageability architecture for these products in data centers. You will work with various component leads internally and externally, drive customer use cases, align architecture with customer requirements and release best products to market. Join us at the forefront of technological advancement. What you’ll be doing: Drive server management for large clusters and data centers deploying GPUs and Grace solution from Nvidia. Work with data center architects and cloud customers to narrow down on requirements for implementation to ensure speed of light product development. Work with internal teams to make sure requirements are designed and implemented in right way with each firmware and software module Collaborate with other leads to design & build data center health management workflow. Drive reliability and optimization in firmware architecture from a data center view point. Work closely with cluster bring up team and resolve is

PythonGitAIProject Management
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA has been redefining computer graphics, desktop gaming, and enhanced computing capabilities for more than 25 years. Today, we are tapping into the unlimited potential of AI to define the next era of computing. As a NVIDIAN, you will work on problems that sit at the boundary of architecture, silicon, firmware, software, and production, where strong judgment matters as much as technical depth. We're the Silicon Design for Productization (DFP) Team, within the broader Silicon Co-Design Group, and we turn power and thermal design into executable productization methodology. Power and thermal are among the most complicated problems we work on at NVIDIA because they sit at the intersection of architecture, workload behavior, silicon variation, firmware policy, platform constraints, and product goals. Small decisions here have an outsized impact on performance, efficiency, reliability, bring-up speed, and ultimately what the product can deliver in the field. We define how features move from concepts to bring-up, characterization, validation, and release. In this role, you will help us build that bridge. We're looking for an engineer who reasons from first principles, flourishes with ownership in a fast-paced environment, and uses AI with sound judgment. What you’ll be doing: Lead the effort across multi-functional teams to keep the program’s power and thermal productization strategy clear, executable, and on track. Create methodology and silicon test plan based controller designs and architecture, including characterization process, debug tools, fuse/firmware settings and lab requirements. Drive resolution for challenging silicon issues through structured hypotheses, measurement plans, and root-cause closure. Steward the Power and Thermal playbook when the existing productization methodology

N
📍 Remote, United States· Remote
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA DGX Cloud is an AI Factory designed to power the next generation of AI and industrial-scale breakthroughs. As the Distinguished Engineer for Security Architecture, within our Security Engineering organization, you will set the security design bar for an AI factory of hundreds of thousands of GPUs, and then build against it alongside the teams. This is the founding architecture seat in a new organization. Security Engineering is a new organization at DGX Cloud, accountable for the security outcome of the platform, and this is the architecture function inside it. You will define the security design standard for DGX Cloud, a bar that sits above the company floor, and hold it from inside the teams doing the building. Security here is fleet horizontal and stack vertical, so your scope runs from the hardware root of trust and the hardened baseline, through tenancy and GPU workload isolation, to the services and APIs built on top, across every DGX Cloud engineering organization. A small team of Principal Engineers will report to you and hold the bar at domain depth. This is still a hands-on seat, and you stay in the design with them. You will also serve as DGX Cloud's technical interface into NVIDIA's central security organization. There is no architecture review board here and no approval queue; the bar holds because the strongest security engineers in the room helped set it and helped ship it. What You Will Be Doing: Set the DGX Cloud Security Bar: Own the security design standard across DGX Cloud (tenancy, GPU workloads, identity, supply chain, and isolation) and make it concrete. Reference architectures, golden paths, and requirements engineers can actually build against, not a policy library. Hold the Bar by Building: Embed with engineering teams on real work: join the design, learn the code, help ship the thing rather than grade it afterward.

KubernetesLinuxArtificial IntelligenceAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI's Industrial Compute organization builds and operates the infrastructure required to train and serve frontier AI models. The Capacity Planning team connects rapidly changing research and product demand with the compute, networking, storage, power, data center, hardware, and operational resources required to make that demand executable. About the Role We are seeking a Technical Program Manager to build and lead capacity planning across OpenAI's large-scale AI infrastructure. You will translate uncertain workload demand into clear infrastructure requirements, allocation decisions, supply commitments, activation priorities, and long-range capacity strategies. This role sits at the intersection of research, engineering, infrastructure, finance, sourcing, deployment, and operations. You will create the planning models, operating cadences, governance mechanisms, and source-of-truth systems that allow teams to understand what capacity is required, what is available, what is at risk, and what decisions must be made. This is not a finance-only forecasting or reporting role. Success requires technical fluency across the infrastructure stack, strong analytical judgment, and the ability to move consequential decisions forward when requirements, timelines, and supply conditions change quickly. Key Responsibilities Own capacity-planning processes across near-term workload allocation, quarterly execution, and longer-range infrastructure horizons. Translate research, training, inference, and product demand into compute, accelerator, cluster, networking, storage, rack, power, and site requirements. Develop scenarios that make assumptions, confidence levels, constraints, sensitivities, and decision points explicit. Reconcile requested demand against contracted, delivered, installed, activated, and workload-usable capacity. Partner with research and engineering teams to understand workload priorities, technical dependencies, utilization patterns, and changing req

PythonSQLAWSRest
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<

PythonAWSAzureGCP
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA builds the silicon behind AI, accelerated computing, and graphics. Every watt of performance and every degree of thermal headroom traces back to decisions made in power, performance, and thermal architecture. We are the Silicon Co-Design Group (SCG). We identify, own, and drive system-level co-design ideas. We start with initial concepts and advance to product differentiation across NVIDIA's roadmap. We are hiring a Principal System Power Management and Performance Architect who operates at the ambiguous boundary where workload behavior, silicon capabilities, firmware policies, and platform constraints collide, and who turns that ambiguity into architecture that survives across multiple silicon generations. SCG scope spans architecture, design, software, operations, platforms, and productization. This role shapes system, platform, and data center features and behavior, and partners with teams across NVIDIA. What You'll Be Doing: The work here is rarely well-defined when it arrives. You will be given problems that appear to be performance gaps or power anomalies and encouraged to build a framework for solving them, not just tackle a single instance. Define the multi-generation roadmap for system-level power and performance features, grounded in prototyping, use-case analysis, and cost/benefit trade-offs across segments. You will decide what problems are worth solving and why. Own the architecture and integration strategy for HSIO power management, DVFS, P-states, and low-power features. Your decisions improve product performance, power, and reliability across product lines — not just the current program. Lead system-level boot and IST architecture defining how power and clock domains initialize, sequence, and recover across complex multi-IP systems where the interaction space is large and the failure modes matter. Drive power management strategy at data

N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

We're looking for a Principal Engineer to join our CSP Engagements team as the technical focal point for end-to-end performance, working directly with engineering teams of key CSP/hyperscale customers to ensure they achieve various performance targets on NVIDIA platforms. In this role, you will augment NVIDIA's performance and benchmark teams with a dedicated CSP-facing focus. You will drive work streams with CSP engineering teams to build shared understanding of platform performance characteristics, gather and incorporate their workload-specific feedback into NVIDIA's optimization priorities, and validate that performance targets are met in customer-representative configurations. Your cross-CSP visibility enables you to identify patterns and drive systemic improvements in documentation, configuration guidance, and tooling. What you'll be doing: Drive performance characterization work streams with engineering teams of key CSP/hyperscale customers — ensuring they understand platform performance expectations, profiling methodology, and tuning options for their specific workloads Gather and synthesize CSP performance feedback — identify gaps between expected and actual throughput, and champion optimization priorities back into NVIDIA's CUDA, NCCL, driver, and firmware teams Ensure key open-source performance and stress tools (e.g., STREAM, GPU Burn, GPU BLAST) are updated and validated for the latest NVIDIA rack-scale systems, GPU architectures, and CPU platforms — so customers and internal teams have reliable baseline measurements from day one Work closely with CSPs to ensure their own performance and validation tooling reflects the latest GPU capabilities, memory hierarchy changes, and platform-specific tuning parameters Conduct cross-CSP performance comparison and pattern analysis — identify configuration, software, or workload differences that explai

PythonArtificial IntelligenceAI
🔔

Get new workload porting and performance engineer jobs in United States by email

Daily job updates · Unsubscribe anytime