About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching
Jobs in United States
Workload Porting And Performance Engineer in United States
382 active opportunities · Updated October 2026
Showing
15 jobs
Explore current workload porting and performance engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation
About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling, and infrastructure that help improve model performance across OpenAI’s training and inference workloads at frontier scale. About the Role We’re looking for an autonomous, high-ownership developer productivity engineer who cares deeply about helping other engineers move faster, safer, and with more confidence. This role will sit within OpenAI’s Model Performance organization, contributing to developer infrastructure, CI systems, testing workflows, tooling, and broader performance infrastructure efforts. There is also a strong opportunity to contribute to the Triton project and help improve the systems that support performance-critical engineering work across OpenAI. In this role you will: Improve development workflows for engineers working on model performance infrastructure Design and improve CI/CD, release, validation, and testing pipelines Build and maintain tools that improve reliability, iteration speed, and engineering confidence Partner closely with engineers to identify friction in testing, debugging, deployment, and development workflows Contribute to infrastructure efforts that support performance-critical training and inference systems Help improve developer experience across Python-heavy codebases and performance-oriented infrastructure Work in a high-context, ambiguous environment where ownership and good judgment matter You might thrive in this role if: You are motivated by enabling the people around you and helping engineers do their best work You have strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build/release workflows You are highly collaborative, empathetic, and comfortable partnering deeply with technical teams You are strong in Python and enjoy building reliable, scalable developer tools and infrastructure You have experience improving large-scale engineering work
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary The Senior Data Engineer will be responsible for delivering high quality modern data solutions through collaboration with our engineering, analysts, data scientist, and product teams in a fast-paced, agile environment leveraging cutting-edge technology to reimagine how Healthcare is provided. You will be instrumental in designing, integrating, and implementing solutions on-premise as well supporting migrations of existing workloads to the cloud. The Senior Data Engineer is expected to have extensive knowledge of modern programming languages, designing and developing data solutions. The position is open in a data engineering team that is responsible for processing payer files into our Data Warehouse. Required Qualifications 6+ years of experience working with SQL and relational database management systems 3+ years of experience in Cloud Data Engineering Platforms such as AWS, GCP, Azure, Databricks, Snowflake etc. 3+ years of experience in on-prem Data Engineering Platforms such as Microsoft SQL Server, Oracle, Teradata etc. Programming and modifying code in languages like SQL, Python, and PySpark to support and implement Cloud based and on-prem data warehousing services.</spa
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Platform Engineer An individual contributor who will serve as a hands-on technical member of the SMAI Platform Engineering team. In this role, you will help build, operate, and maintain the platforms that power Micron's analytics and AI workloads. You will contribute to platform reliability and scalability through day-to-day engineering, collaborative problem solving, and close partnership with solution architects, project teams, and multi-functional partners. Responsibilities: Collaborate with global platform teams, customers, partners, and vendors to deliver effective technical solutions. Know the latest platform roadmaps, emerging technologies, and new service offerings; evaluate and recommend adoption opportunities. Partner with solution architects to design, implement, and optimize solutions across IaaS, PaaS, SaaS, and Infrastructure as Code (Terraform). Document findings, operational procedures, guidelines, and reusable patterns while providing feedback to vendors and internal teams. Deliver high-quality platform support by managing customer requests, maintaining service standards, and implementing controlled platform adjustments. Monitor platform performance, observability, costs, and resource utilization to identify optimization opportunities and improve reliability. Collaborate with multi-functional teams to ensure seamless operations, scalable architectures, and automation, including AI-driven business solutions. Implement and m
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Job Summary We are seeking a motivated engineer to join the DRAM Systems Engineering team, focusing on the development, evaluation, and optimization of next-generation memory systems for AI accelerators. This role emphasizes research and development across hardware architecture, operating systems, and performance analysis to support Agentic AI inference workloads. If you are ambitious and eager to make an impact in the exciting world of AI and memory systems, this is the perfect opportunity for you! Responsibilities Characterize AI inference workloads and examine memory behavior Build and evaluate tiered memory hierarchies for AI accelerators Study KV cache lifecycles, MoE models, and data placement strategies Compare and optimize explicit versus hardware-assisted data movement Develop, test, debug, and detail system-level and OS components Prototype and evaluate agentic AI systems by building agents and multi-agent workflows using modern frameworks and orchestration patterns (planning, tool use, memory, and context management). Apply these technologies both as workloads under study and as accelerators for internal engineering workflows <h2 style="color:!importan
Specialist - Workstation Description - The Advanced Compute Solutions (ACS) Sales Specialist plays a vital role in helping new and existing customers discover and benefit from ACS. You will use your sales expertise and industry knowledge to build strong relationships, identify opportunities, and support customers throughout their journey. This position focuses on expanding our customer base, advancing the ACS sales pipeline, fostering long-term customer relationships, and supporting account managers in delivering outstanding results. ACS is a key division within HP focused on driving innovation for the future of work. Its mission is to deliver high-performance computing, enabling customers to tackle complex workloads and emerging technologies efficiently. Key Responsibilities Drive growth by acquiring new customers and expanding adoption of ACS products. Identify and pursue new ACS opportunities, while also expanding relationships with existing clients. Manage the sales process end-to-end, including building a strong pipeline, guiding product selection, pricing, and closing deals. Develop business use cases and provide expert support to sales teams. Achieve sales targets and contribute to higher customer quotations, diverse business, and increased contract-value renewals. Collaborate with customers, internal teams, and industry specialists. Stay ahead of the market by understanding ACS offerings and competitor strategies to position HP’s solutions effectively. Build trusted, consultative relationships with clients, including senior executives, by understanding their industry-specific needs and challenges. Education & Experience Recommended: Bachelor’s degree is preferred but not required. </l
Specialist - Workstation Description - The Advanced Compute Solutions (ACS) Sales Specialist plays a vital role in helping new and existing customers discover and benefit from ACS. You will use your sales expertise and industry knowledge to build strong relationships, identify opportunities, and support customers throughout their journey. This position focuses on expanding our customer base, advancing the ACS sales pipeline, fostering long-term customer relationships, and supporting account managers in delivering outstanding results. ACS is a key division within HP focused on driving innovation for the future of work . Its mission is to deliver high-performance computing , enabling customers to tackle complex workloads and emerging technologies efficiently. Key Responsibilities Drive growth by acquiring new customers and expanding adoption of ACS products. Identify and pursue new ACS opportunities, while also expanding relationships with existing clients. Manage the sales process end-to-end, including building a strong pipeline, guiding product selection, pricing, and closing deals. Develop business use cases and provide expert support to sales teams. Achieve sales targets and contribute to higher customer quotations, diverse business, and increased contract-value renewals. Collaborate with customers, internal teams, and industry specialists. Stay ahead of the market by understanding ACS offerings and competitor strategies to position HP’s solutions effectively. Build trusted, consultative relationships with clients, including senior executives, by understanding their industry-specific needs and challenges. Education & Experience Recommended: Bachelor’s degree is preferred but not
Specialist - Workstation Description - The Advanced Compute Solutions (ACS) Sales Specialist plays a vital role in helping new and existing customers discover and benefit from ACS. You will use your sales expertise and industry knowledge to build strong relationships, identify opportunities, and support customers throughout their journey. This position focuses on expanding our customer base, advancing the ACS sales pipeline, fostering long-term customer relationships, and supporting account managers in delivering outstanding results. ACS is a key division within HP focused on driving innovation for the future of work . Its mission is to deliver high-performance computing , enabling customers to tackle complex workloads and emerging technologies efficiently. Key Responsibilities Drive growth by acquiring new customers and expanding adoption of ACS products. Identify and pursue new ACS opportunities, while also expanding relationships with existing clients. Manage the sales process end-to-end, including building a strong pipeline, guiding product selection, pricing, and closing deals. Develop business use cases and provide expert support to sales teams. Achieve sales targets and contribute to higher customer quotations, diverse business, and increased contract-value renewals. Collaborate with customers, internal teams, and industry specialists. Stay ahead of the market by understanding ACS offerings and competitor strategies to position HP’s solutions effectively. Build trusted, consultative relationships with clients, including senior executives, by understanding their industry-specific needs and challenges. Education & Experience Recommended: Bachelor’s degree is preferred but not
Specialist - Workstation Description - The Advanced Compute Solutions (ACS) Sales Specialist plays a vital role in helping new and existing customers discover and benefit from ACS. You will use your sales expertise and industry knowledge to build strong relationships, identify opportunities, and support customers throughout their journey. This position focuses on expanding our customer base, advancing the ACS sales pipeline, fostering long-term customer relationships, and supporting account managers in delivering outstanding results. ACS is a key division within HP focused on driving innovation for the future of work . Its mission is to deliver high-performance computing , enabling customers to tackle complex workloads and emerging technologies efficiently. Key Responsibilities Drive growth by acquiring new customers and expanding adoption of ACS products. Identify and pursue new ACS opportunities, while also expanding relationships with existing clients. Manage the sales process end-to-end, including building a strong pipeline, guiding product selection, pricing, and closing deals. Develop business use cases and provide expert support to sales teams. Achieve sales targets and contribute to higher customer quotations, diverse business, and increased contract-value renewals. Collaborate with customers, internal teams, and industry specialists. Stay ahead of the market by understanding ACS offerings and competitor strategies to position HP’s solutions effectively. Build trusted, consultative relationships with clients, including senior executives, by understanding their industry-specific needs and challenges. Education & Experience Recommended: Bachelor’s degree is preferred but not
Specialist - Workstation Description - The Advanced Compute Solutions (ACS) Sales Specialist plays a vital role in helping new and existing customers discover and benefit from ACS. You will use your sales expertise and industry knowledge to build strong relationships, identify opportunities, and support customers throughout their journey. This position focuses on expanding our customer base, advancing the ACS sales pipeline, fostering long-term customer relationships, and supporting account managers in delivering outstanding results. ACS is a key division within HP focused on driving innovation for the future of work . Its mission is to deliver high-performance computing , enabling customers to tackle complex workloads and emerging technologies efficiently. Key Responsibilities Drive growth by acquiring new customers and expanding adoption of ACS products. Identify and pursue new ACS opportunities, while also expanding relationships with existing clients. Manage the sales process end-to-end, including building a strong pipeline, guiding product selection, pricing, and closing deals. Develop business use cases and provide expert support to sales teams. Achieve sales targets and contribute to higher customer quotations, diverse business, and increased contract-value renewals. Collaborate with customers, internal teams, and industry specialists. Stay ahead of the market by understanding ACS offerings and competitor strategies to position HP’s solutions effectively. Build trusted, consultative relationships with clients, including senior executives, by understanding their industry-specific needs and challenges. Education & Experience Recommended: Bachelor’s degree is preferred but not
We are now looking for a dynamic business leader to grow NVIDIA's Host Networking business for AI infrastructure with AI Labs and Hyperscalers! This leader will drive strategic direction, customer engagement, and multi-year growth for networking products such as NVIDIA DPUs, SuperNICs, and their associated software and ecosystem. Success in this role will be measured by the level of adoption and integration of our Host Networking products with our end customers' workflows and workloads. Success is contingent upon building trust with executives, architects, product leaders, and platform teams across NVIDIA and our largest customers. This leader will lead the go-to-market motion, connecting customer AI factory needs to NVIDIA's networking portfolio and aligning product, sales, engineering, architecture, marketing, and partner teams to secure design wins and scale deployments. What you'll be doing: Identify, develop and close strategic design wins for DPU and SuperNIC with top AI labs and Cloud Service Providers! Build and implement the segment sales growth strategy for host networking across hyperscaler and frontier model AI labs building large scale AI infrastructure. Define customer-specific DPU and SuperNIC value propositions and deployment motions, and lead a matrixed team across product, architects, engineering, sales and marketing teams. Promote NVIDIA host networking products externally and internally, positioning their value for AI workloads and other infrastructure products from NVIDIA, in a collection of use-cases in Networking, Security and Storage. Build a robust opportunity pipeline with segment sales and account teams, including account mapping, customer requirements, proof points, executive engagement, and partner alignment. Track and drive quarterly business reporting, forecast accuracy, design-win progress, roadmap asks, and
Other cities to consider
More places hiring for this role
Get new workload porting and performance engineer jobs in United States by email
Daily job updates · Unsubscribe anytime