About the Team OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production, and we partner closely with Research to bring the next generation of models into the world. We're a small, fast-moving team of engineers focused on delivering a world-class developer experience while pushing the boundaries of what AI can do. We’re expanding into multimodal inference, building the infrastructure needed to serve models that handle image, audio, and other non-text modalities. These workloads are inherently more heterogeneous and experimental, involving diverse model sizes and interactions, more complex input/output formats, and tighter coordination with product and research. About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale. You’ll be part of a small team responsible for building reliable, high-performance infrastructure for serving real-time audio, image, and other MM workloads in production. This work is inherently cross-functional: you’ll collaborate directly with researchers training these models and with product teams defining new modalities of interaction. You'll build and optimize the systems that let users generate speech, understand images, and interact with models in ways far beyond text. In this role, you will: Design and implement inference infrastructure for large-scale multimodal models. Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs. Enable experimental research workflows to transition into reliable production services. Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities. Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers. You might thrive in t
Jobs in United States
System Power Engineer in United States
4,976 active opportunities · Updated October 2026
Showing
15 jobs
Explore current system power engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
The Engineering Manager, Memberships leads the team responsible for building and operating Topgolf’s Memberships systems, from guest-facing membership experiences through the backend services that power them. This role owns the people, process, and delivery of the Memberships engineering team, while staying technically credible across a full stack built on Go, Vue.js, and PostgreSQL. Requirements Lead, grow, and manage a team of full-stack engineers building Memberships systems, including hiring, performance management, career development, and mentorship Set clear goals and expectations for the team, run effective 1:1s, and build a culture of ownership, accountability, and continuous improvement Balance workload and staffing across Memberships initiatives, escalating resourcing gaps and continuity risks early Guide architecture and design decisions across Memberships systems, drawing on full-stack experience spanning Go, Vue.js, PostgreSQL, and API design Set engineering standards and best practices for code quality, testing, and release processes, and stay close to the codebase through code reviews and hands-on problem solving on critical issues Own delivery of the Memberships roadmap end to end, from technical planning through implementation, QA, release, and post-launch monitoring, ensuring systems are observable, testable, secure, and built to scale with guest demand Partner with product, design, QA, and platform engineering to translate guest needs and business priorities into a clear, prioritized Memberships roadmap Represent the Memberships team in cross-functional planning and architecture discussions, communicating progress, risks, and tradeoffs to engineering leadership and business stakeholders Critical Skills Strong architectural judgment and the ability to balance technical debt, delivery speed, and long-term maintainability <
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
About us Graphcore is one of the world’s leading innovators in artificial intelligence compute. We are developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and support the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a family of companies responsible for some of the world’s most transformative technologies. Together, we share a bold vision to enable advanced artificial intelligence and ensure its benefits are accessible to everyone. Graphcore brings together AI researchers, silicon designers, software engineers and systems architects to solve complex technical challenges and deliver innovative computing solutions. Job Summary The Principal Electrical Engineer will be a technical authority within Data Center Engineering, leading the architecture and delivery of safe, resilient and scalable electrical infrastructure for high-density AI computing environments. Working with internal teams, data center developers, utilities, consultants and equipment partners, this role will guide projects from early technical studies through design, construction, commissioning, operation and lifecycle improvement. The successful candidate must reside in, or be willing to relocate to, Austin, Texas. Approximately 10% travel may be required. The Team The Data Center Engineering team is responsible for defining and enabling the infrastructure needed to deploy and operate Graphcore’s computing systems at scale. The team works across electrical, mechanical, thermal, controls, systems and operational disciplines, collaborating with external engineering and construction partners to deliver reliable, efficient and maintainable data center environments. Responsibilities and Duties Act as the technical authority for electrical engineering across data center infrastructure projects, from the utility or on-site power source through to the IT rack. Lead electrical archit
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
Join the Team at Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore brings together deep expertise to solve complex problems and deliver meaningful progress in AI compute. Ready to raise the standard for supplier quality in advanced AI manufacturing? Apply now to be part of the journey. Job Summary We are seeking a highly motivated individual with experience in managing supply chain quality in a high-tech manufacturing organisation. The ideal candidate will be hands-on, comfortable with ambiguity and able to manage multiple projects and stakeholders simultaneously. Working across our entire supply chain in a technically challenging, fast paced environment, you will ensure the readiness of our global supply chain to ramp successfully as we develop and manufacture the world’s most advanced AI systems and services. The Team The Graphcore Quality Team delivers outstanding customer experience and champions sustainable excellence throughout Graphcore. We are responsible for customer, supply chain and product quality, as well as organisational compliance, governance and assurance. You will be joining a diverse
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. Plaid’s mission is to unlock financial freedom for everyone by making money movement and access to financial data simple and secure. As a Fullstack Software Engineer, you will design and build the systems and experiences that power how millions of people connect to their finances. You will work across the stack, building scalable backend services and APIs while also crafting intuitive, high-quality frontend experiences that bring those systems to life. This role is ideal for engineers who enjoy switching between backend problem-solving and frontend user experience work, and who are excited to grow their impact across both. You will collaborate closely with product managers, designers, and other engineers to ship products that are reliable, secure, and delightful to use. At Plaid, engineers take ownership early, contribute to architectural decisions, and see their work reach millions of users. Responsibilities: Build across the stack. Design, develop, and maintain scalable backend services and APIs, as well as intuitive, high-quality frontend applications that bring those systems to life. Collaborate cross-functionally. Partner closely with product managers and designers to define
About the Team The Astral team builds high-performance developer tools to power the future of programming, at OpenAI and beyond, including Ruff, uv, and ty. The Astral toolchain sees hundreds of millions of installs per month and powers hundreds of millions of package downloads per day for the Python ecosystem. As a team, we are building on those foundations to continue solving impactful tooling problems as programming evolves. About the Role We are looking for an experienced software engineer to build next-generation programming language tooling. If you like writing high-performance Rust, it could be a good fit; if you like thinking about the future of programming, it could also be a good fit. Strong candidates tend to have deep experience with Rust, Python, open source, compilers, or developer tools — but few candidates are deep in all of these areas, and we've hired candidates without prior Rust or Python experience. In this role, you will: Design and implement features in Astral’s existing open source projects (Ruff, uv, ty, and python-build-standalone, and more). Support Astral’s open source projects as a maintainer, triaging user issues, reviewing pull requests, and participating in community discussions. Evolve the Astral toolchain to accelerate development velocity at OpenAI. Build entirely new tools, in entirely different programming ecosystems, to power the future of agentic software development. Your background might look something like: 5+ years of professional engineering experience, excluding internships, in relevant engineering roles. High agency and comfort operating in a fast-moving environment, with strong ownership of security, reliability, and operational excellence. Strong developer empathy and communication skills, including experience maintaining open source projects. Exceptional systems engineering fundamentals and a track record of leading complex projects from ambiguous problem statements through to user impact. Proficiency in one or more s
From $192K/yr
About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—allowing for seamless collaboration and problem-solving among Dev, Ops and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team: The Revenue Data Engineering Teams designs, builds and runs the data pipelines and helper systems to accurately and in a timely manner quantify our customers’ usage across all Datadog products. This team is at the leading edge of any new product we release. The Revenue Data Processing team builds and operates the data pipelines that does billing, and cost attribution for all Datadog products. We process terabytes of data daily to power revenue-critical systems and are at the center of every new product launch at Datadog. As a Senior Software Engineer, you will own meaningful parts of a large-scale, mission-critical processing platform — driving architectural improvements, building new billing capabilities, and maintaining the high reliability bar our downstream consumers depend on. You Will: Design and build high-throughput data pipelines for billing and cost attribution Drive platform improvements — latency reduction, Spark optimization, sharding, and cross-datacenter reliability Own root-cause investigations on billing accuracy issues in collaboration with Finance and Product teams Contribute to new billing features Work across Python and Scala, with technologies including Spark, Airflow, Trino, and Apache Iceberg Participate in on-call rotation and maintain a high reliability bar for production systems Contribute to engineering standards and help grow the technical culture of the team You Are: You have significant experience building and operating production data pipelines at scale using Spark and Airflow
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s mission is to unlock financial freedom for everyone by making money movement and access to financial data simple and secure. As a Fullstack Software Engineer, you will design and build the systems and experiences that power how millions of people connect to their finances. You will work across the stack, building scalable backend services and APIs while also crafting intuitive, high-quality frontend experiences that bring those systems to life. This role is ideal for engineers who enjoy switching between backend problem-solving and frontend user experience work, and who are excited to grow their impact across both. You will collaborate closely with product managers, designers, and other engineers to ship products that are reliable, secure, and delightful to use. At Plaid, engineers take ownership early, contribute to architectural decisions, and see their work reach millions of users. Responsibilities: Build across the stack. Design, develop, and maintain scalable backend services and APIs, as well as intuitive, high-quality frontend applications that bring those systems to life. Collaborate cross-functionally. Partner closely with product managers and designers to define requirements and de
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s mission is to unlock financial freedom for everyone by making money movement and access to financial data simple and secure. As a Fullstack Software Engineer, you will design and build the systems and experiences that power how millions of people connect to their finances. You will work across the stack, building scalable backend services and APIs while also crafting intuitive, high-quality frontend experiences that bring those systems to life. This role is ideal for engineers who enjoy switching between backend problem-solving and frontend user experience work, and who are excited to grow their impact across both. You will collaborate closely with product managers, designers, and other engineers to ship products that are reliable, secure, and delightful to use. At Plaid, engineers take ownership early, contribute to architectural decisions, and see their work reach millions of users. Responsibilities: Build across the stack. Design, develop, and maintain scalable backend services and APIs, as well as intuitive, high-quality frontend applications that bring those systems to life. Collaborate cross-functionally. Partner closely with product managers and designers to define requirements and de
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Senior Manufacturing Engineer serves as the senior technical authority for assembly process and equipment design within the Industrialization team. This position designs the assembly lines, fixtures, and material flow systems that convert KARNO prototype and development builds into repeatable, scalable production operations — and then systematically removes waste, labor content, and variation from those operations. Working closely with industrialization leadership, design engineering, production, quality, and supply chain, the Senior Manufacturing Engineer owns the most complex assembly value streams end to end: line and station design, fixture and tooling design, material presentation and handling, process qualification, and continuous waste reduction. This position sets the technical standard for how assembly processes are designed and documented at Hyliion, provides mentorship and design review for other manufacturing engineers, and is accountable for measurable improvement in cycle time, first-pass yield, labor content, and ergonomics across the assembly areas. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Assembly Line and Cell Design : Design assembly lines, cells, and workstations for KARNO core, module, and subassembly operations. Establish work sequencing, balance work content to takt, define station layouts and footprints, and design lines that accommodate planned rate increases rather t
From $131K/yr
Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
Other cities to consider
More places hiring for this role
Get new system power engineer jobs in United States by email
Daily job updates · Unsubscribe anytime