Jobiba hiring network

System Engineer Jobs

10,000 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current system engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man

pythonsqlaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i

pythonreactnode.js
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

awskubernetesci/cd
View job →
O
1mo ago

About the team OpenAI’s Forward Deployed Engineering team partners with customers to turn research breakthroughs into production systems. We operate at the intersection of customer delivery and core platform development. About the role Forward Deployed Engineers (FDEs) lead complex end-to-end deployments of frontier models in production alongside our most strategic customers. You will own discovery, technical scoping, system design, build, and production rollout, partnering directly with customer engineering and domain teams. You will measure success through production adoption, measurable workflow impact, and eval-driven feedback that changes product and model roadmaps. You’ll work closely with our Product, Research, Partnerships, GRC, Security, and GTM teams. This role is based in NYC. We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. Travel up to 50% is required. In this role you will Own technical delivery across multiple deployments from first prototype to stable production Build full-stack systems that deliver customer value and sharpen how we learn Embed closely with customer teams, understand their needs, and guide adoption of what you build Scope work, sequence delivery, and remove blockers early Make trade-offs between scope, speed, and quality; adjust plans to protect delivery Contribute directly in the code when progress or clarity depends on it Codify working patterns into tools, playbooks, or building blocks that others can use Share field feedback that helps Research and Product understand where the models succeed and where they can improve Keep teams moving through clarity and follow-through You might thrive in this role if you Bring 5+ years of engineering or technical deployment experience that includes customer-facing work Have scoped and delivered complex systems in fast-moving or ambiguous environments Write and review production-grade code across frontend and backend using Python, JavaScript, or compar

javascriptpythonjava
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system

pythonsqlaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a Senior Actuator Design and Integration Engineer to lead the development of custom electromechanical actuators for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, reliability, and manufacturing teams. This role focuses on the design, integration, and validation of precision electromechanical systems, including motors, transmissions, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will Lead the architecture, design, and integration of custom robotic actuators, including motors, transmissions, sensing, thermal systems, structural components, and packaging. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, backdrivability, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-loop pe

awsrestai
View job →

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role As a Security Engineer on Detection & Response, you’ll help protect OpenAI’s most sensitive assets– including our intellectual property, customer data, and the infrastructure that supports them– by building and operating the systems we use to detect suspicious activity and respond effectively when it matters. You’ll work across endpoints, identity, cloud, hyperscale compute infrastructure, and datacenter-adjacent layers, partnering closely with security teams and infrastructure owners to define the telemetry and response requirements we need and building tooling and automation where it delivers the most leverage. In this role, you will: Build and evolve Detection & Response capabilities across OpenAI’s infrastructure, products, and research environments, with an emphasis on high-signal detection and reliable operational response. Engineer detection pipelines and tooling: develop rule lifecycle management, measurement/quality loops (coverage, precision, latency), tuning processes, and safe rollout patterns. Automate response and investigations by building workflows that reduce toil (triage, enrichment, containment, evidence capture) and improve time-to-understand/time-to-contain. Partner with other Security teams and system/infrastructure owners across the company to ensure new systems ship with the right telemetry, threat models, and response playbooks from day one. Define D&R requirements and drive visibility across endpoin

awsazuregcp
View job →
O
1mo ago

Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role As OpenAI continues to grow, we are looking for experienced, problem-solving engineers to ensure our systems scale. Our success depends on our ability to quickly iterate on products while also ensuring that they are performant and reliable. You will work in a deeply iterative, collaborative, fast-paced environment to bring our technology to millions of users around the world, and ensure it’s delivered with safety and reliability in mind. Successful candidates will play a crucial role in ensuring the reliability, scalability, and performance of our systems as we continue to expand. As a reliability expert, you will be at the forefront of maintaining and enhancing the stability, scalability, and performance of our rapidly evolving infrastructure. You will work closely with cross-functional teams, including software engineers, product managers, and data scientists, to build and maintain resilient systems that can handle our growing user base and workload. In this role, you will: Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands. Build and maintain the load, chaos and synthetic testing software leveraged by development teams to make the systems they design and operate more reliable. Build and maintain automation tools to streamline repetitive tasks and improve system reliability. Build and maintain the platform for CPU/storage, GPU, and network lifecycle management to drive efficiency, accountability and support dynamic optimization of our resources. Implement fault-tolerant and resilient

awskubernetesrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a Senior Mechanical Engineer to lead the design, integration, and sustaining engineering of mechanical subsystems in robotic platforms. You will work closely with experienced engineers and cross-functional partners to set functional requirements, develop and iterate on hardware to meet program expectations. This role is aimed at candidates with strong fundamentals in mechanical system design — including tolerance, alignment, load paths, wear, and failure modes — and robotics. You will contribute to real hardware programs moving from prototype through early production. You will lead both new subsystem development and ongoing improvements to existing systems based on testing, field performance, and manufacturing feedback. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will Lead the design and iteration of mechanical subsystems, including structures, mechanisms, and actuators. Create and maintain CAD models, assemblies, and drawings with appropriate tolerancing and documentation. Build and test prototypes, supporting debugging of mechanical issues such as fit, alignment, friction, and wear. Assist in developing test methods and executing validation to evaluate performance, durability, and failure modes. Work with cross-functional teams to integrate mechanical components with sensors, actuators, and control systems. Support transition of designs from prototype to manufacturable assemblies, incorporating DFM and DFA considerations. Collaborate with manufacturing partners

awsrestai
View job →
O
1mo ago

About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati

pythonsqlaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Privacy Engineering team builds the systems and technical foundations that govern how user data is understood, retained, accessed, and used across OpenAI. We partner with Product, Data, Infrastructure, Security, and Legal to translate policy and trust commitments into durable architecture and enforceable controls. Our work spans data inventory and mapping, classification and lineage, retention and deletion, access governance, purpose and usage controls, auditability, and lifecycle automation. We aim to make policy-aligned data handling the default while giving teams clear, reliable primitives for building and operating products at scale. About the Role We are looking for an experienced Software Engineer to drive the architecture and execution of user data governance across OpenAI. You will define technical direction, build shared platforms and controls, and lead cross-functional programs that make data flows discoverable, policies enforceable, and ownership explicit. This role is well suited to a senior engineer who can move between deep systems design and organization-wide influence, turn ambiguous requirements into pragmatic roadmaps, and operate high-trust systems end to end. This position is based in San Francisco. Relocation assistance is available. In this role, you will: Set the technical strategy and architecture for user data governance across data mapping, classification, lineage, retention, deletion, access, and permitted usage. Design and build shared services, APIs, metadata systems, and policy-enforcement mechanisms that make governance controls consistent, scalable, and auditable. Establish reliable inventories of user data, system ownership, data flows, and policy applicability across products, infrastructure, analytics, and research systems. Partner with Product, Data, Infrastructure, Security, and Legal leaders to define decision rights, translate requirements into controls, and drive adoption across teams. Own governance systems

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work

pythonsqlaws
View job →
O
1mo ago

About the Team The Fleet team builds core components to enable productive research from small to state of the art scale across OpenAI, with the goal of accelerating progress towards AGI. We frequently collaborate with other teams to speed up the development of new state-of-the-art capabilities. About the Role As we scale up with more researchers and engineers joining OpenAI, we seek a pragmatic and passionate engineer with a strong focus on the development experience for both engineers and scientists. In this role, you will be responsible for building and maintaining systems that allow our research + engineering organization to iteratively develop, test, and deploy new features reliably, with high velocity, and with a frictionless and fast development cycle. You will help oversee and drive to the vision of how we should build, test and deploy software. You will drive the design of our continuous integration pipelines, testing infrastructure, training and support around our build system. Our current environment relies heavily on Python, Rust, and C++, which you will take ownership of and strive to transform into a state of the art development experience for research. Ultimately, your role will be to provide the necessary tools and metrics to support our fast-paced culture and ensure a stable, scalable platform for growth, while also fostering a seamless and low friction experience for OpenAI’s research. This role is based in San Francisco, CA. For a San Francisco role, we use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. You might thrive in this role if you: Have supported large monorepo development and deployment before Are a proficient Python programmer working in large monorepos Are proficient with Docker and Kubernetes Experienced in CI/CD About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boun

pythonawsdocker
View job →

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa

awsrestai
View job →
O
1mo ago

About the Team We’re hiring software engineers to make the Workload team more productive. The Workload team maintains the core components of OpenAI’s training and inference frameworks and helps execute frontier experiments. About the Role We’re looking for someone who cares about the developer experience of working in and around OpenAI’s core training and inference frameworks. In this role you will: Be responsible for optimizing the development workflows of the engineers around you Work within various Workload teams to address their specific needs, but collaborate with the centralized teams that own various aspects of development experience Optimize iteration speed, both broadly, and in particular by optimizing specific teams’ CI Improve reliability, for instance, by driving testing strategy for particular components Work through the long tail of things that it takes to build libraries and systems that will delight researchers You might thrive in this role if: You are motivated by helping people. You believe a thing that separates great teams from good teams are the players willing to do whatever work it takes, without ego. You believe in the power of developer experience. Something magical happens when people can quickly and confidently iterate on a simple codebase, but this magic is fragile and must be fought for. When you see someone trip over something, no matter how small, your first instinct is asking yourself what it would take for that to not happen again. Your second instinct is clicking merge on the PR you’ve already written to make it so. You are pragmatic. You have the ability to see the world through a perfectionist’s eyes, but are not yourself a perfectionist. You know which problems to pick and when to switch to making progress on a different problem. You like going end-to-end on things. You love co-design — that feeling when you were only able to find the right solution because you both deeply understand the users that interact with a system and the

pythonawsrest
View job →
🔔

Get new system engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More system engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring System Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.