The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
Jobiba hiring network
Hardware Software Codesign Engineer Jobs
1,301 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current hardware software codesign engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea
About the Team Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running. That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets. About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models
About the Team The Storage Infrastructure team builds and operates the storage foundation behind OpenAI’s most demanding workloads. We work directly with research to design storage systems for rapidly evolving experiments, while also powering production at scale. We own the platform end to end: backend systems, user-facing services and APIs, and the control planes that manage how data is placed, moved, and retained over time. Our stack spans cloud and in-house object stores across very different workload profiles, from GPU-attached systems to dedicated storage hardware. We also build the federation layer that unifies these backends behind a simple interface and routes each workload to the right storage solution. About the Role You will help build the storage platform that powers OpenAI’s research and production systems. This is a hands-on infrastructure role for engineers who want to work on deeply technical systems at scale and own them in production. You’ll work across object storage, cross-region data movement, lifecycle management, and the federation layer that provides a unified interface across multiple backends. Much of our stack runs on Kubernetes, and we primarily build services in Rust. In this role, you will: Build and operate storage services that underpin OpenAI’s research infrastructure Develop object storage systems across cloud and in-house environments Build systems for cross-region data movement, replication, and recovery Design lifecycle management capabilities that keep data durable, available, and cost-effective Evolve the federation layer that unifies multiple backend systems behind a simple interface Improve performance, reliability, and operational excellence across the platform Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads You might thrive in this role if you: Have experience building or operating distributed systems in production Have worked on storage infrastructure, object stores, dist
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches
About the Team We’re hiring software engineers to make OpenAI’s networking teams more productive. These teams build and operate the high-performance networking systems that support OpenAI’s training and inference infrastructure at frontier scale. About the Role We’re looking for someone who cares deeply about the developer experience of engineers working on complex infrastructure systems — especially around build systems, test architecture, release pipelines, and reliable development workflows. This role will be embedded with OpenAI’s networking team: making it faster, safer, and easier for engineers to build, test, validate, and ship changes across multi-server, networked, and hardware-adjacent environments. In this role you will: Improve development workflows for engineers building and operating OpenAI’s networking systems Design and improve continuous deployment, release, and validation pipelines Build and maintain test harnesses for multi-server, networked, and hardware-backed environments Improve iteration speed across C++, Python, and build-system-heavy codebases Partner with engineers to identify friction in CI, testing, debugging, and deployment workflows Drive testing and reliability strategy for infrastructure components that support large-scale training and inference workloads Work closely with centralized developer experience teams while staying deeply embedded with the networking engineers closest to the systems You might thrive in this role if: You are motivated by helping other engineers move faster and with more confidence You have experience with CI/CD, release pipelines, testing infrastructure, or build systems You are comfortable moving between C++, Python, and build systems such as CMake, Bazel, or Blaze You enjoy building test harnesses, automation, and workflow improvements for complex systems You do not need to be a networking expert, but you are excited to learn enough about the domain to make the team meaningfully more effective When you see
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake’s cloud spend is in billions of dollars per year. Hence, it is critical for us to govern and optimize our cloud spend, both for margins and long-term competitive advantage. Cloud Efficiency team’s charter is to build scalable products that enable governance, monitoring and optimization of cloud spend. Think of this as Observability for cloud costs and efficiency. The team’s vision is to “Transform cloud spend into a competitive advantage by empowering teams to continuously optimize the per-unit cost.” In order to improve the overall cloud efficiency (i.e. cost per unit), it is critical to build monitoring products that collate costs with other factors such as utilization, attribution, hardware performance and architecture. Hence, there is an opportunity to build a unified, self-serve cloud efficiency product across Snowflake, that delivers actionable, real-time efficiency datasets through streamlined user experiences. This will enable thousands of engineers at Snowflake and will elevate cloud efficiency at Snowflake for long-term success. When developing these solutions, we think about the problem end-to-end: how do we collect data from different stacks (e.g. costs from AWS, GCP, Azure and CPU, Memory, Utilization) across Snowflake reliably, how do we store it eff
Software Developer in Test Description - This role is responsible for ensuring quality, reliability and performance of software applications throughout the development lifecycle primary through software test automation. The role designs, codes, and implements software test automation using appropriate programming languages, frameworks, and tools. The role works closely with cross-functional teams to gather requirements, provide technical insights, and ensure the successful execution of test automation with the main goal of improve quality of the solution. The role also creates and executes comprehensive test plans, test cases, and test scripts based on project specifications. The role sets and provides design guidance to other developers and SQA engineers regarding test automation and the test framework. *Onsite in Ft Collins 4-days a week Responsibilities • Designs quality assurance and test processes for portions of end-user video conferencing/collaboration application, systems software running on android hardware, local, networked, and Internet-based platforms. • Analyzes design and determines test scripts, coding, automation, and integration activities required based on specific objectives and established project guidelines. • Designs and maintains Test automation framework • Executes and writes portions of testing plans, protocols, and documentation for assigned portion of application; identifies and debugs issues with code and suggests changes or improvements. • Identifies opportunities for performance improvements and optimizes code and application performance. • Utilizes latest AI tools and technologies in speeding up test automation • Executes test cases depending on the needs of the project • Provides valuable input into the development of user stories and acceptance criteria, shaping a quality-oriented d
Software Chief Architect - AvionX Company: The Boeing Company The Boeing Company is looking for a Software Chief Architect - AvionX to join the AvionX Software team located in Long Beach, California . This position will focus on leading the AvionX Software Engineering organization. The Chief Software Architect for AvionX Software Vertical is the highest software technical authority across airborne software programs from architecture definition through certification closure. This role owns the coherence, integrity, and DO-178C/LOR compliance posture of software architectures across commercial and defense platforms, bridging software engineering discipline with system-level safety objectives and regulatory strategy. This role is also responsible to work with Enterprise Software Engineering Organizations to enable SW engineering practices and processes, such as Fabric, Design Practices etc. to be implemented across all AvionX Programs. Position Responsibilities: Advises management on a wide range of topics related to the embedded software engineering environment Consults on code for embedded systems software to run on specific specialized hardware Directs testing and debugging of software for embedded devices and systems. Influences current and emerging technologies, tools, frameworks, and changes in regulations relevant to software development and hardware technologies Directs the design, development, test, debugging and maintenance of software that is integrated into embedded devices and systems and meets industry, customer, safety and regulation standards Directs the review, analyses, and translation of customer requirements into the design of softwar
Job Details: Job Description: The Role and Impact As a Systems and Solutions Engineer, you will drive the design, development, and integration of systems that combine software, firmware, board, and silicon/SoC components to meet specific customer needs. In this role, you will play a key part in defining, implementing, and optimizing solutions to ensure high performance, reliability, and quality across the system lifecycle. Your work will directly impact the seamless functionality and user experience of cutting-edge technologies, enhancing Intel's position in delivering innovative systems to global customers. Business Group You will be joining the Silicon and Platform Engineering Group (SPE), an organization committed to advancing Intel's mission of delivering world-class silicon and platform solutions. The group focuses on developing integrated systems that align with customer needs and support Intel's broader goals of leadership in technology innovation. SPE collaborates across diverse domains to ensure Intel platforms meet performance, reliability, and scalability requirements. Key Responsibilities - Design and develop software, firmware, and hardware solutions that integrate seamlessly across system components. - Lead the definition and implementation of system architecture, translating business opportunities into technical requirements and use cases. - Evaluate technical risk and optimize systems for ease of use, reliability, security, availability, and sustainability. - Drive technical solutions to address customer challenges, deploying systems and conducting benchmarks to validate performance. - Collaborate with cross-functional teams to influence next-generation requirements and solutions, guiding research and academic collaborations as needed. - Conduct lab experiments to simulate real-life environments, analyze prototype performance, and refine system
The NVIDIA PerfTech team is looking for a talented C++ Software Engineer to help build the next generation of AI-powered developer tools. You will apply strong C++ and software-engineering fundamentals while gaining hands-on experience with agentic workflows, retrieval systems, and AI services. In this role, you will contribute to Genie, NVIDIA’s company-wide AI knowledge and developer-productivity service. You will work across C++ tools and AI services to help engineers find information, understand complex systems, and work more effectively. What You’ll Be Doing: Develop production-quality C++ components, APIs, and integrations for NVIDIA’s AI-powered developer-tools ecosystem. Build capabilities connecting native C++ tools with Genie’s retrieval and agentic features. Contribute to agentic workflows, retrieval systems, ingestion pipelines, MCP tools, APIs, and enterprise integrations. Build benchmarks and improve retrieval quality, reliability, performance, and resource usage. Own features from investigation and design through implementation, testing, and delivery. Collaborate with graphics, software, and hardware teams developing performance-analysis and developer tools. What We Need to See: Bachelor’s or Master’s degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience. 5+ years of modern C++ programming skills gained through professional experience, internships, or substantial technical projects. Good understanding of data structures, algorithms, object-oriented design, multithreading, debugging, and testing. Ability and motivation to work across C++ systems and Python-based AI services. Familiarity with AI-powered applications, agentic workflows, retrieval systems, or related technologies. Abil
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Software Engineering Intern on the Systems team at HP IQ, you’ll work on low-level software that sits close to the hardware and helps power intelligent experiences across our products. This role is ideal for students who enjoy understanding how complex systems work under the hood. You’ll have the opportunity to work across multiple layers of the software stack, investigate performance bottlenecks, optimize system behavior, and build software that interacts closely with hardware and system resources. We’re looking for engineers who are curious about more than whether something works — you want to understand how it works, why it performs the way it does, and how to make it better. What You Might Do Build and optimize low-level systems software using languages such as C and C++. Investigate performance bottlenecks and improve the speed, efficiency, and reliability of existing systems. Work on data processing and sensor pipelines that connect software with underlying hardware. Analyze and improve memory usage, resource management, and system performance. Work across multiple lay
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About the Role HP IQ’s Security Team is creating something the world has never seen before. We are attempting something truly impactful — innovating at the deepest levels of hardware and launching a service that will inspire users to experience computing in an entirely new way. Privacy and security are not just priorities; they are fundamental to our product and essential to our success. What You Might Do We are seeking a software engineering intern ready to take on the challenge of securing users’ devices, users’ data, and HP IQ’s infrastructure. We see security as the key to accomplishing what other companies cannot. If you thrive at balancing exceptional user experiences with strong security, this is the place for you. Essential Qualifications Pursuing a degree (Bachelors or Graduate) in Computer Science or related technical field Strong software engineering skills Ability to thrive in a collaborative environment Interest in privacy & security Interest in product design & development Experience with C, C++, Java or Python Preferred Skills An understanding of security concepts
Get new hardware software codesign engineer jobs by email
Daily job updates · Unsubscribe anytime