Jobiba hiring network

Software Reliability Engineer Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b

typescriptpythonaws
View job →

Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability. Key Responsibilities MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and

pythonawsazure
View job →

This is where your work makes a difference. At Baxter, we believe every person—regardless of who they are or where they are from—deserves a chance to live a healthy life. It was our founding belief in 1931 and continues to be our guiding principle. We are redefining healthcare delivery to make a greater impact today, tomorrow, and beyond. Our Baxter colleagues are united by our Mission to Save and Sustain Lives. Together, our community is driven by a culture of courage, trust, and collaboration. Every individual is empowered to take ownership and make a meaningful impact. We strive for efficient and effective operations, and we hold each other accountable for delivering exceptional results. Here, you will find more than just a job—you will find purpose and pride. Your role at Baxter The Principal Systems Engineer will serve as a Product Design Owner (PDO) responsible for technical owner for the design, risk and integration of infusion pump systems and/or projects, which combine electro-mechanical hardware, embedded software, and user interface components, as well as the interface with other related EM and/or digital products. This role drives operational excellence and predictable, consistent execution in our products, and design and risk integrity and may serve as Risk Owner on some projects. The engineer is accountable for product safety, performance, reliability, usability and regulatory compliance, as well as risk. The PDO drives design decisions, manages design control activities, may be accountable for a product risk file, and collaborates across engineering disciplines and cross-functions to deliver robust, reliable, safe and innovative infusion therapy solutions. What you will be doing: Leads interdisciplinary design and development of medical products in compliance with FDA, EU MDR,

recruitment
View job →
JT
Jobiba Technologies
📍 Sahibzada Ajit Singh Nagar
1mo ago

IoT Hardware Engineer Build the Future of Smart Connected Devices Join our growing team as an IoT Hardware Engineer and work on exciting IoT solutions that combine electronics, embedded systems, and smart technologies. We are looking for a passionate, hands-on engineer who enjoys designing, building, testing, and improving real-world connected devices. Key Responsibilities: Design, develop, and prototype IoT devices using ESP32 microcontrollers. Develop and integrate firmware, hardware components, and embedded systems. Select and evaluate sensors, actuators, communication modules, and electronic components for IoT applications. Participate in hardware design reviews and product improvement discussions. Assemble, test, and debug electronic prototypes for performance and reliability. Troubleshoot hardware and software issues using professional debugging techniques. Required Skills & Experience: ITI, Diploma in IT Hardware Engineering, Computer Engineering, Electronics, or related field (highly preferred). Strong experience in programming and working with microcontrollers, especially ESP32. Experience with embedded platforms such as ESP32, STM, and similar systems. Knowledge of sensor interfacing using I2C, SPI, UART, and other communication protocols. Excellent PCB soldering skills and experience building complex electronic prototypes. Strong understanding of electronics fundamentals including resistors, voltage levels, power management, and circuit design. Ability to use testing tools such as oscilloscopes and logic analyzers for debugging circuits. Preferred Skills: Experience with KiCad or other hardware design tools. Knowledge of low-power design techniques for battery-operated IoT devices. Understanding of power optimization and embedded system efficiency. Familiarity with Agile development methodologies. Who We’re Looking For: A creative and skilled IoT enthusiast who loves turning ideas into working hardware products. If you enjoy experimenting with e

Anyscale Platform Engineering Leader About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for an experienced Engineering leader to lead our Infrastructure, SRE and Enterprise Governance Engineering teams. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud using Ray - the popular open source platform used by companies like Netflix, Uber, Instacart and others - seamless. In this position, you will guide the vision, technical direction of the team, and recruit, enable a high-performing engineering team that delivers critical values to developers and Anyscale customers by solving complex distributed systems challenges. You will oversee and drive the strategy and execution of components which includes cluster launcher, cloud providers (AWS/GCP/Azure/etc.), Kubernetes support, cluster autoscaling, control plane, data plane, reliability, billing stack, production database and related components. You will closely work with our customers and our field engineering team to solve their problems, understand their challenges and make sure they are successful. We'd love to hear from you if you have: Solid engineering management experience leading produ

awsazuregcp
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking an Actuator Gear Design Engineer to lead the development of custom gears and gear stages for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, and reliability teams. You will partner with external suppliers and internal manufacturing to create full gearbox assemblies. This role focuses on the design, integration, and validation of precision gearing, including broader knowledge around motor electromagnetics, transmission types, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will: Lead the architecture, design, and integration of custom robotic actuator gearing. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, back drivability, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-lo

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a senior Actuator Electromagnetic Design Engineer to lead the development of custom electromechanical actuators for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, reliability, and manufacturing teams. This role focuses on the design, integration, and validation of precision electromechanical systems, including motors, transmissions, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will: Lead the architecture, design, and integration of custom robotic actuators, including the design, simulation, integration and sourcing of custom electromagnetic components. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-loop performance. Devel

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We're looking for an Optical Interconnect System Engineer to design, qualify, and deploy scalable optical connectivity for large-scale AI infrastructure. This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability, and data-center teams to translate system needs into practical interconnect solutions. This is a hands-on role for someone who can connect design decisions with installation, qualification, troubleshooting, and long-term operational performance. In this role, you will: Define optical interconnect architectures and requirements across hardware platforms and rack-level systems. Design high-density fiber systems for performance, density, reliability, installation, and serviceability. Lead optical-mechanical integration and cross-functional design reviews. Develop test and qualification plans for optical components, modules, switching platforms, and integrated systems. Own optical loss budgets, routing guidelines, handling requirements, and serviceability criteria. Support system bring-up, deployment, troubleshooting, failure analysis, and reliability improvement. Create reusable design guidelines, interface requirements, and qualification methods. You might thrive in this role if you have: Core experience Experience desi

awsrestai
View job →

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b

typescriptpythonaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Release Engineer team is responsible for building and maintaining the systems that power software delivery—from CI/CD pipelines and artifact management to release automation and fleet telemetry. We ensure software across bootloaders, firmware, operating systems, and cloud services is built reproducibly, validated rigorously, and released safely at scale. About the Role As a Release Engineer, you’ll design, build, and operate release infrastructure that enables reliable, secure, and traceable software delivery across complex multi-component systems. You’ll partner closely with embedded, cloud, and QA teams to ensure that every build—from development to OTA deployment—is fast, verifiable, and production-ready. We’re looking for engineers who take pride in automation, build reproducibility, and system reliability—and who enjoy building the connective tissue that allows hardware and software to ship together seamlessly. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and operate CI/CD pipelines for multi-component builds (bootloader, firmware, OS images, backend, companion apps) using hermetic toolchains. Define versioning and branching strategies; automate promotions, changelogs, and artifact retention. Integrate unit, integration, and hardware-in-the-loop (HIL) test results; quarantine flaky tests, auto-bisect failures, and block unsafe promotions. Build A/B OTA update flows with verity and health checks; run staged rollouts and canaries; implement safe rollback and roll-forward strategies. Implement code signing for binaries and firmware, generate SBOMs, run vulnerability scanning, and attach build attestations and provenance. Manage dashboards and alerts for build health, promotion latency, failure rates, and fleet update telemetry. You might thrive in this role if you: Have experience building and operating buil

pythonawsci/cd
View job →

About the role We’re looking for an engineering manager to lead a team building software systems that detect and prevent harmful misuse of frontier AI models—before incidents occur. This is a builder’s role: you’ll lead engineers shipping production services, detection pipelines, and mitigation mechanisms that protect frontier model integrity and reduce high-severity misuse risk. While this work intersects with frontier model development, security and risk, we’re explicitly seeking someone with a software engineering foundation who is comfortable building reliable systems that can operate at billions of users scale. In this role you will: Lead a team of software engineers building detection + mitigation systems for frontier model misuse, with an emphasis on model IP protection / distillation detection and emerging risk surfaces from autonomous agents. Set the technical roadmap and execution strategy: prioritize, design, ship, iterate, measure impact. Build production systems: services, pipelines, tooling, instrumentation, and automation that scale with frontier model usage. Partner deeply with Research and Product to translate evolving model capabilities into concrete tests, signals, and mitigations that can be deployed at scale. Drive strong engineering fundamentals: architecture, reliability, monitoring, performance, and operational excellence. Hire and grow an exceptional team across backend, data systems, and applied ML engineering domains as needed. Anticipate what breaks at scale as agentic workflows become more capable. You might thrive in this role if you: Experience building systems in adversarial, fast-evolving environments Are comfortable with ambiguity and novelty Have experience adjacent to security (e.g., abuse prevention, fraud, integrity, platform defense, auth/identity, malware/spam, adversarial environments) Communicate clearly and build trust quickly with senior stakeholders—pragmatic, collaborative, and calm under scrutiny. Significant experience

awsrestai
View job →

About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage distributed data pipelines, collaborate closely with researchers to translate requirements into robust systems, and harden pipelines that serve as the backbone for OpenAI's rapid iteration cycles. We’re looking for engineers who are detail-oriented, have strong experience with distributed systems, and excel at building reliable infrastructure in high-stakes environments. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain data infrastructure systems such as distributed compute, data orchestration, distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability, and security. Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient. Partner with researchers to deeply understand requirements and translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation. You might thrive in this role if you: Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data. Are detail-oriented and bring rigor to building and maintaining reliable systems. Demonstrate excellent software enginee

awsrestmachine learning
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users. That means working across the systems that make Codex actually function as an agent in the real world: the production performance envelope around tokens, latency, reliability, cost, and capacity; the core execution loop and interfaces that turn models into useful behavior; the shared infrastructure that enables other teams to build on Codex; and the feedback loops that turn real-world usage into better models and better agent behavior over time. About the Role We’re looking for applied AI engineers to help bring Codex agents from impressive demos to dependable tools. This role is about improving agent performance on real software engineering tasks and closing the gap between research capability and real-world usefulness. You’ll work closely with research, infrastructure, and product to ensure agents are not just powerful, but useful, steerable, and reliable in practice. The job is not only to improve model behavior in isolation, but to turn those improvements into measurable gains in solve rate, usefulness, and economic value for users. What You’ll Do Design and iterate on agent behaviors across real-world coding tasks and long-horizon workflows. Work closely with research to develop and run evals to measure agent performance, regressions, failure modes, and edge cases. Improve performance through prompting, tool-use strategies, context construction, and model-facing experimentation. Analyze failures in production and systematically improve robustness and reliability. Build feedback loops and data systems that get better real-task data into evaluation and research. Work with product teams to shape user-facing agent experiences and the interfaces the agent depends on. Help define what “good” looks like for agents completing complex tasks end-to-end. You Might Be a Good Fit If You Ha

pythonawsrest
View job →
G
10 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

pythonartificial intelligenceai
View job →
U
11 days ago

About Ubiquiti At Ubiquiti Inc., we create technology platforms for Businesses, Smart Homes, and Internet Service Providers, driven by our goal to connect everyone, everywhere. To date, Ubiquiti has shipped over 100 million devices worldwide, from ISP networking products to next generation of IT solutions. Our growth is made possible by the dedicated team of hundreds behind the scenes. From software developers and product managers to designers and strategists, Team UI is driven to achieve our common goal: Rethinking IT. At Ubiquiti, you’ll heighten your potential and broaden your horizons - all while shaping the future of connectivity. Responsibilities Lead hardware circuit design, schematic capture, and layout review for next-generation, high-density, and high-throughput networking platforms. Evaluate and integrate advanced switching architectures, management subsystems, and cutting-edge high-speed interconnect technologies. Drive hardware architecture design, system bring-up, high-speed signal integrity (SI) validation, and root-cause failure analysis. Partner with mechanical, thermal, and power engineering teams to address challenges related to high power density, thermal dissipation, and system-level reliability. Own BOM structure and support factory deployment to ensure seamless transition of high-layer-count PCBAs from NPI to mass production. Collaborate with cross-functional software, firmware, QA, and compliance teams throughout the entire product lifecycle. Q ualifications Bachelor’s degree or above in Electrical Engineering or a related discipline. 5+ years of hands-on experience in high-complexity system-level hardware design, ideally focused on enterprise-grade networking or high-performance infrastructure equipment. Deep technical understanding of high-speed Ethernet design, high-speed differential signals (advanced SerDes, PCIe, multi-gigabit/ultra-high-speed interfaces), and high-density PCB design rules. Practical experience with complex power d

🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime