Jobiba hiring network

Software Reliability Engineer Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

S
1mo ago

About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe Terminal helps our users extend their online presence to the physical world. The Terminal team’s mission is to make it as easy for businesses to accept in-person payments as the Stripe API has done for online payments. With Terminal, businesses can unlock in-person payments use cases that are right for their business model—whether it’s creating a superb retail experience, extending their website to a pop-up store, or enabling a mobile point-of-sale at their next event. We’re looking for an experienced product manager to lead and shape the future of the software platform that powers Stripe Terminal’s devices portfolio. Working closely with your engineering partners you will design, build and launch device software capabilities that delight users and differentiate Stripe’s solutions in the market. Working closely with our hardware experts you will build the multi-year strategy for how Stripe will continue to enhance the scalability, reliability, usability, and market differentiation of the software capabilities of our in-person commerce devices. In this role, you will work with a spectrum of users, from our largest platforms to small start ups as well as external technology partners to deeply understand user needs and market trends. You will obsess over stability, scalability, and expanding our product while delivering the best in person experiences in the world. What you'll do: Set a motivating multi-year vision for device softw

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the hardware platforms that support the next generation of AI systems. The Opportunity As a Systems Engineering Intern, you will contribute to projects that combine hardware, firmware, and software engineering for advanced AI compute platforms. You will work with experienced engineers on subsystem design, laboratory testing, system validation, automation, and performance analysis. The internship provides hands-on experience with modern hardware development and system-level engineering. You will own clearly defined technical tasks with guidance from the team and document your methods, results, and conclusions. What You Will Do Support the design and testing of CPU and high-speed input and output subsystems for advanced compute platforms. Run laboratory tests and measurements to help evaluate performance, power, signal behavior, and reliability. Contribute to system-level validation by creating scripts and tools that streamline testing, data collection, and analysis. Explore emerging input and output technologies, including PCIe 6.0 and 800G Ethernet, and learn how they support advanced computing workloads. Assist with investigations into platform power, cooling, and energy efficiency, including liquid-cooling systems for high-performance processors. Use power meters, oscilloscopes, logic analyzers, or comparable lab equipment under appropriate supervision. Analyze test results, identify unexpected behavior, and work with engineers to reproduce and investigate issues. Collaborate across hardware, firmware, software, m

pythonartificial intelligenceai
View job →

Senior Product Manager, Robotics & Autonomy What we're doing isn't easy, but nothing worth doing ever is. At Diligent Robotics, we envision a future powered by robots that work seamlessly with human teams. We build artificial intelligence that enables service robots to collaborate with people and adapt to dynamic human environments. Our robots operate every day in hospitals, helping healthcare staff spend less time on routine work and more time caring for patients. Operating a real-world fleet gives us something few robotics companies have: continuous customer feedback and operational data that directly shapes the next generation of Physical AI. We're looking for a Senior Product Manager, Robotics & Autonomy to define and execute the product strategy for some of the most critical capabilities in our robotics platform. You'll work at the intersection of robotics, autonomy, AI, and software engineering to translate business priorities, customer needs, and technical opportunities into a clear product roadmap that drives measurable outcomes. This role is ideal for someone who understands complex autonomous systems and enjoys working alongside world-class engineers to bring ambitious technology from concept into production. Responsibilities Own the product strategy and roadmap for key Robotics and Autonomy initiatives, balancing customer impact, technical feasibility, and long-term platform investments. Define product requirements for autonomy, navigation, perception, fleet intelligence, simulation, and robotics platform capabilities. Partner closely with Engineering, AI, Robotics, Customer Success, Operations, and Leadership to align priorities across the organization. Translate customer feedback, fleet telemetry, and operational insights into product decisions that improve robot performance, reliability, and user experience. Prioritize investments using data, customer value, technical complexity, and business impact. Drive cross-functional execution from concep

agilemachine learningai
View job →
H
Hyreo
📍 Bengaluru• Full-time
18 days ago

Roles and Responsibilities Own the architecture of Myntra’s new product platforms to drive business results Drive and own the architecture and design of some of the most advanced & complex software systems / products in the industry to create company wide impact Help build, mentor and coach a team of very talented Engineers, Architects, Quality engineers, System Operation Engineers and DevOps engineers in architectural and design best practices Experience in distributed systems, cloud service development, deployment and delivery Accountable for the design, for the ease of evolution, quality of the systems, performance, scaling, and availability characteristics and limitations of the systems Envision and develop the long-term architectural direction, with emphasis on platforms/ reusable components while adopting an agile delivery process. Establish structures and processes that ensure a high level of quality and reliability and extensibility of deliverables Drive the creation of next generation extensible web, mobile and fashion commerce platforms, security protocols, customisation and tools to support continuous scaling, internationalisation and platform extensions Drive code and design reviews of components / systems / products in scope and drives the architectural governance for them Set directional paths for the teams/department for adoption of new technology stacks for solving business problems Represent multiple technology domains and Myntra in external technical forums Work with product management, business stakeholders and other engineering leaders to help define mid-term, long-term roadmaps and shape business directions Initiate and deliver leadership training within the engineering organisation, including training new managers, and drive the growth of leaders to create a strong leadership bench. Qualifications & Experience 8+ years of experience in software product development Must have a d

javasqlagile
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
1mo ago

About Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About the Role We are looking for a TPM to drive development and integration of a range of sensor systems for robotics. This role will drive cross-functional alignment across requirements, engineering design, integration, validation, manufacturing, supply chain, and release processes, helping turn complex sensor-system needs into clear plans, decisions, and milestones. Location and in-person expectations: This role is based in San Francisco, CA and requires in-person presence 4 days a week. In this role you will: Drive requirements alignment across engineering design, integration, testing, and validation for camera modules, LiDAR, IMUs, RADAR, proximity sensors, audio components and the systems they interact with. Coordinate the integration of modules including electrical, mechanical, harnessing, and software interfaces with the full robotic system with deep understanding of timelines to drive the respective PCBAs, enclosures, build and test fixtures, connectors and cables. Establish effective cadences for technical reviews, BOM readiness, change management, production releases, approvals, and decision tracking. Align harnesses, fasteners, assembly fixtures, test fixtures, and documentation so cross-functional teams can execute against a clear plan. Lead validation planning around functional, reliability, NVH failure modes, including testing needs, schedules, and exit criteria. Partner with manufacturing and supply chain to manage handoffs, lead times, dependencies, and production readiness. Drive tradeoff decisions across cost, qua

REMOTEawsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
1mo ago

About Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About the Role We are looking for a Technical Program Manager to own actuator development and integration from system goals through production readiness. The actuator program spans mechanical, electrical, firmware, harnessing, controls, test, reliability, manufacturing, and supply chain, and needs a TPM who can turn cross-functional decisions into clear scope, executable milestones, and timely decisions. In this role, you will help the team converge on the right technical plan, surface risks early, and deliver reliable actuator systems on schedule. This role is based in San Francisco, CA and requires in-person presence 4 days a week. In this role you will: Drive actuator programs end-to-end, aligning scope, milestones, interfaces, dependencies, and exit criteria across engineering teams. Drive scope lock and technical convergence for sprints, MVPs, and stretch goals while connecting component decisions to system performance. Coordinate actuator development across motors, gears, sensing, electronics, and firmware, and align the electrical and mechanical interfaces that connect actuators to the broader robot. Lead validation planning from early prototypes through engineering validation, reliability testing, and production readiness. Drive tradeoff decisions across cost, quality, performance, schedule, and lead time by collaborating cross-functionally and quantifying impact to meet program deliverables. Establish effective mechanisms for technical reviews, change control, design releases, decision tracking, and manufacturing readiness.

REMOTEawsrestai
View job →
O
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
O
OpenAI
📍 United States• Full-time
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
O
OpenAI
📍 Singapore• Full-time
1mo ago

About the Team The Consumer Products team at OpenAI builds the hardware that brings advanced AI technology directly into people’s hands. We design, engineer, and manufacture next-generation consumer devices that combine elegant design, robust engineering, and cutting-edge AI. Our team spans mechanical, electrical, and systems engineering — partnering closely with industrial design, operations, and software to create seamless, high-quality experiences that make AI more accessible and useful. About the Role As a Hardware Engineer on the Consumer Products team, you’ll contribute to the design, development, and integration of complex hardware systems for new AI-driven devices. You’ll work across disciplines to translate early product concepts into reliable, scalable, and beautifully crafted products ready for mass production. This role is based in Singapore. We follow a hybrid model (four days per week in the office) and offer relocation support for new employees. Occasional travel to manufacturing partners may be required. In this role, you will: Design, prototype, and validate mechanical and electrical subsystems from early development through mass production. Collaborate closely with design, operations, and manufacturing teams to ensure quality, reliability, and scalability. Drive DFM/DFA reviews, design iterations, and bring-up of prototype and pilot builds. Debug complex system-level issues and lead root-cause investigations and resolutions. Partner with global suppliers and cross-functional teams to ensure flawless execution from concept to launch. You might thrive in this role if you: Have 6+ years of experience in consumer hardware, consumer electronics, or advanced product development. Bring deep technical fluency in mechanical, electrical, or systems engineering (cross-disciplinary experience a plus). Have taken at least one consumer product from concept to mass production. Are energized by hands-on problem solving, iteration, and building in fast-paced enviro

awsrestai
View job →

About Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About the Role We’re looking for a Technical Program Manager to own and scale the systems that power robotic data acquisition across our development and evaluation environments. This role sits at the intersection of Robotics Engineering, Operations, and Infrastructure, ensuring that DAQ stations and associated workflows reliably produce high-quality data for model training and evaluation. You will drive end-to-end execution of complex, cross-functional programs that integrate robotic platforms, sensing systems, operator tooling, and data pipelines into cohesive, production-ready systems. Success in this role requires strong systems thinking, operational rigor, and the ability to translate ambiguous research needs into scalable infrastructure. In this role you will: Own DAQ Program Delivery Lead the roadmap, execution, and scaling of robotic DAQ systems, ensuring alignment with research, engineering, and operational priorities. Drive Cross-Functional Integration Coordinate across robotics hardware, software, infrastructure, and operations teams to deliver tightly integrated, deployment-ready data collection systems. Operationalize Data Collection Systems Translate experimental and research workflows into repeatable, scalable DAQ processes with clear SLAs, metrics, and reliability targets. System Readiness & Deployment Ensure DAQ stations (robots, sensors, compute, operator interfaces) are fully integrated, validated, and ready for production use across multiple sites. Program Execution & Risk Management Build and manage detai

awsrestai
View job →
IE
12 days ago

About Inspira Education Inspira Education Group is one of the fastest-growing edtech startups in the US. We started with a simple mission to democratize access to high-quality coaching so that every student in the world has an equal opportunity to access the best opportunities. As the world’s leading network of top admissions coaches in medical, legal, business, and college studies, we’re building software and services in one place—disrupting long-entrenched application processes with products and experiences that strive to provide an equal platform for candidates from diverse backgrounds worldwide. As one of the fastest-growing edtech firms in the world, we are backed by some of the leading venture capital firms and investors in the world, including Zeev Ventures, Quiet Capital, Craft Ventures and Jeff Fluhr (Founder of Stubhub). About the role We’re looking for a strong full-stack engineer who can own the complete product development process: understand a business problem, define the solution, design the user experience, build the software, and improve it after launch. You’ll work closely with leadership and business teams, combining hands-on engineering with product management and design responsibilities. You should be highly effective with AI coding tools and have the technical depth to independently review, debug, secure, and maintain everything you ship. This is an in-person role requiring 5 day/week in our NYC office. What you’ll own Translate business needs and user feedback into product requirements, user flows, prototypes, and prioritized development plans. Design and build polished applications across the front end, back end, database, and integrations. Make architecture decisions and scope releases that balance speed, reliability, and future maintainability. Use AI tools throughout development to accelerate implementation, testing, debugging, and documentation. Own deployment, production monitoring, incident resolution, and ongoing improvemen

javascripttypescriptpython
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
18 days ago

Opportunity Overview: We are seeking a Technical, Hands-on Manager to lead a team in building AI-driven healthcare enterprise applications . In this role, you will combine strong people leadership with deep technical expertise to guide the development of scalable, data-intensive solutions that drive meaningful business impact. As a leader who is still technically involved , you will mentor and manage data scientists and analysts while having the ability to guide the team through evaluating, selecting, and implementing the right models , paired with an understanding of end to end workflow and ensure the delivery of scalable AI solutions . This position requires excellent communication and collaboration skills as you partner closely with internal stakeholders and cross‑functional engineering, product, and clinical teams . In our fast-paced environment, adaptability is key—your ability to reprioritize quickly and lead your team through evolving business needs will ensure maximum impact What you’ll do: Lead, mentor, and develop a high‑performing team of Data Scientists and ML Engineers, ensuring strong execution and continuous skill advancement. Contribute to event‑driven architecture design and implementation, enabling asynchronous processing and large‑scale system integration. Work seamlessly across functions—partnering with Data Scientists on model tuning, experimentation, and prompt design; collaborating with Product and Software Engineering to embed AI/ML into user-facing applications; engaging with DevOps/Platform Engineering on environment setup, CI/CD, monitoring, and reliability; and working with Data Engineering on pipeline design and ingestion strategies. Provide technical leadership in the effective use of AWS services such as Lambda, EC2, EMR, S3, Athena, Batch, Textract, Comprehend, Bedrock. Drive the implementation of project scope definition, effort estimation, and planning in close coordination with cross-functional teams. Conduct code reviews, pr

pythonsqlaws
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
1mo ago

About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. Our Go-to-Market team helps organizations understand, adopt, and deploy OpenAI’s technology to solve meaningful business challenges and create lasting value. The Technology team works with leading software, internet, cloud, infrastructure, cybersecurity, semiconductor, and digital-native companies as they build new AI-powered products, transform internal operations, and rethink how they serve their customers. We partner with executives, product leaders, engineers, and go-to-market teams to help organizations integrate OpenAI’s capabilities into their products and businesses responsibly and at scale. The team collaborates closely with Solutions Engineering, Customer Success, Product, Research, Partnerships, Marketing, and Operations to turn customer priorities into successful, durable deployments. About the Role We are looking for an experienced Account Director, Tech to help build and grow OpenAI’s business across the technology industry. You will own relationships with a portfolio of strategic technology companies, helping executive, product, and technical leaders understand how OpenAI’s products can accelerate innovation, improve productivity, and create differentiated customer experiences. You will be responsible for developing account strategies, creating qualified pipeline, navigating complex enterprise sales cycles, and expanding adoption across products, teams, and use cases. This role requires a combination of enterprise sales leadership, technical fluency, commercial judgment, and the ability to operate credibly with both business and engineering stakeholders. You should be comfortable engaging with customers that have sophisticated technical environments, rapidly evolving AI strategies, and high expectations for product performance, security, reliability, and scale. Success in this role will be measured by revenue growth, depth of customer adoption,

REMOTEawsgitrest
View job →

Cloud Platform Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations is currently seeking a Cloud Platform Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking a skilled platform engineer to help build and operate the cloud platform services that host critical enterprise applications and software toolchains. In this role, the selected candidate will focus on the shared platform capabilities that enable teams to deploy, run, and maintain containerized and cloud-hosted solutions in a consistent and supportable manner. As both an individual contributor and technical leader, this position will help define and implement platform standards for Kubernetes, container hosting, deployment automation, configuration management, and operational support. This role is focused on platform reliability, repeatability, scalability, and service enablement, rather than custom application software development. Position Responsibilities: Design, implement, and maintain cloud platform services supporting Kubernetes, containers, ingress, storage integration, secrets management, and service connectivity Build and sustain reusable deployment patterns for Commercial-Off-The-Shelf (COTS), Open Source Software (OSS), and internally customized applications Develop and maintain automation for platform provisioning, upgrades, patching, and lifecycle support Manage cluster lifecycle activities including: Cluster upgrades Node management <

awsazuredocker
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime