Jobiba hiring network

Reliability Engineer Jobs

2,049 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of this API & power-users team, you will improve the capabilities, reliability, and product fit of OpenAI’s agentic models for power users and API developers. You might design evals from real developer workflows, build training environments around production-like tool use, turn qualitative model failures into training data, evals, or post-training interventions, or drive a behavior improvement from discovery through post-training, integration, and launch. This role is intentionally broad. The strongest candidates are comfortable turning ambiguous model behavior problems into concrete progress, whether that means improving tool use, planning, instruction following, recovery from mistakes, or how models behave in API-based workflows. You should be excited to work across research, engineering, data, evals, and product to make models better at acting in real workflows. You will work closely with researchers, engineers, API/product teams, Codex, infrastructure, and safety/align

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and

awsrestai
View job →
O
1mo ago

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Computer Use, you will teach models to operate computers. You will help train models that can navigate browsers and desktops, use tools and applications, reason through complex workflows, collaborate with users and other agents, and complete long-horizon tasks with reliability and judgment. This work sits at the intersection of frontier model training, product behavior, evaluation, and systems engineering, and will directly shape the computer-use capabilities shipped in OpenAI’s next generation of agents. Currently, our models are the best in the world at this behavior! You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might Design and run experiments th

awsrestmachine learning
View job →

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Operations Program Manager (OPM) to serve as the single-threaded operational leader for new hardware introductions (NPI) and production ramps across OpenAI’s AI infrastructure systems. This role combines hands-on execution with strategic ownership. You will be responsible for defining the operating model, aligning cross-functional stakeholders, setting the critical path, making informed tradeoffs, escalating decisively, and ensuring hardware programs deliver on schedule, quality, cost, and scalability. Success in this role requires comfort operating in ambiguity, influencing without authority, and driving alignment across internal teams and external partners—while keeping eyes firmly on long-term system scalability and repeatability. In this role, you will: Strategic & Leadership Ownership Act as the single-threaded owner for operational readiness across NPI and ramp, accountable for outcomes from early bring-up through sustained production Translate OpenAI’s infrastructure strategy and engineering objectives into clear operating plans, execution priorities, and decision frameworks Drive alignment across Engineering, Operations, Strategic Sourcing, Finance, Capacity Planning, and Executive stakeholders by framing tradeoffs, risks, and recommendations Proactively identify inflection points where decisions or investments are required to protect long-term scale, reliability, or cost targets Influence operational strategy with manufacturing par

awsrestai
View job →

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity As a Member of Technical Staff on AI Infrastructure, you will build and maintain the foundational systems and distributed infrastructure that power AI model post training, inference, and data pipelines. You will collaborate with engineering and research teams to ensure performance, scalability, and reliability of critical AI systems. What You’ll Do Design and implement large-scale, distributed AI infrastructure and services Optimize performance for GPU/xPU accelerators and cloud environments Build tools for observability, reliability, and scaling of AI workloads Partner with cross-functional teams to define AI infrastructure requirements and roadmap Contribute to architectural design and system longevity About You Have experience with GenAI infrastructure systems, distributed systems, cloud computing, and high-performance infrastructure Are proficient in programming languages like Python, Go, or similar Understand scaling challenges specific to AI workloads and accelerators Thrive in fast-paced, collaborative engineering environments The reasonably estimated base salary for this role ranges from $256,000.00 to $276,

pythonaigo
View job →
S
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Unistore (Hybrid Tables) is Snowflake's strategic convergence investment—seamlessly combining OLTP and analytical workloads in a single unified database, and one of the company's most strategically differentiated bets. Operating like a startup within Snowflake, you will own revenue goals, define product strategy for the next phase of growth following our recent General Availability (GA) milestones, and coordinate go-to-market as the primary product authority for the area. OLTP-HTAP convergence is a fast-moving, high-stakes space with strong customer demand—making this one of the highest-visibility product leadership roles in the company and a rare opportunity to define the future of transactional and analytical convergence at scale. AS A STAFF PRODUCT MANAGER AT SNOWFLAKE, YOU WILL: Drive the Unistore product area end-to-end—including product strategy, roadmap prioritization, and go-to-market execution. Rally engineering and product leadership around a compelling 2-3 year vision for Unistore, identifying untapped market opportunities for the next frontier of hybrid data. Define and drive clear requirements for data movement, storage lifecycle, reliability, cost, and quality of service. Champion AI-driven workflows by defining how Unistore serves as the low-latency transacti

aigorust
View job →

Become a part of our caring community Humana is seeking a self-driven and collaborative Lead Engineer to join our Interactive Voice Response (IVR) team. In this role, you will deliver innovative IVR solutions and develop robust omnichannel APIs for our enterprise platforms. You will have the opportunity to drive the success of a high-impact, customer-facing application within a Fortune 50 company, working closely with multiple teams throughout the software development lifecycle (SDLC). Lead Engineer –Omnichannel Humana is seeking a self-driven and collaborative Lead Engineer to join our Omnichannel team. In this role, you will design, develop, secure, and enhance enterprise APIs that support high-impact, member facing, applications across Humana's digital and voice channels. This role offers the opportunity to modernize and strengthen existing API capabilities while helping deliver resilient, scalable, and secure omnichannel solutions within a Fortune 50 organization. Key Responsibilities Design, develop, and maintain scalable Omnichannel APIs that support enterprise applications and customer-facing capabilities. Enhance the security, resiliency, performance, and reliability of existing APIs through modernization, improved architecture, observability, testing, and operational controls. Apply AI and AI-assisted engineering practices to accelerate development, improve quality, automate testing, enhance documentation, and identify opportunities for optimization. Partner with architecture, security, cloud, product, engineering, and operations teams to deliver secure, resilient, and enterprise-aligned API solutions. Collaborate with agile teams to plan, track, and deliver API enhancements, platform improvements, and cloud-based capabilities. Develop proofs of

machine learningairecruitment
View job →
A
Amplitude
📍 San Francisco• Full-time• $198K – $299K/yr
16 days ago

Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. The Developer Experience (DX) team at Amplitude builds and maintains the foundations that power how developers integrate, extend, and trust Amplitude across platforms. Our mission is to make Amplitude’s SDKs reliable, easy to adopt, and a joy to build on, so customers can confidently instrument their products and unlock insights at scale. We’re looking for a Staff Software Engineer, iOS to play a key technical leadership role on our DevEx team. In this role, you will lead the design and development of Amplitude’s core iOS SDKs, including Analytics and Session Replay , and serve as the iOS platform expert that other SDK teams, such as Statsig, Guides, and Surveys rely on. As a Staff Engineer, you’ll operate with a wide scope and high impact: setting technical direction for the iOS platform, driving cross-SDK architecture, improving performance and reliabilit

gitaiswift
View job →

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . About the Role SonicWall is hiring a Senior Software Engineer to develop and maintain the Content Filtering Service client embedded in SonicOS — the firmware running on SonicWall firewall appliances worldwide. This module intercepts live traffic, queries the CFS rating servers in real time, and enforces URL-based content filtering policies at the packet level. This is security-critical code: a policy bug is a traffic bypass on a deployed security appliance. The codebase is in C with a split Control Plane / Data Plane architecture, multi-blade cache synchronization, and multiple firmware version upgrade paths. Responsibilities Maintain, harden, and extend the CFS client module across both the control plane and data plane. Own the binary UDP protocol implementation between the firewall and the CFS rating servers, ensuring correctness and compatibility across protocol versions. Develop and maintain complete IPv4 and IPv6 support across all code paths. Own the firmware upgrade and migration paths, ensuring they are atomic, failure-safe, and correctly handle all supported version transitions. Improve reliabilit

HI
HP IQ
📍 San Francisco• Full-time• $149K – $240K/yr
16 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role The AI team is building cutting-edge solutions that bring the power of AI directly to edge devices while seamlessly integrating with cloud infrastructure. We are looking for a Senior Software Engineer to design and develop high-performance, scalable services to support AI workloads across edge and cloud environments. What You Might Do Design, build, and maintain services that power AI-driven applications, ensuring scalability and performance. Develop APIs and microservices that facilitate seamless integration between cloud-based AI models and edge devices. Optimize data pipelines and storage solutions for real-time AI inference and processing. Implement security and privacy best practices for distributed AI systems. Work closely with AI researchers, infrastructure engineers, and frontend developers to deliver end-to-end AI-driven solutions. Build and optimize an agent orchestration runtime that enables tool use, memory management, and multi-step reasoning across LLMs, APIs, and edge-connected systems. Develop robust logging, monitoring, and alerting systems to ensure system reliabilit

pythonjavasql
View job →

Job Title Software Technologist - C# .NET Full Stack Job Description C# .NET Full Stack Developer Philips is a global leader in health technology, dedicated to improving lives through meaningful innovation. One of our core businesses, Connected Care , focuses on delivering smarter, data-driven solutions that connect patients, providers, and systems to improve outcomes and efficiency. This role sits within Hospital Patient Monitoring (HPM) , which provides advanced monitoring solutions for acute care settings. From bedside and transport monitors to centralized systems, HPM helps clinicians identify at-risk patients and intervene quickly. The position is based in Bangalore, a key hub supporting innovation and collaboration across Philips. Your role: Participate in the full software development lifecycle, from requirements analysis to deployment Design and develop scalable, secure, and maintainable applications using C# and .NET technologies Implement front-end components and integrate them with backend services Develop and execute unit, integration, and system tests to ensure quality and reliability Conduct code reviews to maintain coding standards and best practices Collaborate with DevOps teams for deployment and monitoring Diagnose and resolve software defects, ensuring optimal performance Create and maintain technical documentation (architecture diagrams, API specs, user guides) Stay updated on emerging technologies and apply innovative solutions Mentor junior developers and contribute to a culture of continuous improvement. You're the right fit if: Experience: Bachelor's Degree in Computer Science, Software Engineering, Information Technology</p

javascriptreactnode.js
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
16 days ago

Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is experiencing rapid growth. You’ll play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features, while also balancing scalability, reusability, and performance. As a Staff Engineer on the Application Engineering team, you’ll serve as a senior technical leader - responsible for designing and delivering high-quality, scalable software systems that power Cohere Health’s core platform. You’ll act as a multiplier, elevating the technical bar for the team, mentoring engineers, and partnering with product, data, and clinical teams to deliver solutions that meet compliance, quality, and performance standards. This role is ideal for engineers who thrive on solving complex problems in healthcare, have deep expertise in building distributed systems, and want to influence architecture and engineering practices at scale. What you’ll do: Technical Leadership & Architecture Define and drive the architecture of large-scale, distributed application systems across the Cohere platform. Ensure solutions are secure, performant, maintainable, and compliant with NCQA, CMS, and payer requirements. Champion engineering best practices in CI/CD, testing, release management, and observability. Hands-On Engineering Write clean, maintainable, and well-tested code, primarily in modern frameworks (e.g., Python, TypeScript/React, Java/Kotlin). Lead the development of core features and APIs that directly impact providers, payers, and patients. Partner with DevOps and Data teams to ensure seamless integration, scalability, and operational readiness. Quality & Compliance Focus Embed automated testing, monitoring, and release safeguards into the development lifecycle. Proactively address compliance and audit-readiness requirements in application

typescriptpythonjava
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
16 days ago

Opportunity Overview: This is a unique opportunity to join a high-caliber software engineering team that is experiencing rapid growth. You’ll play a key role in building impactful healthcare technology on a modern technology stack, with a focus on our core data and AI platforms. Your work will focus on enhancing the platform's key features while also balancing scalability, reusability, and performance. As a Staff Engineer on the Application Engineering team, you’ll serve as a senior technical leader - responsible for designing and delivering high-quality, scalable software systems that power Cohere Health’s core platform. You’ll act as a multiplier, elevating the technical bar for the team, mentoring engineers, and partnering with product, data, design , clinical and payment teams to deliver solutions that meet compliance, quality, and performance standards. This role is ideal for engineers who thrive on solving complex problems in healthcare, have deep expertise in building distributed systems including data solutions, and want to influence architectural decisions for security and scale, drive cross-collaborations for alignment, establish technical standards for consistency and evolve both application and data engineering best practices at scale. What you’ll do: Technical Leadership & Architecture Define and drive the architecture of large-scale, distributed application systems across the Cohere platform. Ensure solutions are secure, performant, maintainable, and compliant with NCQA, CMS, and payer requirements. Champion platform engineering best practices in CI/CD, testing, release management, and observability. Hands-On Engineering Write clean, maintainable, and well-tested code, primarily in modern frameworks (e.g., Python, TypeScript/React, Java/Kotlin). Lead the development of core features and APIs that directly impact providers, payers, and patients. Partner with DevOps and Data teams to ensure seamless integration, scalability, and operati

typescriptpythonjava
View job →
I
Instacart
📍 Canada - Remote (ON, BC• Full-time• Remote• From C$196K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview The Marketing Enablement & Technology (MET) team sits within Instacart's Marketing organization, building the systems, APIs, and data pipelines powering Instacart's Paid Marketing channels—one of our most substantial growth levers. Operating at the intersection of Marketing, Product, and Engineering, we own the platforms behind campaign feeds, audience targeting, event pipelines, and QA automation. We're seeking a Senior Software Engineer to provide technical leadership for the Paid MarTech pod—driving the team's technical roadmap, making key architectural decisions for reusability and scalability, and enhancing team productivity through process evolution and tooling improvements. You'll lead cross-team initiatives, partner directly with senior Product and Marketing leaders on strategic decisions, and serve as a domain expert for marketing technology a

REMOTEpythonaigo
View job →
I
Instacart
📍 United States - Remote• Full-time• Remote• From $230K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview The Marketing Enablement & Technology (MET) team sits within Instacart's Marketing organization, building the systems, APIs, and data pipelines powering Instacart's Paid Marketing channels—one of our most substantial growth levers. Operating at the intersection of Marketing, Product, and Engineering, we own the platforms behind campaign feeds, audience targeting, event pipelines, and QA automation. We're seeking a Senior Software Engineer to provide technical leadership for the Paid MarTech pod—driving the team's technical roadmap, making key architectural decisions for reusability and scalability, and enhancing team productivity through process evolution and tooling improvements. You'll lead cross-team initiatives, partner directly with senior Product and Marketing leaders on strategic decisions, and serve as a domain expert for marketing technology a

REMOTEpythonaigo
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.