Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
27 days ago

About the Team OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems underpin data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences. As frontier models and agents evolve how they use memory, history and snapshots, the underlying architecture increasingly shapes the capabilities products can deliver—and their latency, reliability, cost and efficiency. About the Role We are looking for a technically deep TPM to independently define and lead multiple programs across data platforms, online databases and storage infrastructure. You will connect model, product and data-consumer requirements to architecture, and work with the relevant engineering teams to take new capabilities through production adoption and repeatable expansion. The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The challenge is not simply forecasting more resources: it is making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment and validation. A data pipeline, database query, file operation or execution snapshot can affect whether a product or agent succeeds; you will connect those outcomes to the systems underneath. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate model, product and data-platform needs into precise access patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency. Partner with engineering to transform data and storage architecture into repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together.

awsazurerest
View job →
O
1mo ago

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT engineering org builds and operates the systems that bring product improvements to users across backend services, web, mobile, and desktop platforms. The Developer Velocity team partners with product engineering, platform, infrastructure, reliability, engineering acceleration, and observability teams to make everyday development faster and releases safer, more predictable, and easier to operate. About the Role We are seeking a Technical Program Manager to improve developer velocity and deployment excellence across ChatGPT. You will lead durable improvements to local development, CI, testing, build systems, release trains, progressive rollout, and post-deployment validation. You will identify the highest-leverage sources of engineering friction, align teams around shared standards and metrics, and drive adoption of tooling and workflows that improve both speed and reliability. This role combines systems thinking, technical program leadership, and hands-on operating rigor across a broad engineering surface. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the cross-functional roadmap for improving local development, CI, testing, build workflows, and release infrastructure. Create a durable intake and prioritization mechanism for developer friction, using evidence to focus teams on the highest-impact improvements. Lead programs that improve deployment speed and safety, including pre-merge confidence, progressive rollout, release guardrails, rollback readiness, and post-deploy valida

awsci/cdrest
View job →
O
OpenAI
📍 Washington• Full-time• Remote• From $2M/yr
1mo ago

About the team The OpenAI for Government team is a dynamic, mission-driven group leveraging frontier AI to transform how governments achieve their missions. Our team works to empower public servants with secure, compliant AI tools (e.g., ChatGPT Enterprise, ChatGPT Gov) and mission-aligned deployments that meet government technical requirements with strong reliability and safety. About the role Our Federal Sales team has a unique mission to help government customers understand the transformative impact that highly capable AI models can bring to their agencies and missions. This role combines technical understanding, strategic vision, partnership management, and value-driven strategy tailored specifically to federal customers. You’ll drive key opportunities through the entire federal sales cycle, from pipeline generation to closure. You’ll collaborate closely with researchers, engineers, and solution strategists to help government customers advance their missions through AI. This role is based in Washington DC. We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. In this role, you'll: Manage a focused set of key federal accounts, developing and executing comprehensive federal account plans. Lead federal customers through their AI adoption journey, from consideration to successful deployment. Partner with solutions and research engineering to build and execute complex government customer programs and projects. Own and manage a federal consumption revenue target. Oversee consumption revenue forecasting and reporting. Analyze key federal account metrics and provide insights to internal and external stakeholders. Closely monitor the federal landscape (agencies, policies, competitors, partners, etc.) to inform product roadmaps and corporate strategies. Collaborate cross-functionally with solutions, marketing, communications, business operations, people operations, finance, product management, and engineering. Support the recruitment

REMOTEawsrestai
View job →
N
1mo ago

NVIDIA Networking is a leading provider of innovative end-to-end InfiniBand and Ethernet connectivity solutions for servers and storage. Our portfolio includes adapter cards, switches, cables, and software designed to optimize Data Center performance with industry-leading bandwidth and scalability. We serve diverse sectors such as high-performance computing, enterprise, cloud computing, and Web 2.0. Our mission is to stay ahead of the market by delivering groundbreaking products and services. Our Ethernet solutions are tailored for industries like Media & Entertainment and any domain that benefits from advanced DataStream and TCP/IP acceleration. What You’ll Be Doing: Lead a team of 8+ mechanical design engineers. Define priorities, create project plans, and allocate resources for mechanical programs in coordination with Product Managers. Drive all electro-mechanical, automated JIG and thermal design aspects, of production test setups, ensure readiness of test and assembly infrastructure for high-volume manufacturing. Develop multiple early design concepts in fast-paced product development cycles. Lead task forces to investigate and resolve production issues, reliability concerns, and conduct failure analysis. Perform risk assessments and implement mitigation strategies during product design. What We Need to See: B.Sc. in Mechanical Engineering or higher. 10+ overall years of relevant experience including 4+ years of experience managing teams of engineers in R&D environment. Proven expertise in developing, testing, and manufacturing of complex automated connection systems with precise moving parts, pneumatic and electro-mechanical systems design. Strong hands-on experience with 3D CAD tools (Creo preferred) static and dynamic mechanical simulation Solid background in designing components and sub systems an

R
Roblox
📍 San Mateo• Full-time• From $263.7K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. WHY DATA SCIENCE & ANALYTICS? The Data Science & Analytics organization's mission is to increase our speed, frequency and acumen of making decisions at scale by instilling a data-influenced approach to building products. We cover a wide area of the data spectrum including analytical data engineering, product analytics, experimentation, causal inference, statistical modeling and machine learning. Aligned and partnering with product groups, we use this vast tool belt to discover new opportunities and unmet use cases, influence and shape the product roadmap and prioritization, build data products and measure the impact on our community of players and developers. WHY Consumer Apps? Roblox is used by tens of millions of people every day across a wide range of devices, including mobile, desktop, console, TV, and VR. The Consumer Apps team owns the app foundation that makes Roblox feel fast, fluid, and reliable for players and is dedicated to delivering superior performance, reliability, and user experience across all platforms Roblox supports. This team ensures a seamless and engaging user interface that facilitates intuitive interactions while enabling efficient, high-quality experiences

pythonsqlaws
View job →
AG
Adani Group
📍 Ahmedabad• Full-time
1mo ago

We are seeking a detail-oriented and analytical Data Steward – Operational Technology (OT) to join our Enterprise Data team. The ideal candidate will have a minimum of 3 to 5 years of experience in data management, data quality, and operational technology datasets. In this role, you will be responsible for managing data quality, cataloguing, asset modeling, and data standardization across operational systems (PLC, DCS, SCADA, IoT) and modern cloud data platforms like Databricks. You will work closely with plant operations, IT, data engineering, and analytics teams to ensure high data reliability, consistency, and alignment with enterprise data governance standards. Source: Adani Group | Job ID: 56028

AG
1mo ago

The Mechanical Foreman is responsible for overseeing the installation, maintenance, and repair of all mechanical systems and equipment within the mine. This role ensures that all mechanical work is conducted in compliance with the Coal Mines Regulation of 2017, maintaining high standards of safety, reliability, and efficiency. The Mechanical Foreman coordinates with engineers, technicians, and mine management to implement mechanical projects, troubleshoot issues, and ensure the continuous operation of mining machinery and equipment. Source: Adani Group | Job ID: 51950

AG
Adani Group
📍 Singrauli• Full-time
1mo ago

The Electrical Foreman is responsible for overseeing the installation, maintenance, and safe operation of all electrical systems and equipment within the mine. This role ensures that all electrical work is conducted in compliance with the Coal Mines Regulation of 2017, maintaining high standards of safety, reliability, and efficiency. The Electrical Foreman coordinates with engineers, technicians, and mine management to implement electrical projects, troubleshoot issues, and ensure continuous power supply to support mining operations. Source: Adani Group | Job ID: 51897

P
Pendo
📍 Raleigh• Full-time• $150K – $165K/yr
1mo ago

The Team + The Role The Product Design team at Pendo works closely with senior product leaders and sits at the center of product strategy and delivery. The team operates close to both the problem and the build, using LLM-assisted exploration, code-level prototyping, and customer validation before work moves into Figma for systems and polish. Designers ship across a multi-app platform with a high bar for craft, reliability, and product impact. As a Staff Product Designer, you will lead design across the platform and move toward the highest-impact work. You will own projects end-to-end, from problem framing through shipping, partnering with product and engineering to make sophisticated experiences feel simple. AI-assisted workflows are part of how this team works every day, and this role is expected to use, improve, and share those workflows. This role is based in our Raleigh office. What this looks like day-to-day Platform design leadership: Lead design across multiple product areas and own projects from problem framing through shipping. Partner with product and engineering to reduce complexity without losing capability, starting from customer need rather than feature requests. Prototyping and validation: Build interactive prototypes using code-adjacent tools and validate them with real users before engineering investment. Use findings to sharpen flows, surface edge cases, and improve the work before full systemization. AI-integrated workflow: Use AI tools actively across the design process, including LLMs for problem framing and brief development, Cursor and Claude for prototyping, and generative tools for faster interface exploration. Document and share prompts, patterns, and workflows with design and engineering partners. Design systems: Translate and evolve Pendo's design principles into reusable components and patterns. Identify when to extend existing systems and when to push them forward, then drive those decisions with stakeholders. Cross-functional alignment

D
1mo ago

As a Senior Platform Product Manager focused on AI SDLC Trusted Throughput, you will define and drive the product strategy for enabling safe, reliable software delivery at AI-native scale across Datadog’s Internal Developer Platform. As AI accelerates development velocity and system complexity, you will help evolve SDLC systems from human-supervised workflows to platforms with built-in safety, observability, and correctness guarantees. You will partner closely with engineering, security, and developer platform teams to improve deployment reliability, operational visibility, and governance while enabling both engineers and AI agents to move quickly with confidence. This role offers the opportunity to shape foundational developer infrastructure and influence how AI-powered software delivery operates across Datadog. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the product strategy, roadmap, and execution for AI-native SDLC throughput and reliability initiatives across Datadog’s Internal Developer Platform Define and drive platform outcomes aligned to DORA metrics, balancing deployment velocity with reliability, change failure reduction, and operational safety Partner with engineering, infrastructure, security, and developer experience teams to build automated validation, auditability, and risk-scoring capabilities into deployment workflows Deliver actionable SDLC observability and diagnostic capabilities that connect executive-level metrics to operational signals across the software delivery lifecycle Drive systems that monitor and validate AI-generated or AI-attributed changes to ensure correctness, compliance, and trustworthy automation Serve as a cross-functional product leader across SDLC Foundations, Security Engineering, and compl

aigorust
View job →
S
1mo ago

About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe Terminal helps our users extend their online presence to the physical world. The Terminal team’s mission is to make it as easy for businesses to accept in-person payments as the Stripe API has done for online payments. With Terminal, businesses can unlock in-person payments use cases that are right for their business model—whether it’s creating a superb retail experience, extending their website to a pop-up store, or enabling a mobile point-of-sale at their next event. We’re looking for an experienced product manager to lead and shape the future of the software platform that powers Stripe Terminal’s devices portfolio. Working closely with your engineering partners you will design, build and launch device software capabilities that delight users and differentiate Stripe’s solutions in the market. Working closely with our hardware experts you will build the multi-year strategy for how Stripe will continue to enhance the scalability, reliability, usability, and market differentiation of the software capabilities of our in-person commerce devices. In this role, you will work with a spectrum of users, from our largest platforms to small start ups as well as external technology partners to deeply understand user needs and market trends. You will obsess over stability, scalability, and expanding our product while delivering the best in person experiences in the world. What you'll do: Set a motivating multi-year vision for device softw

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The People Technology team builds and operates the systems that support how OpenAI hires, develops, and supports its people. The team brings together People Systems and People Innovation Labs, a product engineering group focused on rethinking how we find and retain exceptional talent and help employees do their best work. People Systems owns the company’s core people-technology ecosystem, including platforms such as Workday and Ashby. People Innovation Labs builds new employee and People Team experiences on top of that foundation, including OpenHouse, our internal employee hub, and AI-powered products and automations. Together, we are working toward a model in which our enterprise systems provide reliable data, controls, and core business logic, while employees and managers can complete more of their work through simple, integrated, and AI-native experiences. About the Role We are looking for a People Systems Lead to manage the People Systems team and shape how our core systems evolve. You will be responsible for the reliability and effectiveness of our current environment while helping us move beyond the constraints of traditional enterprise software. This includes stabilizing and improving platforms such as Workday and Ashby, designing the integrations that connect them to the broader technology ecosystem, and partnering with People Innovation Labs to surface workflows through OpenHouse, Slack, and AI-powered experiences. This role requires someone who is comfortable moving between strategy, technical design, and team leadership. You should understand People systems deeply, be able to work through integration and architecture decisions with engineers, and translate complex organizational needs into scalable solutions. You will also manage vendor relationships, develop the People Systems team, and drive alignment across People, Engineering, Finance, Security, Legal, and other partners. This role could be a fit for someone who has grown up in People S

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Role As a Sales Manager, Energy, you will build and lead a team of Account Directors focused on strategic growth across utilities, oil and gas, renewables, power generation, and energy services. The team will partner with complex organizations modernizing operations, improving reliability, accelerating the energy transition, and adopting enterprise AI responsibly at scale. You’ll help the team navigate regulated enterprise sales cycles, deepen relationships with business, technology, operations, engineering, security, and risk leaders, and drive adoption of OpenAI’s platform across safety-conscious, asset-intensive organizations. Key Responsibilities Recruit, develop, and lead a high-performing team of Energy Account Directors. Create a strong coaching culture through deal reviews, account strategy sessions, ride-alongs, and structured 1:1s. Define the Energy GTM strategy, including subsector segmentation, account prioritization, partner strategy, executive engagement, and territory planning. Drive disciplined pipeline generation, forecast accuracy, and operational rigor. Guide multi-stakeholder opportunities involving operations, engineering, digital, data, security, legal, risk, procurement, and executive leadership. Help customers translate AI and API capabilities into measurable outcomes across asset and field operations, grid and generation planning, engineering knowledge, customer service, commercial workflows, and enterprise productivity. Partner with Product, Solutions Architecture, Technical Success, Legal, Security, Finance, and policy experts to support responsible deployment. Provide structured feedback on customer requirements, integration blockers, reliability and governance needs, and emerging industry trends. What We’re Looking For 15+ years of enterprise sales, GTM, or sales leadership experience. Proven experience building and scaling enterprise sales teams responsible for complex strategic accounts and large revenue targets. Deep underst

awsgitrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Core Models team helps shape how OpenAI’s frontier models are built, measured, and launched. We work across Research, Engineering, Model Design, Data Science, and Product to turn advances in model capabilities into reliable, useful experiences for people. Our scope includes model planning and launches as well as building data flywheels, evaluations and measurement systems to ensure our models have strong capabilities and behavior. About the Role As a Product Manager for the Core Models team, you'll be at the forefront of defining and guiding the future of how our AI models work in real-world applications. You will connect user needs to model and systems decisions: how prompts are understood; how information is aggregated and made useful for training and evaluation data; and how capabilities move from research prototypes into the mainline model and launch stack. You will operate comfortably across research, infrastructure, and consumer product surfaces, creating clarity where ownership and technical boundaries are still emerging. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate user and product goals into clear model requirements, system architecture choices, and research priorities across query understanding, indexing, retrieval, ranking, tool boundaries, data, training, inference, and evaluation. Build closed learning loops that turn product usage, explicit feedback, and other user signals into datasets, evaluations, experiments, training priorities, and launch decisions. Define success across offline evaluations and online product metrics, balancing model quality, usefulness, latency, safety, reliability, and cost. Partner closely with post-training research, applied product engineering, Model Design, and Data Science to integrate capabilities into the mainline model stack. Create reusable platforms and operatin

awsrestai
View job →
O
1mo ago

About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation

awsrestai
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.