Jobs in United States

Reliability Engineer in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 Chicago, Illinois, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $90K/yr

Quick readStrong listing-quality and freshness signals

As an Enterprise Customer Success Manager, you will proactively drive new product attachment and effective strong relationships across our largest and most strategic customers. You’ll advocate for the customer internally and focus on a positive customer experience. Interactions are rooted in relationship-management, first and foremost, while also advocating for growth opportunities. Enterprise Customer Success Managers follow a well-defined methodology that helps them identify the customer's unique needs and clearly convey the value of the Datadog product. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as a strategic partner to customers, orchestrating cross-functional internal teams and engaging executive, technical, and business stakeholders to understand customer goals and translate them into a clear, deliverable Datadog value narrative. Proactively build and maintain executive relationships to deliver clear, outcome-driven value stories that connect Datadog technical use cases to measurable business results. Lead QBRs and strategic reviews as a forum to demonstrate impact, align on priorities, and define next-step initiatives. Analyze adoption and usage trends to quantify value delivered, extract insights from large datasets, identify gaps, and drive financially grounded commercial recommendations and strategic opportunities. Position Datadog as a critical observability platform that enables reliability, efficiency, and informed decision-making. Own and project manage the on-boarding process for new customers Collaborate cross-functionally with AEs, SEs, TAM, Product, Support, Enablemen and other technical teams to ensure consistent value delivery and messaging. Who You Are: Customer-centric with 3+ years in a Customer Success

AIGoRustDevOps
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s Account Verification team builds the foundation of trust for open finance. We help fintechs and financial institutions connect and verify bank accounts securely so that money can move safely and instantly. Account Verification is the entry point for all of Plaid’s payment-focused consumer experiences and is a critical building block for our customers. The team obsesses over creating seamless verification journeys that balance speed, reliability, and security, enabling consumers to confidently connect to the financial ecosystem. As a PM for Account Verification, you’ll own one of Plaid’s most critical and high-impact product areas. You’ll lead the evolution of our verification platform across Auth, Balance, and Identity Match, defining how millions of people and businesses connect their financial accounts every day. We are looking for a high-ownership builder who thrives in ambiguity, loves building with customers, and is excited to define what’s next for one of Plaid’s most established and strategically important product lines. You’ll set vision and strategy, drive execution across a cross-functional team, and shape how Plaid competes in an increasingly complex and, eventually, AI-driven ver

AWSAIGoRust
M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Own day-to-day operations for the scanner program and scanner builds in the spa: purchasing/procurement, vendor management, receiving, inventory, and logistics. Stand up lightweight production operations as we move from prototypes to repeatable builds: build planning, kitting, work instructions, and readiness checklists. Partner with Quality to ensure the operational system supports compliance: traceability, document control, training records, NCR/CAPA workflows, and audit readiness. Drive cross-functional execution for the physical spa build-out and scanner integration: schedules, dependencies, risk register, and weekly coordination with vendors and internal teams. Own the “integration glue” across facilities + device ops: commissioning plans, acceptance criteria, and operational handoff (runbooks, maintenance, spares, escalation paths). Build and track operational metrics: cost, budget, lead times, vendor performance, build throughput, and reliability of critical subsystems. Audit of import/export documentation, management of contract renewals, regulatory compliance. What we’re looking for Proven operations leadership in hardware/medical/robotics (or similarly complex electromechanical products), including procurement and vendor management. Strong program management instincts: can run schedules, unblock cross-functional dependencies, and keep priorities clear under ambiguity. Comfort operating in quality/regulatory environments and building processes that are rigorous without slowing a small team. High ownership and bias to action: can jump between spreadsheets, docks, and the lab/site to keep the program moving. Useful experience Experience running prototype-to-production transitions (NPI, EVT/DVT/PVT-style builds, CM/EMS collaboration). Facilities / construction operations experience (GC coordination, MEP commissioning, site readiness). Familiarity with inventory systems and procurement tooling (even “simple but disciplined”).

G
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -97.9%

From $126.4K/yr

Quick readStrong listing-quality and freshness signals

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Staff Human Resources Information Systems Analyst, People Technology - Workday, you won't just manage our people systems—you'll own them. You'll drive how our people technology evolves, partner deeply with stakeholders to solve root problems, and build scalable solutions that reduce friction and improve reliability, usability, and insight across the People technology landscape. This is a strong fit if you think like a product owner, refuse to accept requests at face value, and can move with speed and precision within public company compliance requirements, including SOX ITGCs. In this role, you'll le

GitRestAIGo
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $280.5K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As the Senior/Principal Product Manager for Engine Systems Foundations, you will drive the vision and strategy for the most foundational parts of the Roblox game engine and be hands-on with the execution and delivery of products that impact over 130 million players every day. This team is responsible for the core performance, reliability, and efficiency of the engine, and it owns key features like our memory allocation library, thread/work dispatch system, and the backing APIs that power our creator performance tooling. If you are a visionary product leader who thrives on deeply technical challenges to improve the speed and quality of a system, you’ll be a great fit! The role is based in San Mateo, CA (hybrid with Tues-Thurs onsite). You will: Define the long-term vision and strategy for Systems Foundations, ensuring we have plans in place to continually invest in the core pieces of a high-performance, realtime game engine. Take ownership of the engine-related content in the public Creator Analytics creators use to monitor the experiences on Roblox, ensuring we’re delivering actionable insights. Work with the Creator organization to define and drive end-to-end performance workfl

AWSGitAIC++
O
📍 Washington, District of Columbia, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's mission is to ensure that artificial intelligence benefits all of humanity. OpenAI for Government works with U.S. and allied government institutions to support the responsible adoption of AI across defense, intelligence, federal civilian, state and local, and international public-sector missions. We work at the intersection of technology, policy, operations, security, and delivery to help public servants use frontier AI in ways that are effective, trusted, aligned with democratic values, and grounded in real-world consequences. The team helps government agencies transform how they work through secure, compliant AI tools and mission-aligned deployments, including ChatGPT Enterprise, ChatGPT Gov, APIs, Codex, and emerging frontier capabilities. We partner with government leaders, operators, technologists, and policy stakeholders to translate cutting-edge AI into measurable mission impact while meeting government requirements for safety, reliability, security, compliance, and trust. Cyber is a critical government mission area. OpenAI's government cyber work brings together OpenAI for Government, Product, Research, Security, Product Policy, Legal, Global Affairs, and Communications to help trusted public-sector defenders responsibly use AI to protect government networks, critical infrastructure, and national-security systems. About the Role We are seeking a senior government cyber leader to serve as Head of Government Cyber Integration for OpenAI for Government. This role will integrate OpenAI's government cyber work across strategy, testing and evaluation, trusted access, deployment, policy, security, and external engagement. This is a matrix leadership role, not a replacement for line management. Product teams still own product roadmaps. Research and Safety still own model capability measurement and risk mitigation. Security still owns OpenAI's security posture and customer security requirements. Product Policy, Legal, and Global Affairs still

AWSRestAIGo
O
📍 United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI is evaluating multiple infrastructure pathways, including powered land, colo/BTS, and NeoCloud opportunities. The Site Readiness & Development team provides the diligence layer needed to compare opportunities, identify risk, and support credible deployment decisions across those pathways. About the Role The NeoCloud & Colo Due Diligence Lead will evaluate third-party infrastructure opportunities where OpenAI is considering deployment through NeoCloud, colo, or BTS structures. This role will focus on facility and deployment readiness, including MEP readiness, rack strategy, developer capability, facility design, power deliverability, schedule credibility, and operating assumptions. Unlike the land diligence team, this role is centered on technical and operational readiness of third-party infrastructure rather than greenfield site master planning, civil development, and entitlement strategy. This is an individual contributor lead role and does not have direct reports initially. The role determines whether each opportunity is fit-for-use and fit-for-service against OpenAI facility, rack, power, network, reliability, and operational standards; identifies material deficiencies and tracks remediation with developers/operators; and evaluates commissioning, validation, AHJ/code, and deployment interfaces such as structured cabling, network readiness, and high-density rack support where relevant. Key Responsibilities Lead diligence on NeoCloud, colo, and BTS opportunities across technical and operational readiness dimensions. Assess each opportunity against OpenAI facility, rack, power, network, reliability, and operational standards to determine deployment fit. Validate MEP readiness, rack deployment strategy, facility design assumptions, power deliverability, and schedule credibility. Identify material deficiencies and work with developers/operators to define remediation plans, owners, timing, and residual risk. Review reliability, availabilit

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s People Experience & Technology (PXT) team owns the core people platform that powers worker, recruiting, contingent, approvals, and lifecycle workflows across the company. PXT is responsible for operating Workday, Ashby, and related people systems as governed, reliable sources of truth, while building the controls, monitoring, documentation, and auditability required to support scale. About the Role We’re hiring an Enterprise Systems Manager, Recruiting Systems to help own and harden OpenAI’s recruiting platform, with a focus on Ashby and its connected workflows. This is a hands-on systems role for someone who can translate recruiting process problems into governed, durable fixes through configuration, workflow design, access controls, documentation, reporting guardrails, and integration partnership. You will work at the boundary of Recruiting, HR Operations, Legal, Compensation, Analytics, IT, and PXT to improve the reliability and control health of recruiting workflows. The right person is comfortable going deep in system design while also driving rollout, adoption, and operational clarity. In this role you will: Own specific recruiting workflow domains in Ashby and adjacent tools, including stages, fields, permissions, approvals, templates, and configuration standards. Partner on high-priority remediation work across start dates, offers, approvals, integrations, auditability, data integrity, and workflow controls. Design and implement governed workflow changes that balance recruiter usability with reporting trust, downstream integration reliability, and control requirements. Establish and maintain guardrails such as required and conditional fields, stage definitions, role-based permissions, approval logic, validation patterns, and change standards. Drive durable fixes for recurring operational issues by identifying root causes and resolving them through configuration, automation, documentation, or process redesign. Partner with PXT, IT,

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Infrastructure team is central to this mission, setting the core strategy and implementing the vision. From site selection to deployment to operations, this team sits at the intersection of commercial, technical, and operational domains, interacting with experts and executives inside and outside of OpenAI. We design and operate mission-critical facilities that support cutting-edge AI workloads at scale. About the Role We are seeking a Facilities Operations Lead to support the commissioning, deployment, and long-term operation of our next-generation AI data centers. This role bridges the interface between data center construction and hardware landing, ensuring seamless integration of mission-critical infrastructure with hardware deployment timelines. You will define and execute commissioning plans, support infrastructure bring-up, and take ownership of operations and maintenance for cutting-edge, large-scale, AI data centers. You will collaborate closely with design, construction, and hardware teams to define repeatable processes for new data center builds and lead hands-on operations to uphold the performance and reliability of our deployed infrastructure. Key Responsibilities Define and execute sequences of operations, commissioning steps, and bring-up processes for mission-critical data center facilities. Interface with the design and hardware teams to define deployment procedures tailored to each data center and hardware configuration. Oversee installation, commissioning, and operational readiness of large-scale data center campuses. Manage monitoring, maintenance, and quality control of the data center infrastructure, including high-performance liquid cooling systems. Develop on-site operations staffing strategy. Develop and enforce procedures for planed and unplanned downtime and SLAs for critical

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions tailored to the demands of advanced AI workloads. We work across the full stack—from silicon to system integration—partnering closely with internal teams and external vendors to define and deliver next-generation AI infrastructure. Our team focuses on defining scalable, high-performance system architectures and reference designs that balance performance, cost, and operational efficiency across rapidly evolving technologies. About the Role We are seeking a 3P Architect to define and drive rack- and cluster-level reference designs in collaboration with external partners. This role is responsible for translating workload requirements and system-level goals into concrete architectures, aligning partners on critical design attributes, and ensuring vendor roadmaps meet our infrastructure needs. You will work closely with performance modeling and internal architecture teams to evaluate tradeoffs, while owning the end-to-end definition and execution of third-party system designs. This includes identifying gaps in current technologies, driving vendor development, and shaping future infrastructure capabilities. This role requires strong system intuition, cross-functional leadership, and the ability to operate effectively across internal teams and external ecosystems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Define rack- and cluster-level reference architectures for AI infrastructure deployments. Translate workload requirements into clear system design specifications and partner deliverables. Collaborate with performance modeling teams to evaluate architectural tradeoffs and system behaviors. Align internal stakeholders and external partners on critical system attributes (performance, cost, power, reliability, scalability). Identify gaps in current technology offerings and dr

AWSRestAIGo
🔔

Get new reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime