Jobs in United States

Reliability Engineer in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Industrial Compute team is central to this mission, setting the core infra strategy and implementing this vision. From site selection to the buildout process, this team sits at the intersection of commercial, technical, strategy, and operations, interacting with teams and executives inside and outside of OpenAI. About the Role The Clean Energy and New Technology Lead will own infrastructure clean energy and emerging energy technology strategy and execution, identifying and deploying scalable solutions that enable resilient, low-carbon compute and data center growth. The role will work closely with regulatory and policy teams to align infrastructure expansion with OpenAI’s long-term environmental and operational objectives. This is an individual contributor lead role and does not have direct reports initially. The role will evaluate where emerging energy technologies can materially improve reliability, cost, carbon, speed, or resilience; translate those options into practical deployment pathways; and help ensure OpenAI’s infrastructure growth remains aligned with sustainability considerations. Key Responsibilities Evaluate emerging energy solutions such as clean firm power, advanced storage, grid flexibility, low-carbon backup power, heat reuse, water-related energy efficiency, and other scalable technologies where relevant. Identify pilot opportunities and deployment pathways that can move promising energy technologies from concept to commercially and operationally credible execution. Translate technical options into clear reliability, cost, schedule, carbon, regulatory, and operational implications for infrastructure decision-making. Partner with energy regulatory, policy, procurement, engineering, deployment, finance, legal, and site-readiness teams to align technology and sustainability choices with

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, you will improve the capabilities, reliability, and product fit of OpenAI's agentic models. You might own a research direction, build the infrastructure that makes large training runs faster and more trustworthy, create evals that reveal where models fail, or drive a capability from an idea through experimentation, integration, and launch. This role is intentionally broad. The strongest candidates are not defined by one method or subfield; they are people who can take an ambiguous capability problem and make progress across research, engineering, data, evals, and product. You should be excited to work on models that act in the world: writing and debugging code, using tools, calling functions, operating computers, collaborating with other agents, and completing valuable work on behalf of users. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, meas

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of this API & power-users team, you will improve the capabilities, reliability, and product fit of OpenAI’s agentic models for power users and API developers. You might design evals from real developer workflows, build training environments around production-like tool use, turn qualitative model failures into training data, evals, or post-training interventions, or drive a behavior improvement from discovery through post-training, integration, and launch. This role is intentionally broad. The strongest candidates are comfortable turning ambiguous model behavior problems into concrete progress, whether that means improving tool use, planning, instruction following, recovery from mistakes, or how models behave in API-based workflows. You should be excited to work across research, engineering, data, evals, and product to make models better at acting in real workflows. You will work closely with researchers, engineers, API/product teams, Codex, infrastructure, and safety/align

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Computer Use, you will teach models to operate computers. You will help train models that can navigate browsers and desktops, use tools and applications, reason through complex workflows, collaborate with users and other agents, and complete long-horizon tasks with reliability and judgment. This work sits at the intersection of frontier model training, product behavior, evaluation, and systems engineering, and will directly shape the computer-use capabilities shipped in OpenAI’s next generation of agents. Currently, our models are the best in the world at this behavior! You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might Design and run experiments th

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Operations Program Manager (OPM) to serve as the single-threaded operational leader for new hardware introductions (NPI) and production ramps across OpenAI’s AI infrastructure systems. This role combines hands-on execution with strategic ownership. You will be responsible for defining the operating model, aligning cross-functional stakeholders, setting the critical path, making informed tradeoffs, escalating decisively, and ensuring hardware programs deliver on schedule, quality, cost, and scalability. Success in this role requires comfort operating in ambiguity, influencing without authority, and driving alignment across internal teams and external partners—while keeping eyes firmly on long-term system scalability and repeatability. In this role, you will: Strategic & Leadership Ownership Act as the single-threaded owner for operational readiness across NPI and ramp, accountable for outcomes from early bring-up through sustained production Translate OpenAI’s infrastructure strategy and engineering objectives into clear operating plans, execution priorities, and decision frameworks Drive alignment across Engineering, Operations, Strategic Sourcing, Finance, Capacity Planning, and Executive stakeholders by framing tradeoffs, risks, and recommendations Proactively identify inflection points where decisions or investments are required to protect long-term scale, reliability, or cost targets Influence operational strategy with manufacturing par

AWSRestAIGo
P
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity As a Member of Technical Staff on AI Infrastructure, you will build and maintain the foundational systems and distributed infrastructure that power AI model post training, inference, and data pipelines. You will collaborate with engineering and research teams to ensure performance, scalability, and reliability of critical AI systems. What You’ll Do Design and implement large-scale, distributed AI infrastructure and services Optimize performance for GPU/xPU accelerators and cloud environments Build tools for observability, reliability, and scaling of AI workloads Partner with cross-functional teams to define AI infrastructure requirements and roadmap Contribute to architectural design and system longevity About You Have experience with GenAI infrastructure systems, distributed systems, cloud computing, and high-performance infrastructure Are proficient in programming languages like Python, Go, or similar Understand scaling challenges specific to AI workloads and accelerators Thrive in fast-paced, collaborative engineering environments The reasonably estimated base salary for this role ranges from $256,000.00 to $276,

PythonAIGoRust
S
📍 Bellevue, Washington, United States· Full-time· Hybrid
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Unistore (Hybrid Tables) is Snowflake's strategic convergence investment—seamlessly combining OLTP and analytical workloads in a single unified database, and one of the company's most strategically differentiated bets. Operating like a startup within Snowflake, you will own revenue goals, define product strategy for the next phase of growth following our recent General Availability (GA) milestones, and coordinate go-to-market as the primary product authority for the area. OLTP-HTAP convergence is a fast-moving, high-stakes space with strong customer demand—making this one of the highest-visibility product leadership roles in the company and a rare opportunity to define the future of transactional and analytical convergence at scale. AS A STAFF PRODUCT MANAGER AT SNOWFLAKE, YOU WILL: Drive the Unistore product area end-to-end—including product strategy, roadmap prioritization, and go-to-market execution. Rally engineering and product leadership around a compelling 2-3 year vision for Unistore, identifying untapped market opportunities for the next frontier of hybrid data. Define and drive clear requirements for data movement, storage lifecycle, reliability, cost, and quality of service. Champion AI-driven workflows by defining how Unistore serves as the low-latency transacti

H
📍 Louisville, United States
✓ Quality checkedCompany trend +310%

Become a part of our caring community Humana is seeking a self-driven and collaborative Lead Engineer to join our Interactive Voice Response (IVR) team. In this role, you will deliver innovative IVR solutions and develop robust omnichannel APIs for our enterprise platforms. You will have the opportunity to drive the success of a high-impact, customer-facing application within a Fortune 50 company, working closely with multiple teams throughout the software development lifecycle (SDLC). Lead Engineer –Omnichannel Humana is seeking a self-driven and collaborative Lead Engineer to join our Omnichannel team. In this role, you will design, develop, secure, and enhance enterprise APIs that support high-impact, member facing, applications across Humana's digital and voice channels. This role offers the opportunity to modernize and strengthen existing API capabilities while helping deliver resilient, scalable, and secure omnichannel solutions within a Fortune 50 organization. Key Responsibilities Design, develop, and maintain scalable Omnichannel APIs that support enterprise applications and customer-facing capabilities. Enhance the security, resiliency, performance, and reliability of existing APIs through modernization, improved architecture, observability, testing, and operational controls. Apply AI and AI-assisted engineering practices to accelerate development, improve quality, automate testing, enhance documentation, and identify opportunities for optimization. Partner with architecture, security, cloud, product, engineering, and operations teams to deliver secure, resilient, and enterprise-aligned API solutions. Collaborate with agile teams to plan, track, and deliver API enhancements, platform improvements, and cloud-based capabilities. Develop proofs of

Machine LearningAIRecruitment
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA’s Silicon Co-Design Group is the team that gets every GPU, SoC, and CPU silicon program from first power-on to high-volume production. We are hiring a Senior Manager to lead our Test, Manufacturability, Reliability & Quality (TMRQ) organization. This is not a coordination role . Your work decides if a product can be built at scale and trusted in the field. These include production test development (SLT, BLT), control run flow, system reliability stress (HTOL), platform- and board-level manufacturing issue closure, and field diagnostic test development. You lead a team of individual contributors and a first-line manager at the layer where silicon, platform, and software collide with manufacturing reality. Decisions you make show up in yield curves, production ramp , and customer escapes. You are the leader the program turns to when a build is stuck, a control run is fallout-heavy, or a field return points back at silicon . The exceptional hire also uses AI deliberately — with demonstrated workflow impact and the judgment to know where it compresses real work and where it introduces risk. What you will be doing: Keep programs moving. Own the technical execution and velocity of SLT, BLT, Board/Chip/Rack CR, and system reliability stress (HTOL) across every GPU, SoC, and CPU silicon program. Close the hardest multi-functional failures. Resolve Vmin and binning escapes, performance shortfalls, and power anomalies by driving root-cause across design, methodology, DFT, ATE, package, software/firmware, and manufacturing — and own the WARs and productized fixes through to confirmation. Give leadership the clarity to act. Convert raw integration signals — CR fallout, BLT/SLT yield, SHTOL/CHTOL data, RMA trends, customer escalations — into decision-ready options that enable executive leadership to act with confidence on POR, QS/PS gates, a

O
📍 Washington, District of Columbia, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The OpenAI for Government team is a dynamic, mission-driven group leveraging frontier AI to transform how governments achieve their missions. Our team works to empower public servants with secure, commercial and specialized national-security models that address government's mission needs, meet their technical requirements, and provide the highest degree of reliability and safety. As part of the OpenAI for Government team you will work with U.S. intelligence community and international public sector customers, helping to accelerate adoption through hands‑on support, tailored training, and early insights—so civil servants can spend less time on red tape and more on meaningful work. If you're passionate about responsibly deploying frontier AI to uplift public institutions and transform how government works, we want you on our team. About the Role Our OpenAI for Government team has a unique mission to help government customers understand the transformative impact that highly capable AI models can bring to their critical missions. This role combines technical understanding, strategic vision, partnership management, and value-driven strategy specifically tailored to government customers. As a Capture Manager, you’ll identify and drive key opportunities through the government sales cycle, from pipeline generation to closure. You’ll collaborate closely with industry partners, potential customers, policy makers, and internal experts to help government customers advance their missions through AI. This role is based in Washington, DC. In this role, you will: Pipeline Development : Partner with O4G GTM teams to identify, build, and maintain a multi-year pipeline of strategic opportunities. Capture Management : Lead the full capture lifecycle — from opportunity identification and qualification through win strategy, teaming, proposal, and award. Content Development : Write and edit proposal sections, weaving together inputs from subject matter experts into cohesive

AWSRestAIGo
F
📍 Mclean, Virginia, United States
✓ High-confidence listingCompany trend -26.7%
Quick readStrong listing-quality and freshness signals

At Freddie Mac, our mission of Making Home Possible is what motivates us, and it’s at the core of everything we do. Since our charter in 1970, we have made home possible for more than 90 million families across the country. Join an organization where your work contributes to a greater purpose. Position Overview: Are you an experienced HR technology leader passionate about operational excellence, workforce technology strategy, and delivering exceptional employee experiences? Freddie Mac is seeking a Workforce Technology Operations Manager to lead the team responsible for the technology operations, workforce data, governance, and service delivery that support critical workforce programs across the enterprise. In this leadership role, you will oversee the reliability, security, and performance of workforce technology services while driving continuous improvement, managing operational risk, and enabling scalable solutions that support business priorities and employee outcomes. Our Impact: We enable the technology, data, and operational capabilities that support Freddie Mac's workforce and employee experience. Through strong governance, reliable service delivery, and trusted workforce data, we help ensure workforce technology solutions operate securely, efficiently, and effectively across the enterprise. Partnering across HR, Technology, Risk, and business teams, we drive operational excellence, strengthen controls, improve employee experiences, and deliver technology solutions that support organizational success. Your Impact: In this role, you will: Lead and develop a team responsible for workforce technology operations, security administration, workforce data governance, and service delivery. Manage the production operations of workforce technology applications, including platform reliability, integrations, releases, issue resolution, and ongoing support.

H
📍 Dallas, TX, United States
✓ High-confidence listingCompany trend +310%
Quick readStrong listing-quality and freshness signals

Become a part of our caring community We are looking for a highly motivated Senior Technology Leadership professional to join our IT Operations team. You will support the Lead of the Application Operations Center and Enterprise Post Production Validation teams, with a focus on process automation, innovation, and continuous improvement. You will work with both onshore and offshore resources, ensuring in daily tasks, enhancing application support, and driving improvements in monitoring and validation processes. Main Responsibilities: Collaborate with the Lead to provide operational and strategic support for Application Operations Center and Post Production Validation teams. Identify, evaluate, and implement automation opportunities to increase efficiency and reduce manual workload. Drive process innovation by recommending and deploying advanced tools and methodologies for application operations and validation activities. Analyze existing workflows and develop documentation for standard operating procedures and best practices. Oversee daily application support and monitoring activities to ensure system stability, performance, and reliability. Partner with onshore and offshore teams to coordinate task execution and promote consistent adoption of new processes and technologies. Develop and maintain dashboards and reports to track key performance indicators and present findings to leadership. Ensure automation and process improvements comply with organizational standards and regulatory requirements. Facilitate knowledge sharing, training sessions, and change management activities to support team development and successful project implementation. Engage with stakeholders to gather requirements, understand challenges, and communicate progress on automation initiatives. Ability to create

PythonJavaAzureAI
JI
📍 Shrewbury, United States
✓ High-confidence listingCompany trend +1800%
Quick readStrong listing-quality and freshness signals

JLL empowers you to shape a brighter way . Our people at JLL are shaping the future of real estate for a better world by combining world class services, advisory and technology for our clients. We are committed to hiring the best, most talented people and empowering them to thrive, grow meaningful careers and to find a place where they belong. Whether you’ve got deep experience in commercial real estate, skilled trades or technology, or you’re looking to apply your relevant experience to a new industry, join our team as we help shape a brighter way forward. Facilities Technician – JLL What this job involves: We are seeking a proactive and skilled Facilities Technician to join our team in a role that is essential for the daily operation, upkeep, and long-term reliability of building systems. This position primarily supports vivarium operations, ensuring all equipment, utilities, and facility conditions meet strict compliance, performance, and environmental standards. While not stationed in the vivarium full-time this work will take place within laboratory, facility and vivarium environment, including preventative maintenance, corrective maintenance, environmental checks, and vivarium-specific work orders. This hands-on role requires strong technical ability, attention to detail, professionalism, and effective collaboration across Facilities and Research departments to maintain operational excellence and compliance. What your day-to-day will look like: Conduct regular vivarium walkthroughs to ensure cleanliness, operational readiness, and compliance with facility and environmental standards Complete all preventative and corrective maintenance for vivarium equipment, including HVAC components, air handling units, lighting systems, water supply systems, alarms, monitoring devices, and specialized mechanical systems Respond promptly to daily maintenance requests and vivari

Artificial IntelligenceAIRecruitmentCustomer Service
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s mission is to ensure that general-purpose artificial intelligence benefits all of humanity. We believe that achieving our goal requires effective engagement with public policy stakeholders and the broader community impacted by AI. Accordingly, our Global Affairs team builds authentic, collaborative relationships with public officials and the broader AI policymaking community to inform and support our shared work in these domains. We ensure that insights from policymakers inform our work and – in collaboration with our colleagues and external stakeholders – seek to shape policy so that it aligns with and supports our mission. About the role OpenAI is looking for an Applied Risk Standards Specialist to lead standards work for applied risk in applications, including mental health, youth safety, age assurance, functional efficacy, reliability, and privacy-related risks. You will help turn internal research, safety policies, and evaluation methods into technically sound, measurable, and adaptable standards that build public trust and support responsible deployment. This role sits within our AI standards and global assurance function in Global Affairs. Working closely with Safety Systems, Research, Product, Model Policy, Legal, and third-party evaluators, you will author proposals, negotiate requirements, and represent OpenAI in priority standards bodies. You will take on the drafting and external engagement needed to advance this work, enabling technical experts to focus on the underlying methods and evidence. You will also bring emerging requirements back into internal planning before they become established expectations. Alongside this primary portfolio, you will support the AI Standards and Global Assurance Lead on frontier AI assurance, contributing to risk-management standards, evaluation and benchmarking requirements, and approaches to independent assessment. The role connects technical practice with external standards. It does not replace t

Artificial IntelligenceAI
🔔

Get new reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime