Jobs in United States

Reliability Engineer in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's mission is to ensure that AGI benefits all of humanity. The Business Systems team helps make that mission possible by building the internal products and platforms that allow OpenAI to operate with speed, reliability, and care. We build internal applications and workflows for Finance and Supply Chain. Our work spans product discovery, React and TypeScript interfaces, Python services and APIs, data models, workflow orchestration, enterprise integrations, and the systems that connect people to systems of record. We work directly with the people who use these products and care about correctness, permissions, auditability, and production reliability. Examples of our work include building an integration platform for supply chain integrations, integrations with Oracle Fusion and Zip, contract intelligence applied to B2B revenue recognition, and Temporal-based agentic workflows for credit checks, duplicate bank detection, and invoice triaging. We turn these efforts into reusable patterns that can support many workflows, rather than one-off automations. About the Role We are looking for Product Engineers to build internal applications end to end. This role spans product discovery, user experience, frontend, backend services, data models, workflow orchestration, and integrations with order management, fulfillment, and supply chain systems. You will take a problem from a first conversation with a Finance or Supply Chain partner through design, implementation, rollout, and production support. Strong candidates combine product judgment with engineering depth. You should be comfortable moving between a React interface, a Python API, a durable workflow, and an integration with an enterprise system. You should be able to ship a useful first version quickly while building the foundations for reuse, security, and long-term maintainability. Direct AI experience is helpful, but the core requirement is strong product engineering judgment and reliable execution. I

TypeScriptPythonReactSQL
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

PythonAWSAzureGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Cloud Agents team builds product infrastructure for long-running agents in the cloud: orchestration, sandboxing and isolation, secure environment connectivity, secrets and identity, observability, reliability, and cost controls. These agents securely connect to diverse developer and customer environments and use tools to accomplish goals. We partner closely with product, research, and infrastructure teams to turn agentic capabilities into dependable platforms for OpenAI products and developers building on OpenAI. About the Role We are looking for an experienced software engineer to help build and scale our cloud agent platform. You will design and operate systems for orchestrating agents at scale. You will work closely with product engineers on ChatGPT, API, and Codex to define the right abstractions and enable them to ship products quickly. Strong backend or infrastructure experience is important; experience with Python, Rust, distributed systems, cloud infrastructure, or product platforms is especially helpful. In this role, you will: Design and scale the orchestration, sandboxing and storage systems that run agentic workloads for Codex, ChatGPT, and the OpenAI API. Partner with product engineers to build a platform that enables them to ship quickly and turn feedback into robust abstractions. Improve reliability, security, performance, and cost efficiency for long-running agents. Deploy services that can operate across different environments and clouds. Your background might look something like: 9+ years of professional engineering experience, excluding internships, in relevant roles at technology and product-driven companies. Experience leading large-scale backend, platform, or infrastructure projects from ambiguous problem statements to production systems. Proficiency in one or more backend languages such as Python, Go, Rust, TypeScript, or similar, and the ability to move across service, platform, and product boundaries. Strong understanding

TypeScriptPythonAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team: The Database Systems team specializes in high-performance distributed databases. Our team built Rockset, the real-time search, analytics, and vector database that powers all vector search and retrieval augmented generation (RAG) at OpenAI. In addition to retrieval, as an online database, Rockset powers core functionality across all of OpenAI's product lines and many critical internal use cases. About the Role : We are looking for engineers passionate about distributed systems, close-to-the-metal performance optimization (our core engine is written in C++), and building scalable database infrastructure from the ground up. As an engineer on the Database Systems team, you'll contribute to the core database engine, driving improvements across ingestion, query execution, indexing, and storage. You'll partner with teams across OpenAI to unlock new product capabilities and help scale online database reliability and throughput as usage grows by orders of magnitude. In this role you will: Design, build, and operate high-performance distributed systems Identify and resolve performance bottlenecks to scale infrastructure to the next order of magnitude Define long-term technical direction and guide system evolution Collaborate with product, engineering, and research teams to deliver scalable and reliable infrastructure Dig deep into complex production issues across the stack Contribute to incident response, postmortems, and best practices for system reliability You might thrive in this role if you: Have significant experience building, scaling, and optimizing distributed systems at scale Are curious about database internals, storage engines, or low-latency query systems Enjoy debugging challenging performance issues in complex, high-throughput systems Have experience operating production clusters at scale (e.g., Kubernetes or other orchestration systems) Think rigorously about scalability, correctness, and reliability Thrive in fast-paced environments with high

AWSAzureGCPKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Safety Systems team is dedicated to ensuring the safety, robustness, and reliability of AI models and their deployment in the real world. Learn more about OpenAI’s approach to safety. Building on the many years of our practical alignment work and applied safety efforts, Safety Systems addresses emerging safety issues and develops new fundamental solutions to enable the safe deployment of our most advanced models and future AGI, to make AI that is beneficial and trustworthy. About the Role At OpenAI, we're dedicated to advancing artificial intelligence, and we know that creating a secure and reliable platform is vital to our mission. That's why we're seeking a software engineer to help us build out our trust and safety capabilities. In this role, you'll work with our entire engineering team to design and implement systems that detect and prevent abuse, promote user safety, and reduce risk across our platform. You'll be at the forefront of our efforts to ensure that the immense potential of AI is harnessed in a responsible and sustainable manner. Your Responsibilities: Architect, build, and maintain anti-abuse and content moderation infrastructure designed to protect us and end users from unwanted behavior. Work closely with our other engineers and researchers to utilize both industry standard and novel AI techniques to measure, monitor and improve AI models’ alignment to human values. . Diagnose and remediate active incidents on the platform and build new tooling and infrastructure that address the root causes of system failure. You might thrive in this role if: You have built and run production services in a high growth, rapidly scaling environment. You can debug live issues and restore systems quickly. You have worked on content safety, fraud, or abuse, or are motivated and excited to work on present-day (“now-term”) AI safety. You have experience with Python or with modern languages such as C++, Rust, or Go, and are able to quickly ramp up on Py

PythonAWSAzureKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Safety Systems team is dedicated to ensuring the safety, robustness, and reliability of AI models and their deployment in the real world. Building on the many years of our practical alignment work and applied safety efforts, Safety Systems addresses emerging safety issues and develops new fundamental solutions to enable the safe deployment of our most advanced models and future AGI, to make AI that is beneficial and trustworthy. Learn more about OpenAI’s approach to safety About the Role As an Analytics Engineer in Safety Systems, you will play a pivotal role in building a data-centric culture, enhancing decision-making processes, and driving strategic initiatives through analytics. You will partner closely with Engineering, Research, and Data Science to develop and maintain canonical data sources and source-of-truth dashboards that enable both people and AI agents across the organization to derive trustworthy, actionable insights. You will own the consumption layer for safety metrics: defining intuitive, reliable ways for stakeholders across Safety Systems, partner teams, and leadership to understand the safety of our products, answer safety-related questions independently, and inform product decisions and company strategy. Most importantly, you will be a core member of the Safety Systems team, collaborating with researchers and engineers to advance our goals of safe, robust, and reliable AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and maintain canonical datasets that serve as sources of truth for safety metrics. Develop and refine data products such as dashboards, reports, agent-enabled workflows, and machine-readable interfaces that empower stakeholders to extract and analyze data independently. Work closely with stakeholders in Engineering, Research, and Data Science to understand their decision-making n

SQLAWSRestAI
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Chemical and Slurry Systems Engineer within Micron’s Global Facilities and Construction Team, you will deliver engineering expertise supporting the planning, design, construction, operation, and maintenance of facilities systems across Micron’s global manufacturing network. You will participate in capacity and scenario planning, evaluate impact to facilities infrastructure, and serve as a technical domain expert throughout all project phases. Your work ensures environmental, safety, regulatory, and code compliance while driving system reliability and operational excellence. Responsibilities Develop safe, reliable designs for Chemical and Slurry systems and ensure compatibility of materials of construction with all chemicals used. Create, review, and maintain corporate equipment specifications, engineering standards, and design guidelines. Advise sites on system capacity, load projections, and scenario‑driven impacts. Provide technical feedback on infrastructure additions or modifications to support manufacturing and technology roadmap changes. Supply Micron standards to design partners and conduct timely reviews of design packages, construction documents, and engineering work. Support project cost and schedule development and validate scope alignment with collaborator requirements. Coordinate with Facilities, Design, Construction, Procurement, and Manufacturing to ensure clear communication and project alignment. Provide technical guidance during construction,

AIProcurementRecruitment
H
📍 Colorado, United States of America, United States
✓ High-confidence listingCompany trend +103.7%

$59.4K – $89.6K/yr

Quick readStrong listing-quality and freshness signals

Software Quality Engineer Description - This role is responsible for maintaining the quality, reliability, and performance of software applications throughout the development lifecycle. The role identifies and rectifies defects, ensures adherence to established quality standards, and contributes to the overall improvement of the software development process. The role involves various activities aimed at preventing and detecting issues, thereby enhancing the end user experience. The role creates and executes comprehensive test plans, test cases, and test scripts based on project specifications. *Onsite in Ft. Collins 5-days a week Responsibilities • Executes established test plans and protocols for assigned portions of code for end-user applications, systems software, and firmware running on hardware, local, networked, and Internet- based platforms; identifies, logs, and debugs assigned issues. • Perform Functional and Solution Testing of Video/Collaboration Software • Additionally, codes and programs test scripts, automation, and integration activities based on specific test requirements. • Conducts functional, integration, regression, and performance testing to validate software functionality. • Automates testing processes using appropriate tools and frameworks to improve efficiency and repeatability. • Monitors and enforces adherence to established coding standards, design guidelines, and best practices. • Monitors software performance and conducts load and stress testing to identify bottlenecks and performance issues. • Prepares and maintains QA-related documentation, including test plans, test matrices, and testing reports. • Develops understanding of and relationship with internal and outsourced development partners on software applications design and development. • Participates as a member of project

PythonAIJenkins
H
📍 Colorado, United States of America, United States
✓ High-confidence listingCompany trend +103.7%

$59.4K – $89.6K/yr

Quick readStrong listing-quality and freshness signals

Software Quality Engineer Description - This role is responsible for maintaining the quality, reliability, and performance of software applications throughout the development lifecycle. The role identifies and rectifies defects, ensures adherence to established quality standards, and contributes to the overall improvement of the software development process. The role involves various activities aimed at preventing and detecting issues, thereby enhancing the end user experience. The role creates and executes comprehensive test plans, test cases, and test scripts based on project specifications. *Onside in Ft Collins 5-days a week Responsibilities • Executes established test plans and protocols for assigned portions of code for end-user applications, systems software, and firmware running on hardware, local, networked, and Internet- based platforms; identifies, logs, and debugs assigned issues. • Perform Functional and Solution Testing of Video/Collaboration Software • Additionally, codes and programs test scripts, automation, and integration activities based on specific test requirements. • Conducts functional, integration, regression, and performance testing to validate software functionality. • Automates testing processes using appropriate tools and frameworks to improve efficiency and repeatability. • Monitors and enforces adherence to established coding standards, design guidelines, and best practices. • Monitors software performance and conducts load and stress testing to identify bottlenecks and performance issues. • Prepares and maintains QA-related documentation, including test plans, test matrices, and testing reports. • Develops understanding of and relationship with internal and outsourced development partners on software applications design and development. • Participates as a member of project t

PythonAIJenkins
O
📍 Atlanta, Georgia, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge As a Senior Staff Software Engineer, you will serve as a technical leader for OneTrust’s AI Governance (AIG) platform, driving the design, scalability, and reliability of systems that enable enterprises to deploy and govern AI and LLM-powered applications responsibly. You will deeply understand how customers build, deploy, and operate AI systems, and translate those needs into secure, compliant, and observable platform capabilities. Your Mission Development Lead the design and development of Java/Python microservices and shared libraries integrating with AI platforms for OneTrust’s AI Governance product. Design, build, and test cloud-native applications deployed on Microsoft Azure using Core Java, REST, and the Spring ecosystem. Lead the architecture and development of reusable AIG reporting and dashboard capabilities that integrate governance data from SQL databases and analytical platforms with runtime observability signals. Design reusable semantic-layer and metric-abstraction capabilities, including dataset contracts, metric defini

PythonJavaSQLAWS
O
📍 San Francisco, California, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge As a Senior Staff Software Engineer, you will serve as a technical leader for OneTrust’s AI Governance (AIG) platform, driving the design, scalability, and reliability of systems that enable enterprises to deploy and govern AI and LLM-powered applications responsibly. You will deeply understand how customers build, deploy, and operate AI systems, and translate those needs into secure, compliant, and observable platform capabilities. Your Mission Development Lead the design and development of Java/Python microservices and shared libraries integrating with AI platforms for OneTrust’s AI Governance product. Design, build, and test cloud-native applications deployed on Microsoft Azure using Core Java, REST, and the Spring ecosystem. Lead the architecture and development of reusable AIG reporting and dashboard capabilities that integrate governance data from SQL databases and analytical platforms with runtime observability signals. Design reusable semantic-layer and metric-abstraction capabilities, including dataset contracts, metric defini

PythonJavaSQLAWS
JI
📍 Kenosha, WI, United States
✓ High-confidence listingCompany trend +1800%
Quick readStrong listing-quality and freshness signals

JLL empowers you to shape a brighter way . Our people at JLL are shaping the future of real estate for a better world by combining world class services, advisory and technology for our clients. We are committed to hiring the best, most talented people and empowering them to thrive, grow meaningful careers and to find a place where they belong. Whether you’ve got deep experience in commercial real estate, skilled trades or technology, or you’re looking to apply your relevant experience to a new industry, join our team as we help shape a brighter way forward. Automation Engineer – JLL What this job involves: We are seeking an experienced Automation Engineer to design, develop, and implement automation control systems for industrial processes and warehouse distribution equipment. The role requires strong knowledge of engineering principles, programming, and control system technologies, with a focus on improving the reliability and performance of conveyors, sortation systems, scanners, cameras, print-and-apply systems, and SCADA devices. All work must follow established policies and procedures, with safety as a top priority. What your day-to-day will look like: Serve as site technical expert in automation control systems and mentor Apprentices to meet safety and technical standards. Design, develop, implement, and optimize control systems and software; maintain and troubleshoot equipment including PLC/PC controllers and industrial networks. </

Artificial IntelligenceAIWarehouseRecruitment
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Department Intro: The Advanced Packaging Technology Development (APTD) department at Micron Technology is shaping the future of memory and storage technology through advanced packaging innovation. Our team develops the processes, automation, and manufacturing capabilities that enable next-generation semiconductor products. Working closely with global R&D and manufacturing partners, we turn new ideas into scalable solutions that help maintain Micron’s leadership in the industry. Position Overview: As a Process Engineer, you will play a key role in improving the quality, reliability, and manufacturability of future memory products. This position combines hands-on problem solving, process optimization, and cross-functional collaboration in a fast-paced development environment. You will help drive technical advancements, influence technology roadmaps, and contribute directly to the success of next-generation semiconductor solutions. Responsibilities: Optimize automation, equipment, and process performance to improve product quality, reliability, and cost Develop and enhance host, automation, and data systems to accelerate learning and decision-making Perform root cause analysis and failure investigations using methodologies such as FMEA, 8D, SPC, and FDC Partner with equipment, process development, and process integration teams to develop and implement new solutions Support pilot line operations and technology transfer to high-volume ma

Artificial IntelligenceAIRecruitment
D
📍 New York, New York, United States
✓ High-confidence listingCompany trend -84.7%
Quick readStrong listing-quality and freshness signals

Datadog's Software Engineers with Systems depth leverage their experience with systems and tooling to build software that ensures Datadog remains reliable, performant, and secure. For this track, their Software Engineering experience may resemble the Distributed Systems track, but is typically applied in combination with their systems experience to build and run internal platforms and tools that our products are built on. These people typically have deep experience building and managing large cloud infrastructure deployments, or leading reliability efforts for orgs similar to ours, or building release machinery to allow hundreds or thousands of devs to do their jobs without stepping on each others' toes. The systems and tooling where they may have experience depth may include (but not limited to): bazel, build tooling, cassandra, CDN, chef, configuration management, container orchestration, consul, docker, elasticsearch envoy, haproxy, kafka, kubernetes, load balancing, network architecture, postgres, redis, release management, RPC frameworks, service discovery, spinnaker, terraform, zookeeper. Bonus: You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. #LI-KM5 Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan. Th

PostgreSQLRedisDockerKubernetes
MT
📍 Richardson, TX, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. The HBM Design Technology Package Co-Optimization (DTPCO) organization is seeking a Staff Engineer! Lead the design and development of test vehicles that enable technology learning and risk reduction for future HBM products! This role serves as a key technical contributor within DTPCO, partnering closely with Architecture, Design, Layout, Product Engineering, Reliability, and Advanced Packaging teams to translate emerging product requirements and technology challenges into actionable test vehicle solutions. You will apply expertise in semiconductor design, physical implementation, silicon characterization, and advanced packaging technologies to define test structures, develop validation strategies, and generate data that drives product decisions across future HBM generations. Unlike a traditional product role centered on ownership of product blocks, this position emphasizes enabling learning in development and engineering. It does so by developing representative test vehicles and characterization structures. Key Responsibilities Serve as a technical lead for HBM test vehicle design and structure development within DTPCO. Partner with HBM product, technology and reliability team. Define and implement test structures targeting key learning areas such as package interaction and reliability Drive test vehicle content definition, structure specifications, measurement requirements, and validation objectives. Analyze silicon characterization and qualification results to identify optimization opportunit

AIRecruitment
🔔

Get new reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime