Jobs in United States

Reliability Engineer Iii in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer iii jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The ChatGPT organization at OpenAI supports our mission by building products that bring cutting-edge AI capabilities to hundreds of millions of users worldwide. The Image Generation team is responsible for one of the fastest-growing experiences in ChatGPT, enabling users to create, edit, and transform images through natural language. Recent advances in our multimodal models have dramatically improved image quality, instruction following, editing precision, consistency, and text rendering, unlocking entirely new creative and professional workflows. We work at the intersection of research and product, partnering closely with researchers, designers, product managers, and platform engineers to bring state-of-the-art image generation capabilities to life across ChatGPT and our mobile applications. Millions of users rely on these experiences every day to create, communicate, learn, and build. About the Role We are seeking an experienced Android Software Engineer to build and improve image generation experiences within the ChatGPT Android app. You will help define how users create, edit, and interact with visual content powered by the latest multimodal AI models. This is an opportunity to work on a highly visible product area, translating cutting-edge AI capabilities into intuitive, performant, and delightful mobile experiences used by millions around the world. ChatGPT's Android app already enables users to generate and transform images directly from their devices, and we're just getting started. In this role, you will: Build and ship new Android features that power image generation and image editing experiences. Create intuitive user experiences that make advanced AI capabilities feel seamless and accessible. Collaborate closely with Product, Design, Research, and Engineering teams to bring new multimodal capabilities to production. Drive improvements in app performance, reliability, architecture, testing, and developer tooling. Optimize media-heavy workfl

AWSRestAIKotlin
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI develops models that can reason through complex problems and hardware designed for the demands of advanced AI. AI for Chips connects these efforts: applying increasingly capable AI systems to the work of semiconductor engineering. Our goal is to help engineers develop better chips and shorten design cycles. This work brings research, model training, and hardware expertise together to build tools that engineers can use on real designs, with correctness and measurable performance at the center. About the Role We’re hiring a Software Engineer to build the research infrastructure and tooling that help OpenAI models design silicon. You’ll turn chip-design workflows into reliable environments for reinforcement learning and evaluation, and make it easier for researchers to run experiments and iterate on new ideas. You’ll move between software engineering, tool integration, and open research problems. We value strong coding fundamentals, clear technical judgment, and independent execution. Prior chip-design experience is helpful, but you can learn the domain alongside the team’s hardware specialists. In this role, you will: Build and maintain infrastructure for reinforcement learning environments, evaluations, and long-running experiments. Integrate electronic design automation (EDA) tools into workflows for RTL generation, verification, and physical design optimization. Improve experiment reliability, reproducibility, observability, and performance; debug failures across tools, services, and infrastructure. Develop tooling and model harnesses that let researchers test ideas quickly and measure correctness and power, performance, and area (PPA). Collaborate with researchers and engineers to turn successful experiments into reusable systems and training workflows. Own ambiguous projects end to end, communicate progress, and use results to guide the next iteration. You might thrive in this role if you: Have strong software engineering fundamentals, with

PythonAWSRestAI
O
📍 Seattle, Washington, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The ChatGPT Library team is building the place where people can save, organize, rediscover, and build on the content they create with ChatGPT. Our goal is to make ChatGPT more useful over time by helping users seamlessly return to important files, images, conversations, and other content across their devices. The team works at the intersection of product engineering, design, and AI research to create intuitive, personalized experiences that make users’ content easy to find and act on. On Android, we are focused on delivering fast, reliable, and deeply native experiences that put a user’s evolving body of work at their fingertips. About the Role We are looking for a senior Android engineer to help build the future of ChatGPT Library on mobile. You will own high-impact product experiences across the Android stack, shaping how millions of people save, organize, discover, and interact with their content in ChatGPT. In this role, you will: Build and ship new Android experiences that help users easily access, organize, and build on the content they create with ChatGPT. Own features end to end—from early product exploration and technical design through implementation, experimentation, launch, and iteration. Develop scalable, maintainable foundations that allow Libraries experiences to evolve quickly as new AI capabilities emerge. Improve the architecture, performance, reliability, and responsiveness of content-rich experiences across a wide range of Android devices. Thoughtfully integrate Android platform capabilities to create experiences that feel intuitive and native to mobile. Partner closely with product, design, research, data science, and engineering teams to translate emerging AI capabilities into useful, polished products. Help establish technical direction and raise the quality bar for Android development across the team. How We Work We care deeply about building products that are intuitive, useful, and trustworthy. We move quickly from ideas to wo

RedisAWSRestAI
W
📍 New York City, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend +8.1%
Quick readStrong listing-quality and freshness signals

🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need. This is a hybrid position, based out of our New York City, San Francisco, Seattle, or London hubs. You'll report to our director of engineering. 🦸🏻‍♀️

PythonAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -80.2%

About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal and VMs dependable & scalable foundations for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—inc

AWSRestAIGo
S
📍 New York, NY, United States· Full-time
✓ High-confidence listingCompany trend -87.9%

$190.4K – $285.6K/yr

Quick readStrong listing-quality and freshness signals

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Lead the technical design and architecture of major platform initiatives, author design documents and build consensus across engineering teams. Define technical roadmaps for complex, multi-quarter projects that span multiple teams. Make critical architectural decisions for company documentation infrastructure, balancing scalability, reliability, and developer experience. Evaluate and set direction for integrating emerging technologies, including AI/LLM capabilities, into company documentation platforms and authoring tools. Establish and evolve engineering standards, best practices and technical guidelines for the team and broader organization. Partner with engineering teams across the company to understand documentation needs and design integrated solutions. Design, build and maintain scalable, reliable and performant services and systems. Contribute high-quality code across the full stack and navigate codebases with different languages and tools. Debug and resolve complex production issues and improve system reliability. Take ownership of system health and incident response. Who you are Minimum requirements Must have a Bachelor's degree or foreign equivalent in Computer Science, Software Engineering, Engineering, or a related field, plus four (4) years of experience in Software Engineering. Must have four (4) years of experience in each of the following: - Working in a full stack environment with a foc

TypeScriptJavaMongoDBAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Safety Systems org is responsible for various safety work to ensure our best models can be safely deployed to the real world to benefit the society and is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Safety Engineering team builds the platforms and tools that make OpenAI’s models safe to use in the real world. We partner closely with researchers, product teams, and policy to turn safety ideas into reliable, scalable systems: measuring risk, enforcing safeguards, and continuously improving how models behave in production. Our work sits at the intersection of product engineering, data, and AI, and directly shapes how millions of people experience OpenAI’s technology. About the Role We’re looking for a self-starter engineer who loves building products in an iterative, fast-moving environment—especially internal tools that unlock real-world impact. In this role, you’ll build full-stack tooling for our Safety Systems teams that directly improves the safety and reliability of OpenAI’s models, including in sensitive areas like mental health and other vulnerable-user protections. Your work will increase the team’s velocity in identifying and fixing safety issues and help tighten the feedback loop between policy, data, and the model training cycle. In this role, you will: Own the end-to-end development of internal tools that help improve the safety of OpenAI’s models (with a focus on areas like mental health and other vulnerable-user protections) Partner closely with Safety Systems researchers, engineers, and model policy creators to understand workflows, pain points, and requirements—and translate them into durable product solutions Build full-stack experiences to support core model policy workflows, such as labeling and inspecting data, analyzing and reviewing failure cases, and surfacing insights for iteration Optimize internal applications f

JavaScriptPythonJavaReact
O
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Astral team builds high-performance developer tools to power the future of programming, at OpenAI and beyond, including Ruff, uv, and ty. The Astral toolchain sees hundreds of millions of installs per month and powers hundreds of millions of package downloads per day for the Python ecosystem. As a team, we are building on those foundations to continue solving impactful tooling problems as programming evolves. About the Role We are looking for an experienced software engineer to build next-generation programming language tooling. If you like writing high-performance Rust, it could be a good fit; if you like thinking about the future of programming, it could also be a good fit. Strong candidates tend to have deep experience with Rust, Python, open source, compilers, or developer tools — but few candidates are deep in all of these areas, and we've hired candidates without prior Rust or Python experience. In this role, you will: Design and implement features in Astral’s existing open source projects (Ruff, uv, ty, and python-build-standalone, and more). Support Astral’s open source projects as a maintainer, triaging user issues, reviewing pull requests, and participating in community discussions. Evolve the Astral toolchain to accelerate development velocity at OpenAI. Build entirely new tools, in entirely different programming ecosystems, to power the future of agentic software development. Your background might look something like: 5+ years of professional engineering experience, excluding internships, in relevant engineering roles. High agency and comfort operating in a fast-moving environment, with strong ownership of security, reliability, and operational excellence. Strong developer empathy and communication skills, including experience maintaining open source projects. Exceptional systems engineering fundamentals and a track record of leading complex projects from ambiguous problem statements through to user impact. Proficiency in one or more s

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. In this role you will: As a Hardware Test Engineer, you will work on Machine Learning/AI hardware system projects to craft the solutions for current and future data center deployments. You will bring a strong understanding of hardware system testing, excellent project management skills, and the ability to collaborate across multiple teams to ensure efficient lab operations. You will be responsible for designing, implementing, and executing comprehensive test plans that ensure the reliability, performance, and scalability of our supercomputing hardware systems. You will develop detailed test plans and methodologies tailored to hardware components, including processors, memory modules, custom accelerators and interconnects. You will collaborate with hardware design, manufacturing, firmware teams and vendors to identify, analyze, and resolve issues affecting hardware, power, thermal and high-speed interconnects. You will perform in-depth debugging on the hardware system Excellent analytical skills to diagnose hardware issues, troubleshoot problems, and propose solutions. Ability to interpret complex test data, identify trends, and draw meaningful conclusions. High-speed links, with a focus on SerDes (Serializer/Deserializer) technology to assess signal integrity, error rates, and overall link performance. You will collaborate with the lab manager to maintain the equipment and hardware systems, including oscilloscopes, thermal test chambers, liquid cooling systems, and other mea

PythonAWSRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. About the role We are looking for an Operating Systems Engineer to build and harden the OS foundations for OpenAI products. We are especially interested in experienced, passionate, and innovative operating systems developers who thrive on building foundational platform software and solving hard problems in security, privacy, performance, power, and reliability. You will work across the OS kernel, core OS services, security and privacy primitives, performance and power, and the frameworks that connect applications and UI to the system. This role emphasizes deep debugging and systems ownership from development through production. You will collaborate closely with embedded, firmware, hardware, application, and product engineering teams. Experience with hardware bring-up is a plus, but not required. What you will do Work on end-to-end OS capabilities spanning the OS kernel, userspace services, application frameworks, UI toolkits, and application-facing APIs. Develop, integrate, and maintain OS components, both kernel-bound and in userspace, including scheduling, memory management, filesystems, drivers, IPC/RPC mechanisms, and security-relevant subsystems. Build and maintain core OS services and daemons (init, service management, device discovery, networking primitives, time, logging, update hooks, crash handling, and so on). Design and implement security and privacy mechanisms: Secure boot and measured boot integration points (where applicable). Mandatory access control and sandboxing. Secrets management, secure storage, key handling, and least-privilege service design. Privacy-preserving telemetry, data minimization, and user-consent oriented system behaviors. Establish a perfo

AWSLinuxRestAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI's People team helps hire, develop, and support the people building safe and beneficial AGI. Within that team, People Systems builds the technical foundation that enables our HR, recruiting, payroll, benefits, and performance operations to scale with quality, speed, and rigor. We work at the intersection of HR systems, software engineering, and internal tooling. Our goal is not just to keep core systems running, but to build durable technical leverage for the company. About the Role We're hiring a Workday Engineer to help design, build, and operate the systems that power critical people workflows at OpenAI. This is a highly technical role for someone who combines strong Workday expertise with real engineering fluency. You'll build reliable integrations, improve system architecture, automate complex workflows, and help connect Workday to internal tools, external platforms, and emerging AI-driven systems. You should be comfortable going beyond configuration work. We're looking for someone who can reason through ambiguous systems problems, write and debug technical solutions, work effectively in Git-based environments, and use modern developer workflows, including CLI-driven tooling, to build and operate with speed and discipline. You'll partner closely with cross-functional teams across People, Finance, Security, and Engineering, including our People Innovations team, to build systems that are secure, scalable, and practical. Some work will involve improving mature production infrastructure; some will involve building entirely new workflows and capabilities from scratch. In this role, you will Design, build, and maintain Workday integrations, applications, and workflow automations across domains such as payroll, benefits, recruiting, performance, and case management Improve the reliability, quality, and scalability of People systems through strong engineering, testing, and operational practices Build technical solutions that connect Workday with i

AWSGitRestAI
M
📍 O Fallon, Missouri, United States
✓ Quality checkedCompany trend +212.5%

Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Software Engineer II Overview Who is Mastercard? Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. With connections across more than 210 countries and territories, we are building a sustainable world that unlocks priceless possibilities for all. The Fraud Products team (part of O&T) is developing new capabilities for MasterCard's Decision Management Platform, which serves as the core for multiple business solutions to combat fraud and validate cardholder identity. Our patented Java-based platform processes billions of transactions per month in tens of milliseconds using a multi-tiered, message-oriented approach for high performance and availability. MasterCard software engineering teams leverage Agile development principles, advanced development, design and test automation practices, and an obsession over security, reliability, and perfo

JavaSQLJenkinsRecruitment
A
📍 San Francisco, CA, United States
✓ Quality checkedCompany trend +90.9%

Job Requisition ID # 26WD100611 Position Overview Autodesk is a global leader in design and make technology, with expertise across architecture, engineering, construction, design, manufacturing, and entertainment. Our software and services empower innovators everywhere to solve challenges big and small—from greener buildings to smarter products to more compelling media and entertainment. At Autodesk, we believe that when you have the right tools to work and think flexibly, you have the power to transform what actually needs making. We provide our customers with technology to help them achieve better outcomes for their products, businesses, and the world. We are looking for a highly motivated software engineer to join the PSET-Access group. The team builds and operates systems that enable secure, reliable, and scalable access to product capabilities and resources, and now in the process to develop the next generation to modernize the access management capabilities. In this role, you will work with engineers and cross-functional partners to design, build, test, deploy, and operate production software. You will primarily develop backend services using Java or Go, while owning well-defined components and projects, contributing to technical decisions, and helping improve the reliability, maintainability, and developer experience of the systems the team owns. This is a strong opportunity for an engineer who has solid software engineering fundamentals and is ready to grow their technical depth and ownership. Responsibilities Design, implement, test, deploy, and maintain production-quality backend software, prim

N
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -86%

$213K – $320K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: As a Developer Platform engineer, you will directly contribute to the foundational pieces that make Notion extensible and connected. You will build tools, APIs, and platform experiences that help customers connect Notion to the world — bringing their apps, data, and workflows into one workspace. Your work will make it easier for developers, admins, and builders to create reliable integrations and automations on top of Notion. You will be a key player in building the robust technical foundation that allows Notion to achieve the connected workspace vision. Your work will include both internal platform contributions that accelerate other Notion engineering teams and end-user-facing functionality that enables toolmaking ubiquity. You will be presented with challenging technical problems, as Notion’s product needs are complex. You’ll play a key role in identifying and executing against technical investments that ensure the long-term quality, reliability, and performance of Notion’s platform as we scale. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursd

TypeScriptRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Infrastructure Quality team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, general contractors, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. Our work spans from vendor qualification through commissioning, ensuring operational readiness across our global portfolio. About the Role We are seeking an experienced Manufacturing Quality Engineer (MQE) to establish, implement, and manage a manufacturing-focused quality program for datacenter infrastructure. This role will be responsible for vendor oversight, quality assurance, process improvement, and issue resolution for all critical systems. You will lead vendor audits, monitor performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced risks, and operational reliability. By partnering with vendors, construction teams, and internal stakeholders, you will help ensure OpenAI’s datacenters are delivered on time and built to the highest operational standards. Travel Domestic and international travel as needed (estimated 40–60%) to manufacturing sites, datacenter locations, and partner facilities. Key Responsibilities Vendor Oversight & Performance Management Conduct manufacturing evaluation, audits, and improve vendor performance across production, inspection, testing, and delivery phases. Develop and track quality metrics to assess manufacturing performance and identify trends. Partner with vendors to refine processes, training, and quality controls to mitigate risks before shipment. Program Development & Execution Develop and maintain a datacenter-focused m

AWSRestAIGo
🔔

Get new reliability engineer iii jobs in United States by email

Daily job updates · Unsubscribe anytime