About the Team The ChatGPT organization at OpenAI supports our mission by building products that bring cutting-edge AI capabilities to hundreds of millions of users worldwide. The Image Generation team is responsible for one of the fastest-growing experiences in ChatGPT, enabling users to create, edit, and transform images through natural language. Recent advances in our multimodal models have dramatically improved image quality, instruction following, editing precision, consistency, and text rendering, unlocking entirely new creative and professional workflows. We work at the intersection of research and product, partnering closely with researchers, designers, product managers, and platform engineers to bring state-of-the-art image generation capabilities to life across ChatGPT and our mobile applications. Millions of users rely on these experiences every day to create, communicate, learn, and build. About the Role We are seeking an experienced Android Software Engineer to build and improve image generation experiences within the ChatGPT Android app. You will help define how users create, edit, and interact with visual content powered by the latest multimodal AI models. This is an opportunity to work on a highly visible product area, translating cutting-edge AI capabilities into intuitive, performant, and delightful mobile experiences used by millions around the world. ChatGPT's Android app already enables users to generate and transform images directly from their devices, and we're just getting started. In this role, you will: Build and ship new Android features that power image generation and image editing experiences. Create intuitive user experiences that make advanced AI capabilities feel seamless and accessible. Collaborate closely with Product, Design, Research, and Engineering teams to bring new multimodal capabilities to production. Drive improvements in app performance, reliability, architecture, testing, and developer tooling. Optimize media-heavy workfl
Jobs in United States
Software Reliability Engineer in United States
2,007 active opportunities · Updated October 2026
Showing
15 jobs
Explore current software reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent Senior Engineering Manager to lead a large firmware engineering organization delivering end-to-end manageability firmware for NVIDIA's next generation Data Center Compute Systems. This role owns HGX product line and OpenBMC-based management firmware and MCU firmware components in data center platforms, including architecture, execution, quality, reliability, telemetry, and customer readiness. We are seeking an experienced senior leader with strong technical depth, broad system perspective, and a proven ability to lead large teams through complex product cycles. This role is onsite in Santa Clara, CA, USA. If you're creative and autonomous, we want to hear from you! What you'll be doing: Lead a large firmware engineering organization delivering OpenBMC based firmware and MCU firmware for next-generation Data Center Compute Systems. Own HGX platform as a lead for Firmware and System software readiness working across the organization. Define and drive the long-term firmware roadmap, balancing architectural innovation with product execution and delivery milestones. Drive architecture strategy across BMC, MCU, platform software, manageability, health management, and data center firmware interfaces. <spa
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 We are seeking a highly technical and experienced Engineering Manager to one of our business critical engineering team. This role is ideal for a hands-on leader who thrives in a fast-paced environment, has a deep understanding of collaborative document editing technologies, and is passionate about building scalable, high-performance software. As the Manager, you will oversee the development and delivery of innovative features, ensure technical excellence, and mentor a team of talented engineers. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our product. Operational Excellence : Ensure the stability and performance of the Docs platform, proactively addressing technical debt and optimizing system architecture. Qualifications: Technical Expertise : Proficiency in
From $88.1K/yr
About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Platform Engineer, you will contribute to building and improving Stitch Fix’s cloud-native infrastructure and internal developer tooling. You’ll work on tools and automation that help product engineers deploy, operate, and debug services more easily, while learning modern platform engineering practices alongside experienced teammates. This role is ideal for engineers who enjoy improving developer experience and want to grow their skills in cloud infrastructure and CI/CD systems. Responsibilities: Contribute to the development and evolution of our internal platform-as-a-service used by application and service developers Build and maintain tooling that improves developer workflows, deployment reliability, and day-to-day productivity Collaborate with platform and application engineers to identify friction points and implement incremental improvements Learn and apply best practices around Infrastructure-as-Code, containerized workloads, and CI/CD pipelines Use, or are eager to adopt, AI-assisted development tools to improve productivity, and are excited to help explore and integrate LLM-powered solutions that automate internal support and operational workflows Have opportunities to propose ideas and improvements, with support and mentorship from the team Things you’ll get exposure to (and we don’t expect experience with everything): AWS Terraform, Pulumi CircleCI Docker, ECS, EKS Ruby, Golang, Python About You 2+ years of software development and infrastructure experience with significant contribut
From $137K/yr
The Code Gen team is tasked with building AI-powered code transformation tools that transform rigid, legacy applications that suffer from poor scalability and high operating costs into modern, microservices-based architectures that are built on top of MongoDB. Join our team and be at the forefront of innovation and creativity. We are looking for a Staff Engineer with domain expertise and years of experience in modernizing legacy applications that are based on traditional database systems. A significant advantage is profound prior experience in leveraging AI, particularly LLMs and GenAI capabilities, to enable reliable, self-driving automation of the code transformation, iterative build, and test processes. In this role, you will be instrumental in initiating technical strategies and ideas, lead the Code Gen team in designing, building, and optimizing our code transformation workflow and tools. You will work on critical components that ensure the scalability, efficiency, and reliability of our services. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation, build and test. This role will be based remotely in North America. A strong candidate for this position will have Extensive experience (8+ years) in software development and operations, with a proven track record of delivering high performance, correctness, and architectural excellence in fast-paced environments Experience using Relational Databases such as Oracle, MySQL, Microsoft SQL Server or PostgreSQL Experience with tools and methodologies for code analysis, refactoring, and automated testing Experience in designing and implementing complex software systems, collaborating effectively with engineers of all experience levels to achieve high reliability and performance Practical knowledge of integrating GenAI into large-scale, complex systems, including a clear unde
About the Team The SaaS and Software Governance team sits within Corporate IT and helps OpenAI scale from startup-speed tooling to mature enterprise architecture. The team owns practical governance for software, SaaS, integrations, APIs, identity, data access, and agent-enabled workflows, with a mandate to improve security, reduce software sprawl, and help business teams move faster through better foundations. About the Team OpenAI is scaling from an emerging, high-velocity startup into a mature enterprise operating model. The IT Software Architect will help shape the software, SaaS, Data integration, and agent-enabled architecture that lets the company move quickly while improving security, compliance, data quality, and customer, partner, and employee experience. This role is not a traditional ivory-tower architecture function. It is a hands-on governance and enablement role that partners with business teams, IT operations, Security, Procurement, Business Platforms, Applied teams and Data Engineering to guide software decisions, reduce unmanaged sprawl, and build reusable enterprise foundations. Why This Role Matters OpenAI’s software footprint is expanding rapidly across SaaS, internally built tools, agents, integrations, APIs, third-party platforms, and application systems that OpenAI. The company needs a stronger tools architecture layer that can help teams make good decisions early, avoid duplicate tools, govern sensitive data and identities, and identify where OpenAI should build instead of buy. The person in this role will help turn software governance from an approval checkpoint into an enterprise capability: a system that improves speed, reliability, security, and business outcomes. What You'll Do Own the target architecture for enterprise software, SaaS, integrations, APIs, and agent-enabled business systems across Corporate IT. Drive deprecation and consolidate targets for enterprise software. Build lightweight governance patterns that guide teams before
About the Team OpenAI’s Applications Engineering organization builds and operates the products that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role You will be a Senior Android engineer on OpenAI’s Applied Foundations team, building the core mobile experiences that power how users sign up, manage their account, family features, pay for services, stay safe, and interact with OpenAI’s products with confidence. This role is about creating high-quality products as well as reusable Android foundations that product teams across different OpenAI apps depend on to ship quickly while meeting the highest standards for security, reliability, and user trust. You’ll own complex client-side systems spanning UI, networking, local state, payment integrations and Apple platform integrations, and work closely with backend, product, and safety partners to shape the architecture that supports OpenAI’s mobile ecosystem at global scale. You might thrive in this role if you: Have 4+ years of professional software engineering experience. Have a proven track record of building high-quality Android applications in production. Are fluent in Kotlin (and/or Java) and familiar with Android development tools and architecture components. Prioritize performance, security, and user experience in mobile development. Enjoy working cross-functionally to bring ambitious product ideas to life. Care deeply about performance, security, and user experience. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We
About the Team OpenAI’s Applications Engineering organization builds and operates the products that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role You will be a Senior iOS engineer on OpenAI’s Applied Foundations team, building the core mobile experiences that power how users sign up, manage their account, family features, pay for services, stay safe, and interact with OpenAI’s products with confidence. This role is about creating high-quality products as well as reusable iOS foundations that product teams across different OpenAI apps depend on to ship quickly while meeting the highest standards for security, reliability, and user trust. You’ll own complex client-side systems spanning UI, networking, local state, payment integrations and Apple platform integrations, and work closely with backend, product, and safety partners to shape the architecture that supports OpenAI’s mobile ecosystem at global scale. In this role, you will: Build and ship new experiences on iOS that showcase the power of AI. Optimize app performance, reliability, and responsiveness at global scale. Design and maintain shared iOS frameworks and primitives for account, trust, and commerce flows that are used across OpenAI’s mobile apps. Establish robust testing frameworks and refine app architecture for long-term maintainability. Collaborate with product, design, research, and backend teams to deliver high-impact features. Provide technical leadership to shape the future of OpenAI’s iOS platform. You might thrive in this role if you: Have 4+ years of professional software engineering experience. Hav
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We’re hiring a talented Software Engineering Manager to lead the Snowtrail infrastructure team at Snowflake. Snowtrail is the infrastructure that enables Snowflake to deliver dedicated coverage for customer-specific workloads. Its innovative approach allows Snowflake to precisely test and measure the impact of changes on individual customers, making it essential for ensuring the platform’s reliability, correctness, and performance. Through query replay, Snowtrail helps us catch regressions early. By leveraging machine learning models to intelligently sample queries and workloads, we continuously optimize for both cost and performance. Evolving Snowtrail to incorporate new engine features while improving scalability, efficiency, and reliability is central to our continued success OUR IDEAL MANAGER WILL HAVE : Strong passion and proven track record for shipping quality software in high code velocity environments 10+ years industry experience designing and building distributed data systems. Excellent problem solving skills, and strong CS fundamentals including data structures, algorithms, and distributed systems. Fluency in SQL, Java, C++, Python or Go. Ability to collaborate well across teams, build high-performing teams and mentor junior engineers. Excellent interpersonal co
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 ClickUp is looking for an experienced Engineering Manager to lead our fullstack team responsible for building and scaling our flagship products. As the leader of the team that owns the APIs and core experiences powering ClickUp, you will play a pivotal role in shaping the future of our platform. You will guide engineers working across the stack, from frontend experiences to backend infrastructure. Your focus will be on driving the development of new features, addressing performance and reliability challenges, and ensuring operational excellence as we continue to grow. This is an opportunity to make a significant impact on our core product while fostering a culture of technical excellence and collaboration. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our produ
About the Team The Finance & Supply Chain Engineering organization includes two complementary teams. Software Engineering builds internal full-stack applications, durable agentic workflows, plugins, MCPs, and measurable AI-enabled engineering practices. Data Engineering builds trusted analytics data assets for Finance and Supply Chain. The teams have distinct charters, with important shared dependencies and broad cross team partnerships across Engineering, Applications, Finance, and Supply Chain. About the Role We are looking for a hands-on senior technical leader who will report alongside the Software Engineering and Data Engineering managers. This is an individual-contributor role with no immediate people-management responsibility. The Tech Lead will raise the technical bar across both teams, participate in important cross-team or high-risk design decisions, and directly own and ship high-impact work. The role should improve team judgment and autonomy rather than act as a floating architect or universal approval gate. In this role, you will: Partner with the Software Engineering and Data Engineering managers as a peer technical leader; managers retain accountability for people, staffing, priorities, performance, and delivery commitments. Directly own the architecture, implementation, launch, and operation of one or more high-impact initiatives, remaining accountable for real outcomes rather than advisory output alone. Guide important design decisions that are cross-team, difficult to reverse, or material to security, financial controls, reliability, data quality, or long-term cost of ownership. Establish pragmatic engineering standards across architecture, APIs and data contracts, testing, security, reliability, observability, lineage, data quality, and operational ownership. Advance engineering standards for building with AI, including agentic workflows, evaluation, telemetry, adoption, and outcome measurement. Work across backend, full-stack, data, and big-d
NVIDIA has been redefining computer graphics, desktop gaming, and enhanced computing capabilities for more than 25 years. Today, we are tapping into the unlimited potential of AI to define the next era of computing. As a NVIDIAN, you will work on problems that sit at the boundary of architecture, silicon, firmware, software, and production, where strong judgment matters as much as technical depth. We're the Silicon Design for Productization (DFP) Team, within the broader Silicon Co-Design Group, and we turn power and thermal design into executable productization methodology. Power and thermal are among the most complicated problems we work on at NVIDIA because they sit at the intersection of architecture, workload behavior, silicon variation, firmware policy, platform constraints, and product goals. Small decisions here have an outsized impact on performance, efficiency, reliability, bring-up speed, and ultimately what the product can deliver in the field. We define how features move from concepts to bring-up, characterization, validation, and release. In this role, you will help us build that bridge. We're looking for an engineer who reasons from first principles, flourishes with ownership in a fast-paced environment, and uses AI with sound judgment. What you’ll be doing: Lead the effort across multi-functional teams to keep the program’s power and thermal productization strategy clear, executable, and on track. Create methodology and silicon test plan based controller designs and architecture, including characterization process, debug tools, fuse/firmware settings and lab requirements. Drive resolution for challenging silicon issues through structured hypotheses, measurement plans, and root-cause closure. Steward the Power and Thermal playbook when the existing productization methodology
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team builds next-generation AI-native silicon and systems while working closely with software, research, and manufacturing partners to co-design hardware tightly integrated with AI models. In addition to delivering systems for OpenAI’s supercomputing infrastructure, the team develops the tools, methodologies, and strategic partnerships needed to accelerate hardware innovation. About the Role We’re seeking an experienced Hardware Strategic Sourcing Manager to own sourcing strategy and supplier partnerships for fiber and optical interconnect components across OpenAI’s next-generation AI infrastructure. Reporting to the Head of Partnerships & Strategic Sourcing, you will lead sourcing across fiber cable assemblies, internal optical harnesses, fiber shuffles, optical backplane assemblies, connectorized and standalone passive optical assemblies, fiber-array units (FAUs), fiber-to-chip and coupling interfaces, detachable connectors, optical routing, and assigned optical packaging, assembly, and test services. You will work closely with electrical engineering, optical engineering, systems engineering, mechanical and packaging engineering, quality, rack integration, data-center deployment,manufacturing, supply chain, finance, legal, and program management teams to translate demanding bandwidth, signal integrity, reliability, and scale requirements into resilient supplier partnerships and scalable commercial strategies. Your work will directly support the performance, reliability, manufacturability, and scale of the high-speed optical connectivity required for OpenAI’s next-generation AI systems. In this role, you will: Develop and execute a comprehensive sourcing strategy for fiber and optical interconnect components supporting high-bandwidth AI systems and infrastructure. Own sourcing across optical fiber cable assembli
About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. The Systems Integration team is critical in this mission, turning complex hardware-software development into reliable product signals. We build the shared infrastructure, tooling, and lab environments that let teams test quickly, understand failures, and ship with confidence. About the Role As a Systems Integration Manager , you will lead the team responsible for device validation infrastructure, test automation, developer tooling, and lab operations. This is a player-coach leadership role: you’ll set technical and operational direction, build and develop a team of engineers and lab operations professionals, and stay close to the architecture and hardest systems problems. You will partner closely with device software, OS, firmware, hardware, reliability, QA, and release infrastructure teams to define validation strategy, improve release readiness, and ensure our test environments and quality signals scale with the product. Because this is a new category of devices, you’ll have the rare opportunity to build the validation foundation early—shaping the systems, standards, and operating model that will support products from prototype through launch. We’re looking for a leader who combines strong technical judgment with people leadership, operational rigor, and experience building reliable systems for complex hardware-software products. This role is b
About the Team OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions. Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency. About the Role We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments. This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment. The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries. Key Responsibilities Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure. Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios. Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize
Other cities to consider
More places hiring for this role
Get new software reliability engineer jobs in United States by email
Daily job updates · Unsubscribe anytime