Jobiba hiring network

Back End Td Reliability Lab Manager Jobs

1,726 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current back end td reliability lab manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.

S
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Issuing team owns the Stripe card issuing product end to end from the APIs and backend services that power spend cards, debit cards, and charge cards, to the dashboard and embedded components that businesses use to manage their card programs. We build the infrastructure and interfaces that let platforms and businesses instantly create, distribute, and control payment cards at scale, and we're responsible for the full cardholder lifecycle across commercial and consumer issuing. We're at an inflection point. Issuing is expanding into new geographies, new card categories (stablecoin, healthcare, consumer credit), and deeper partnerships with major platforms. This is a high-impact, high-visibility role at the intersection of financial infrastructure, developer-facing APIs, and end-user product experience, and it requires someone who can lead technically across a complex, fast-moving domain. What you'll do Define and drive the technical strategy for Issuing, including multi-year architecture decisions across backend services, APIs, and user-facing surfaces. Own end-to-end technical solutions for critical systems, authorization flows, spending controls, cardholder management, and card program configuration, ensuring they are reliable, scalable, and operationally sound. Identify and lead infrastructure investments that unlock new issuing models, including geographic expansion, new card program types, and platform-level extensibility

S
Stripe
📍 Us Remote• Full-time• Remote
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe’s mission is to build the economic infrastructure for the internet. Risk and Compliance Engineering builds products to lower Stripe’s financial and regulatory risk at scale, while retaining a best in class user experience. We build user-facing products, and backend platforms to ensure Stripe’s users are compliant with regulatory and financial partner requirements. We protect Stripe’s brand while also protecting the company from financial losses that can put Stripe’s business at risk. The Risk TPM team is crucial to driving programs to address these challenges. What you’ll do As a Technical Program Manager in the Risk and Compliance space, you will play a key role across engineering, product, and strategy to drive programs that span across the Risk organization and Stripe. You will be responsible for the successful definition, cross-functional strategy, planning and execution of large-scale technical programs that help to solve complex problems and empower Stripe and our users. Responsibilities Execute on technical programs that require deep systems and engineering knowledge. Partner with Engineering Managers, Product Managers and other cross functional partners to define, scope, and drive large programs to conclusion. Develop, implement, and iterate on program management techniques, frameworks, and KPIs to achieve goals with well defined success criteria. Make substantial improvements to how the teams you work with communi

REMOTEpythonsqlai
View job →
S
Stripe
📍 Seattle• Full-time• From $156.8K/yr
1mo ago

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Design, develop, maintain, and support application programming interfaces (APIs), backend services, and distributed systems using Ruby, Java, Go and Scala. Build, test, deploy, and maintain software infrastructure to support large-scale production systems. Design and develop software systems using analytical techniques and mathematical models to evaluate system performance, predict outcomes, and assess design tradeoffs. Develop and execute software testing, validation, debugging, and documentation procedures to ensure system reliability and performance. Diagnose and resolve production issues across multiple services and layers of the technology stack. Analyze user needs and software requirements to determine technical feasibility, design approaches, and implementation timelines within cost and resource constraints. Collaborate with cross-functional engineering teams to design and implement new features for large-scale systems. Design and build systems to securely store, manage, and modify production data and services. Improve engineering standards, development tools, and software development processes to enhance system quality and efficiency. Who you are Minimum requirements Must have a Bachelor's degree or foreign equivalent in Computer Science, Software Engineering, or a related field, plus 2 years of software development experience. Must also have 2 years of experi

javaairuby
View job →
S
Stripe
📍 New York• Full-time
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Secure Devices is responsible for ensuring that every client endpoint at Stripe adheres to our rigorous security standards. Our services play a crucial role in detecting and preventing data loss, restricting software execution to only approved software, and providing attestation capabilities to securely manage device identities. We operate both on-device and backend services across multiple platform types. Our users-first approach ensures that we’re empowering Stripes to be as productive as possible while protecting user data. What you’ll do As a software engineer on Secure Devices, you will work at the intersection of software development, security, and client platform engineering. You will work with teams across Security, Infrastructure and Corporate Engineering to drive strategic projects to better secure Stripe endpoints, build infrastructure for supporting new platforms, and operate services critical to securing over 10,000 Stripe devices. Responsibilities Contribute to the secure design and implementation of Stripe’s mobile expansion initiative Act as the subject matter expert on iOS security by advising partner teams on iOS security best practices and secure-by-design architectures Design, build and maintain Stripe’s endpoint security software. This includes developing telemetry and prevention capabilities via macOS system extensions that run on all Stripe macOS devices Collaborate closely with partner teams to define and measure the

awslinuxrest
View job →
O
1mo ago

About the Team API Enterprise Controls is part of the API Infrastructure organization and owns the platform capabilities that help developers, startups, and enterprises adopt the OpenAI API securely and confidently. We build the systems underneath our APIs and developer platform across authentication and identity, service accounts and key management, secure networking, compliance, auditability, observability, and operational controls. Our users are developers and teams running critical applications on OpenAI, and we partner closely with Product, go-to-market, security, and infrastructure teams to turn their most important needs into reliable, intuitive platform capabilities. About the Role We are looking for an exceptional backend software engineer to help define and ship the enterprise capabilities our API Platform needs to scale.; this is a product-engineering role grounded in deep backend systems. You will work across databases, streaming systems, request routing, authentication, and developer-facing APIs while bringing strong product judgment, developer empathy, and attention to the small details that make a platform easier to understand, trust, and operate. You will lead large cross-functional initiatives, work closely with Product and go-to-market teams, engage directly with sophisticated users, and carry ambiguous needs from discovery through design, launch, and iteration. In this role, you will: Own backend product capabilities end to end across authentication and identity, service accounts and key controls, secure networking, compliance, observability, and operational workflows. Partner with Product, go-to-market, security, infrastructure teams, and sophisticated customers to identify needs, shape the roadmap, and lead large cross-functional projects from design through launch. Design developer-facing APIs, system behavior, configuration, error handling, safe defaults, auditing, and notifications with exceptional care for the details that define a great dev

typescriptpythonaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team API Agents builds the shared agent harness, tools, and infrastructure that turn OpenAI’s frontier models into systems that can reliably complete real work. We carry the capabilities behind Codex into a much broader set of products and workflows across software engineering, research, finance, healthcare, enterprise operations, and more. Our work spans search and connected context, computer use, memory, delegation and multi-agent coordination, and safe execution. Sitting at the intersection of Research, Codex, infrastructure, and applied product teams, we build reusable agent capabilities that compound across the ecosystem. About the Role We are looking for an experienced backend software engineer to build the core systems behind the next generation of agents. You will design reliable services and abstractions that help agents find the right context, use tools and computers, retain knowledge, coordinate over long-running workflows, and take action safely. The role combines deep backend and infrastructure work with strong product judgment, with opportunities to work across agent runtimes, orchestration, search, execution environments, identity and permissions, observability, and evaluations. This is software and systems engineering rather than model training: success comes from strong backend fundamentals, high agency, and the ability to turn fast-moving research capabilities into dependable production primitives. In this role, you will: Design, build, and operate the shared agent harness and backend infrastructure that power long-running, high-value workflows across OpenAI and third-party products. Build reusable capabilities across search and connected context, computer use, memory, tool execution, delegation, subagents, and multi-agent orchestration. Establish the foundations agents need to operate safely in production, including secure execution environments, identity and permissions, observability, evaluations, reliability, and cost and latency effi

typescriptpythonaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team API Multimodal builds the developer-facing products and infrastructure that bring OpenAI’s image, audio, and real-time model capabilities into the world. We are responsible for high-scale APIs for image generation, speech transcription, speech generation, and low-latency voice interactions. We partner closely with Research and Inference to bring frontier model capabilities to developers and use customer feedback to improve our models. About the Role As a software engineer on API Multimodal, you will build and operate the products and distributed systems behind OpenAI’s image, audio, and real-time APIs. You will work across model integration, API design, and production infrastructure to turn new research capabilities into reliable developer experiences. This hands-on role combines backend and systems depth with product judgment: you will own projects end to end, partner with Research, Inference, and Safety, and help make multimodal AI useful at scale. Model training experience is not required. In this role, you will: Design, build, and ship developer-facing APIs and backend services that serve frontier models. Architect low-latency streaming, request, session, and model integration systems that make complex multimodal interactions reliable and intuitive at scale. Work directly with Research to bring new model capabilities into production, shape the systems around them, and incorporate feedback from real-world developers and customers. Own the availability, latency, scalability, and cost efficiency of the services you build. Own projects from technical design and implementation through launch and ongoing iteration, while raising the team’s engineering standards. Your background might look something like: 7+ years of professional experience, excluding internships, in backend, infrastructure, platform, or product engineering roles. A track record of designing, building, and operating production backend services, developer-facing APIs, or distributed syste

typescriptpythonaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI's research training infrastructure powers how our frontier models are trained and evaluated. The Simulation team sits at the intersection between the agentic harness that powers OpenAI's products and the research infrastructure where GPT-next is trained, ensuring that our model's training environment is as realistic as possible. This team owns the integration layer that connects our production harness capabilities into the training stack. The work is highly cross-functional and high leverage: researchers depend on it to run experiments and evaluations reliably as well as to develop the next generation of harness capabilities. Failures in this surface can materially affect training velocity and correctness. About the Role We're looking for a Principal Software Engineer to lead the architecture and evolution of the Simulation Platform. You'll own a critical interface between research and engineering, building the systems, APIs, and operational patterns that let researchers use agentic coding infrastructure safely and effectively in training environments. This role is ideal for a senior backend or infrastructure engineer with strong technical judgment, product sense for highly technical users, and the ability to drive execution across multiple teams. The highest-leverage work is building robust infrastructure that supports and accelerates research without compromising engineering quality. In this role, you will Design, build, and evolve the integration between the Codex harness that powers OpenAI's products and research training infrastructure used for training GPT-next Build a platform for our LLMs to train and be evaluated in simulated environments that mimic their deployment setting as closely as possible, on every axis: agentic harness, compute substrate, timing, tools, data sources, humans in the loop, and more Own major integration surfaces end-to-end, from architecture and API design through rollout, operations, and long-term maintenance Bu

pythonawsrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

Overview: The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Software Engineer to join our Data Acquisition team. Responsibilities: Own and lead engineering projects in the area of data acquisition including web crawling, data ingestion, and search. Collaborate with other sub-teams, such as Data Processing, Architecture, and Scaling, to ensure smooth data flow and system operability. Work closely with the legal team to handle any compliance or data privacy-related matters. Develop and deploy highly scalable distributed systems capable of handling petabytes of data. Architect and implement algorithms for data indexing and search capabilities. Build and maintain backend services for data storage, including work with key-value databases and synchronization. Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks. Conduct and analyze experiments on data to provide insights into system performance. Qualifications: BS/MS/PhD in Computer Science or a related field. 4+ years of industry experience in software development. Experience with large web crawlers a plus Strong expertise in large stateful distributed systems and data processing. Proficiency in Kubernetes, and Infrastructure-as-Code concepts. Willingness and enthusiasm for trying new approaches and technologies. Ability to handle multiple tasks and adapt to changing priorities. Strong communication skills, both written and verbal. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an

awskubernetesrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team ChatGPT is evolving from answering questions to becoming a deeply personalized assistant that helps people discover, create, and make decisions across everyday life. We're building new multimodal product experiences that combine language, images, personalization, and interactive interfaces to help millions of users accomplish tasks in entirely new ways. This team sits at the intersection of AI research, product engineering, design, and consumer experiences. We move quickly, ship frequently, and work on products that define how people interact with AI every day. About the Role We're looking for exceptional full stack product engineers who love building polished consumer experiences from the ground up. You'll work across frontend, backend, AI-powered workflows, and rich interactive interfaces to create new product experiences that blend conversation, visual understanding, personalization, and commerce. You'll collaborate closely with designers, researchers, product managers, and model teams to rapidly prototype, launch, and iterate on experiences used by millions of people. This is an opportunity to help invent entirely new interaction paradigms—not just build traditional web applications. In This Role, You Will Design and build end-to-end product experiences across web services, APIs, and modern frontend applications. Partner closely with product, design, and research to rapidly prototype and launch new AI-native experiences. Build intuitive, performant interfaces that make advanced AI capabilities feel simple and delightful. Develop scalable backend systems that power personalized, real-time product experiences. Work with multimodal capabilities including text, images, and interactive UI components. Iterate quickly using user feedback, experimentation, and product metrics. Help define engineering standards, architecture, and technical direction for a fast-growing product area. You Might Thrive If You Have significant experience building consumer-facin

typescriptpythonreact
View job →
O
OpenAI
📍 New York• Full-time
1mo ago

About the team OpenAI’s mission is to build safe artificial general intelligence (AGI) which benefits all of humanity. This long-term undertaking brings the world’s best scientists, engineers, and business professionals into one lab together to accomplish this. In pursuit of this mission, our Go To Market (GTM) team is responsible for helping customers learn how to leverage and deploy our highly capable AI products across their business. The team is made of Sales, Solutions, Support, Marketing, and Partnership professionals that work together to create valuable solutions that will help bring AI to as many users as possible. About the role As an Account Director focused on Insurance you will own executive-level relationships with leading Insurance organizations. You’ll help these companies safely and effectively deploy OpenAI’s technology to accelerate financial data analysis, automate backend operations, drive AI-powered research, and personalize customer engagement. This role blends literacy, technical depth, business acumen, and relationship-driven enterprise sales. You will collaborate closely with researchers, engineers, and financial services solution strategists to design secure, compliant, and high-impact AI deployments. This role is based in New York City. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you’ll: Manage a focused portfolio of Financial Services, specifically Insurance accounts, developing long-term strategic account plans. Lead complex, multi-stakeholder sales cycles. Partner with solutions and research engineering to design pilots that demonstrate measurable business impact. Collaborate with compliance, privacy, and security teams to ensure responsible deployment of AI in regulated environments. Own a revenue and consumption target; manage forecasts and pipeline reporting. Monitor industry and regulatory trends to guide customer and product strategy. Represent Ope

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s API Multicloud team is responsible for extending OpenAI’s API platform into strategic cloud environments, starting with AWS . The team’s mission is to distribute OpenAI’s API broadly and safely by enabling key API technologies in cloud-native environments, in close partnership with Amazon and internal teams across Codex, Research, Safety Systems, and Applied. The team is focused on bringing core developer and enterprise capabilities into cloud-native environments, including cloud-hosted Codex, model customization / post-training as a service, and new stateful runtime environments for agentic workloads. This work sits at the intersection of production ML systems, developer platforms, model behavior, and large-scale infrastructure. About the Role We’re looking for a backend engineer who can quickly understand OpenAI’s models, products, and systems, then adapt first-party deployments for other cloud platforms. You’ll build backend services, APIs, SDK integrations, authentication flows, and cloud service infrastructure that let developers use OpenAI capabilities in the cloud environments where they already build. This role involves working across teams, sometimes embedded with partner product groups, to ship products quickly and across multiple platforms at the same time. It’s a strong fit for engineers who have built developer tools, especially AI-powered tools, communicate clearly across technical boundaries, and can shape architectures that support different deployment models; experience building cloud services is a strong plus. In this role, you will: Build backend and infrastructure systems that extend OpenAI’s API platform into cloud-native environments, like AWS. Design and ship cloud-contained products that allow customers to use OpenAI capabilities while keeping workloads and data within cloud environments. Help stand up cloud-hosted Codex experiences powered by the OpenAI Responses API. Build the infrastructure and runtime abstractions

typescriptpythonaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role We’re building the observability product for OpenAI—from scalable infrastructure to a rich, AI-powered UI. Our systems ingest over petabytes of logs and billions of time series metrics across our fleet. We're now layering intelligence on top—think agents that summarize SEVs, auto-generate dashboards, or help engineers debug through notebook-like UIs. We’re hiring software engineers across the stack—infra, backend, and product. You’ll join a small, gritty team building both foundational infra and novel internal tools to make OpenAI's production systems reliable, performant, and observable. What You’ll Do Own core observability infrastructure, including distributed logging, time series, and trace storage Build AI-native tools that help engineers detect, understand, and resolve issues autonomously. Contribute to UI experiences like dashboards, notebooking, or interactive debugging Collaborate closely with engineers, researchers, user ops, and other teams across the company to build the next generation observability product You Might Be a Fit If You: Have operated large-scale distributed systems in production. ( especially logging systems or some other time series databases) Thrive in ambiguous environments and roll up your sleeves to solve unscoped problems. Have full-stack chops or product sensibilities—you're excited to build real tools people use. Have strong fundamentals in systems, networking, and cloud infra (Kubernetes, AWS, etc). Bonus : built or contributed to observability systems (e.g. Prometheus, OpenTelemetry, etc). Why This Team We’re b

awskubernetesrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Cloud Agents team builds product infrastructure for long-running agents in the cloud: orchestration, sandboxing and isolation, secure environment connectivity, secrets and identity, observability, reliability, and cost controls. These agents securely connect to diverse developer and customer environments and use tools to accomplish goals. We partner closely with product, research, and infrastructure teams to turn agentic capabilities into dependable platforms for OpenAI products and developers building on OpenAI. About the Role We are looking for an experienced software engineer to help build and scale our cloud agent platform. You will design and operate systems for orchestrating agents at scale. You will work closely with product engineers on ChatGPT, API, and Codex to define the right abstractions and enable them to ship products quickly. Strong backend or infrastructure experience is important; experience with Python, Rust, distributed systems, cloud infrastructure, or product platforms is especially helpful. In this role, you will: Design and scale the orchestration, sandboxing and storage systems that run agentic workloads for Codex, ChatGPT, and the OpenAI API. Partner with product engineers to build a platform that enables them to ship quickly and turn feedback into robust abstractions. Improve reliability, security, performance, and cost efficiency for long-running agents. Deploy services that can operate across different environments and clouds. Your background might look something like: 9+ years of professional engineering experience, excluding internships, in relevant roles at technology and product-driven companies. Experience leading large-scale backend, platform, or infrastructure projects from ambiguous problem statements to production systems. Proficiency in one or more backend languages such as Python, Go, Rust, TypeScript, or similar, and the ability to move across service, platform, and product boundaries. Strong understanding

typescriptpythonaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Codex Web Layer team provides the web-based systems and user experiences for Codex across the entire stack, from the Electron-like application framework that powers the application, to the user-facing in-app browser. About the Role In this role, you will be responsible for designing and implementing infrastructure and features end-to-end for the Codex desktop client application. You will help define what it means to be a hybrid agentic/interactive web browser. The team embodies “full stack” development from the lowest-level OS integration to the highest-level interaction design. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role you will: Partner closely with product and design to conceive, design, and build features for Codex web browsing features on macOS and Windows. This role will focus mostly on the backend C++ layer and Chromium, but many features cross the full stack including some TypeScript. Partner with the wider Codex team to deliver a high-performance, stable, and secure application platform for client development. This includes API design and implementation (mostly in C++) and the infrastructure that supports deploying it (in Python, TypeScript, and agentic skills). Work with a small, experienced team of engineers on this critical and rapidly growing product. You might thrive in this role if you: Have significant experience building technically complex features end-to-end. Are a strong C++ developer, especially with experience in browser environments like Chromium and Electron. Since this role is more backend focused, general knowledge of web development and TypeScript is helpful but not required. Thrive in a fast-paced, ambiguous environment. Communicate clearly and concisely across many different roles in the organization. Are self-directed, identifying important work and executing it end-to-end. About OpenAI OpenAI is an AI

typescriptpythonaws
View job →
🔔

Get new back end td reliability lab manager jobs by email

Daily job updates · Unsubscribe anytime