About the Team The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. This role is based in San Francisco, CA. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change This role is exclusively based in our San Francisco HQ. We offer relocation assistance to new employees. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely depl
Jobs in United States
Cloud Technical Architect Dap in San Francisco
131 active opportunities · Updated October 2026
Showing
15 jobs
Explore current cloud technical architect dap jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity As a Senior Backend Engineer on the Cloud Platform team, you will play a key role in building the core systems and services that power Postman’s internal platform. You’ll help create new backend services that manage how we deploy, scale, and operate our infrastructure and product services, leveraging Java, Spring Boot, and Hibernate (JPA) on top of cloud-native technologies like Kubernetes, ArgoCD, Istio, and Terraform. This role is highly impactful: the systems you build will be used across Postman engineering, enabling faster delivery, better scalability, and a stronger developer experience. You’ll also have the opportunity to contribute to open source, shaping tools that extend beyond Postman’s boundaries. What You’ll Do Design and develop backend services in Java and Spring Boot to support Postman’s internal Cloud Platform. Architect new services that manage service deployment, lifecycle, and scaling across Kubernetes clusters. Implement GitOps workflows (ArgoCD) to support continuous delivery. Integrate with cloud-native tooling such as Istio, Helm, and Terraform. Apply strong soft
About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Order to Cash (OTC) team oversees the complete flow of commercial transactions from order intake and provisioning through billing, collections, and cash application — ensuring accuracy, compliance, and operational excellence in support of OpenAI’s mission to ensure artificial general intelligence benefits all of humanity. About the Role We are seeking a Senior Manager to build and scale Order to Cash across OpenAI’s API and ChatGPT businesses, cloud marketplaces, and strategic partnership channels. This role will own end-to-end Order Management and Billing operations across API and ChatGPT while leading OTC readiness and execution for AWS Marketplace, Google Cloud Marketplace, Oracle Cloud Marketplace, GovCloud, and future partner channels. The role will also own the end-to-end Order Management and Billing close, setting the close calendar, readiness standards, review and sign-off expectations, while leading the team responsible for execution. You will oversee the Order Management and Billing lifecycle across API, ChatGPT, marketplace, and partner transactions, spanning commercial readiness, order intake, provisioning, usage and transaction data, pricing validation, invoicing, credits, settlements, and product and partner reporting. Your work will ensure transactions are accurate, timely, complete, and supported by audit-ready controls. This is a leadership role that combines strategic ownership, cross-functional leadership, and hands-on operational execution. You will define the target operating model, lead first-of-kind launches, shape product and systems roadmaps, and oversee the resolution of complex contract modifications, non-standard pricing, usage disputes, reconciliation breaks, settlement variances, and customer- or partner-impacting escalations. You will collaborate with teams across Finance, GTM, Product, Engineering, Legal, Tax, Reven
About the Team The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. Our mission is to push the frontier of code generation and agentic reasoning, and deploy these capabilities in real-world products such as ChatGPT and the API, as well as in next-generation tools specifically designed for agentic coding. We operate across research, engineering, product, and infrastructure—owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. About the Role As a Performance & Systems Engineer on the Codex team, you will be responsible for whole-system optimization across a complex, evolving stack. Codex spans LLM inference, cloud orchestration, agentic work management, and multiple product surfaces. Your job will be to identify and land high-leverage changes—across infrastructure, modeling, and product layers—that make Codex agents significantly faster and cheaper to serve. We’re looking for generalists who thrive in ambiguity and love chasing performance bottlenecks to ground. This is a high-ownership role where your work will directly improve the experience of millions of users. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Hunt down and address inefficiencies across the Codex system stack, from agent behavior to LLM inference to container orchestration, and beyond. Build tooling to measure, profile, and optimize system performance at scale. Collaborate with researchers and engineers to land high-ROI changes that improve latency and cost. You might thrive in this role if you: Have experience operating across both ML systems and cloud infrastructure. Enjoy diving into messy, ambiguous problems and emerging with clear wins. Think holistically about performance, balancing spee
At Render, we’re building the modern cloud platform for developers creating AI-native, full-stack, multi-service applications. Our mission is to eliminate the tradeoff between the power of hyperscalers and the simplicity of developer-friendly platforms—so teams can ship fast, scale reliably, and focus on their product, not infrastructure. Unlike complex hyperscalers or ephemeral edge/serverless solutions, Render offers a developer-first experience with persistent compute, dynamic autoscaling, built-in orchestration, and observability, allowing teams to launch, scale, and manage real-world applications without writing infrastructure code or managing servers. Whether you're building LLM-powered applications, scalable SaaS products, or async processing pipelines, Render empowers teams to move fast and scale confidently from MVP to millions of users. Our platform is trusted by over 6.5 million developers worldwide and continues to grow rapidly. In February 2026, we raised an additional $100M in Series C financing, bringing our total funding to $257M, to accelerate our vision of making cloud infrastructure both powerful and intuitive—designed for the speed of modern AI development. We’re a diverse and talented team that values craft, velocity, and user experience. If you’re excited to help shape the future of the intelligent cloud and empower developers everywhere, we’d love to hear from you. Applying to Render We're seeking candidates who possess high integrity, humility, and an insatiable drive to learn. Through reasoned discussions and continuous feedback, we strive to improve both individually and collectively. We foster an environment of mutual trust and respect, empowering effective debate to achieve the best outcomes for our customers and team. We especially encourage members of underrepresented groups in the tech community to apply and understand that not all successful candidates will meet each requirement listed. Our interview process is unique to each role, an
At Render, we’re building the modern cloud platform for developers creating AI-native, full-stack, multi-service applications. Our mission is to eliminate the tradeoff between the power of hyperscalers and the simplicity of developer-friendly platforms—so teams can ship fast, scale reliably, and focus on their product, not infrastructure. Unlike complex hyperscalers or ephemeral edge/serverless solutions, Render offers a developer-first experience with persistent compute, dynamic autoscaling, built-in orchestration, and observability, allowing teams to launch, scale, and manage real-world applications without writing infrastructure code or managing servers. Whether you're building LLM-powered applications, scalable SaaS products, or async processing pipelines, Render empowers teams to move fast and scale confidently from MVP to millions of users. Our platform is trusted by over 6.5 million developers worldwide and continues to grow rapidly. In February 2026, we raised an additional $100M in Series C financing, bringing our total funding to $257M, to accelerate our vision of making cloud infrastructure both powerful and intuitive—designed for the speed of modern AI development. We’re a diverse and talented team that values craft, velocity, and user experience. If you’re excited to help shape the future of the intelligent cloud and empower developers everywhere, we’d love to hear from you. Applying to Render We're seeking candidates who possess high integrity, humility, and an insatiable drive to learn. Through reasoned discussions and continuous feedback, we strive to improve both individually and collectively. We foster an environment of mutual trust and respect, empowering effective debate to achieve the best outcomes for our customers and team. We especially encourage members of underrepresented groups in the tech community to apply and understand that not all successful candidates will meet each requirement listed. Our interview process is unique to each role, an
At Render, we’re building the modern cloud platform for developers creating AI-native, full-stack, multi-service applications. Our mission is to eliminate the tradeoff between the power of hyperscalers and the simplicity of developer-friendly platforms—so teams can ship fast, scale reliably, and focus on their product, not infrastructure. Unlike complex hyperscalers or ephemeral edge/serverless solutions, Render offers a developer-first experience with persistent compute, dynamic autoscaling, built-in orchestration, and observability, allowing teams to launch, scale, and manage real-world applications without writing infrastructure code or managing servers. Whether you're building LLM-powered applications, scalable SaaS products, or async processing pipelines, Render empowers teams to move fast and scale confidently from MVP to millions of users. Our platform is trusted by over 6.5 million developers worldwide and continues to grow rapidly. In February 2026, we raised an additional $100M in Series C financing, bringing our total funding to $257M, to accelerate our vision of making cloud infrastructure both powerful and intuitive—designed for the speed of modern AI development. We’re a diverse and talented team that values craft, velocity, and user experience. If you’re excited to help shape the future of the intelligent cloud and empower developers everywhere, we’d love to hear from you. Applying to Render We're seeking candidates who possess high integrity, humility, and an insatiable drive to learn. Through reasoned discussions and continuous feedback, we strive to improve both individually and collectively. We foster an environment of mutual trust and respect, empowering effective debate to achieve the best outcomes for our customers and team. We especially encourage members of underrepresented groups in the tech community to apply and understand that not all successful candidates will meet each requirement listed. Our interview process is unique to each role, an
About the Team The Ona team at OpenAI is helping build the software factory for the enterprise. We build infrastructure that enables AI agents to work in secure, customer-controlled cloud environments, with the context, tools, and controls they need to make progress across the software lifecycle—beyond a single developer’s laptop or active session. Our focus is helping enterprises move from experimenting with agents to using them reliably in production. That means solving challenging problems in cloud environments, orchestration, security, and collaboration, while making the experience straightforward for the people directing and reviewing the work. We’re a team that values initiative, close relationships with customers, and exceptional engineering craft. We take ownership, learn quickly, and communicate directly and kindly. About the Role We’re hiring backend-focused Product Engineers across our platform and security product teams. You’ll build infrastructure and customer-facing workflows that let developers and AI agents work reliably in parallel. You’ll work primarily in Go on APIs, complex networking, development environments, and orchestration for long-running tasks. You’ll own outcomes from understanding a user’s problem and choosing an approach through shipping, operating, and improving the solution, working closely with frontend, infrastructure, and security engineers. In this role, you will: Work directly with customers to build developer and security workflows, from getting a project running to investigating findings, reviewing agent-generated changes, and verifying fixes. Build Go services and APIs for provisioning cloud environments, running agents in customer infrastructure, and integrating with source control, CI, and other developer tools. Design reliable orchestration for long-running, parallel work, including durable state, retries, cancellation, and recovery. Build security into execution workflows through clear permissions, credential handling, is
About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure: Owner Furnished Equipment to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastr
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: Modal builds the infrastructure that lets engineers run AI workloads without the usual pain. To do this well, we need exceptional people – and that’s where you come in. As the first dedicated GTM recruiter on our Talent team, you’ll own sales, GTM, and other G&A searches end-to-end. You’ll work closely with our Head of Talent, founders, and GTM leads to shape how we hire and help bring in the people who will define what Modal becomes. What you’ll do: Drive full-cycle recruiting for key hires across GTM and G&A functions (sourcing, pitching, guiding interviews, and closing candidates) Partner with GTM leaders to understand the real work and calibrate on what great looks like Help set our hiring bar and how we evaluate talent Execute creative top-of-funnel strategies that resonate with a strong community of experienced GTM talent Deliver a fast, respectful, h
AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. About the Role We're seeking a Revenue Operations Manager with a strong track record, a builder's mindset, and a bias for action to join our in-person team in New York or SF. This is a high-impact, hands-on role. You'll own the entire revenue operations function, from top-of-funnel lead routing through deal close and commission administration. You'll work closely with our Head of Finance & People Ops and sales leadership to build the systems, dashboards, and processes that scale our go-to-market motion. What You'll Do: Own the lead routing process from inbound and partnering with marketing to ensure proper attribution Run effective territory management & strategy for Geo based decisioning Support & strategise every aspect of revenue operations in your territory Own the strategy for capacity forecasting, inputs, throughputs & outputs being the conduit back to finance in
About Supabase Supabase is the open-source backend platform trusted by millions of developers to build web apps, SaaS products, and AI-native applications. Our Partnerships team works across the ecosystem - technology partners, cloud platforms, VCs, accelerators, and startup programs - to make Supabase the default backend for the next generation of companies. About the Role We're looking for a Startup Partner Enablement & Programs Manager to bring our partner ecosystem to life through enablement, events, and programs to raise awareness of Supabase's value proposition to early stage startups. You'll own the materials, playbooks, and activations that help our partners - VCs, accelerators, incubators and entrepreneurship programs - successfully promote and build with Supabase. That includes organizing Supabase's presence at partner and VC-related events, running office hours, hackathons and VIP dinners, and turning relationships into repeatable, high-impact programs. This role sits on the Partnerships team and reports to the Head of Startups, working closely with teammates across the startup program, GTM, developer relations, and marketing. Because the bulk of our events and partner activity is concentrated in San Francisco and the Bay Area, we strongly prefer candidates based there who can regularly attend and host events in person. It's a great fit for someone who loves the operational craft of enablement and programs and wants to work at the intersection of startups, venture, and developer tools. What You'll Do Build and maintain partner enablement materials - decks, one-pagers, playbooks, and activation guides - that help partners across the ecosystem promote and build with Supabase Manage the operational cadence of partner programs: collaborating on strategy, tracking activations, and keeping program logistics running smoothly across partner types Organize Supabase's presence at partner and VC-related events — demo days, founder summits, portfolio workshops, p
About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A
About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system
About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastructure, or similarly comple
Other cities to consider
More places hiring for this role
Get new cloud technical architect dap jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime