ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg
Jobs in United States
Technical Design Architect in San Francisco
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current technical design architect jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Applications Engineering organization builds and operates the products (such as ChatGPT & Codex) that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role We’re hiring Backend Software Engineers to design and implement safe services and infrastructure that power our core products. What You’ll Do Architect, build, and improve scalable backend systems and APIs. Drive performance, reliability, and safety across distributed services. Implement data storage, retrieval, compute, and integration solutions. Participate in long-term architectural planning and technical design reviews. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. You Might Thrive Here If You: Have strong experience with distributed systems, APIs, and backend languages (e.g., Go, Python, Rust, C++). Have experience setting up and maintaining production backend services and data pipelines. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Enjoy building resilient services that handle large scale and complexity. Are self-directed and enjoy figuring out the best way to solve a particular problem Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek
About the Team OpenAI’s Applications Engineering organization builds and operates the products (such as ChatGPT & Codex) that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role We’re hiring Full-Stack Software Engineers to design and implement safe services, systems and infrastructure that power our core products. In this role, you will: Architect, build, and improve frontend and backend systems. Work across the full stack to build products and systems from initial exploration through launch readiness. Participate in long-term architectural planning and technical design reviews. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. You might thrive in this role if you: Have built and shipped full-stack apps or systems end-to-end — in fast-moving, startup-like environments. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Enjoy building resilient products and services that handle large scale and complexity. Are self-directed and enjoy figuring out the best way to solve a particular problem Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool t
About the Team API Multimodal builds the developer-facing products and infrastructure that bring OpenAI’s image, audio, and real-time model capabilities into the world. We are responsible for high-scale APIs for image generation, speech transcription, speech generation, and low-latency voice interactions. We partner closely with Research and Inference to bring frontier model capabilities to developers and use customer feedback to improve our models. About the Role As a software engineer on API Multimodal, you will build and operate the products and distributed systems behind OpenAI’s image, audio, and real-time APIs. You will work across model integration, API design, and production infrastructure to turn new research capabilities into reliable developer experiences. This hands-on role combines backend and systems depth with product judgment: you will own projects end to end, partner with Research, Inference, and Safety, and help make multimodal AI useful at scale. Model training experience is not required. In this role, you will: Design, build, and ship developer-facing APIs and backend services that serve frontier models. Architect low-latency streaming, request, session, and model integration systems that make complex multimodal interactions reliable and intuitive at scale. Work directly with Research to bring new model capabilities into production, shape the systems around them, and incorporate feedback from real-world developers and customers. Own the availability, latency, scalability, and cost efficiency of the services you build. Own projects from technical design and implementation through launch and ongoing iteration, while raising the team’s engineering standards. Your background might look something like: 7+ years of professional experience, excluding internships, in backend, infrastructure, platform, or product engineering roles. A track record of designing, building, and operating production backend services, developer-facing APIs, or distributed syste
From $240K/yr
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It's an environment where engineering is a craft, and builders become leaders. What you'll do The Bill Pay team is a high-functioning, product led team supported by product and design, building and leading the next generation of Accounts Payable (AP) software. We’re looking for an engineering leader to have strong opinions on technical design and the willingness and eagerness to jump on customer calls. This is one of the most important roles in Brex engineering as it perfectly sits between our Expense Management and Banking products. This is a founder-shaped problem: the domain is es
About the Team The People Technology team builds and operates the systems that support how OpenAI hires, develops, and supports its people. The team brings together People Systems and People Innovation Labs, a product engineering group focused on rethinking how we find and retain exceptional talent and help employees do their best work. People Systems owns the company’s core people-technology ecosystem, including platforms such as Workday and Ashby. People Innovation Labs builds new employee and People Team experiences on top of that foundation, including OpenHouse, our internal employee hub, and AI-powered products and automations. Together, we are working toward a model in which our enterprise systems provide reliable data, controls, and core business logic, while employees and managers can complete more of their work through simple, integrated, and AI-native experiences. About the Role We are looking for a People Systems Lead to manage the People Systems team and shape how our core systems evolve. You will be responsible for the reliability and effectiveness of our current environment while helping us move beyond the constraints of traditional enterprise software. This includes stabilizing and improving platforms such as Workday and Ashby, designing the integrations that connect them to the broader technology ecosystem, and partnering with People Innovation Labs to surface workflows through OpenHouse, Slack, and AI-powered experiences. This role requires someone who is comfortable moving between strategy, technical design, and team leadership. You should understand People systems deeply, be able to work through integration and architecture decisions with engineers, and translate complex organizational needs into scalable solutions. You will also manage vendor relationships, develop the People Systems team, and drive alignment across People, Engineering, Finance, Security, Legal, and other partners. This role could be a fit for someone who has grown up in People S
About the Team The Cooperative AI team is scaling to devices and embedded operations and user experiences. Our model-powered scaled workforce and knowledge system are moving on to the edge and powering our devices and edge experiences. By leveraging OpenAI’s state-of-the-art models and technologies, in production and in the lab, we develop systems that reason and work autonomously with customers and with our workforce responsible for operational work. We carry real workloads for critical systems across finance, sales, customer support, integrity, product insights, internal operations, and now devices to drive insights into product and industry. We partner closely with internal teams and external customers globally, operating in a hyper-fast feedback loop where many of our users are just a few steps away. This proximity allows us to iterate quickly, validate impact in real time, and accelerate industry impacting learnings and systems builds. We are a highly multidisciplinary, self-contained team focused on transforming the workplace via smart systems, knowledge, scalable and reliable primitives that apply world-class AI capabilities across domains. Our mission is to learn fast and transform how humans collaborate with AI at scale. About the Role We are looking for a Technical Lead Manager to lead a team of engineers building AI-native embedded experiences and operations-forward systems. In this role, you will perform both hands-on technical leadership and small team management. You will drive business outcomes, architecture and technical strategy for complex systems, contribute directly to implementation, and help grow a high-performing team. You will work closely with internal stakeholders to understand operational challenges, identify high-leverage opportunities for automation, and deliver solutions that create measurable impact. This role is ideal for someone who enjoys moving between technical design, coding, mentoring engineers, and working directly with users t
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity As a Member of Technical Staff on AI Infrastructure, you will build and maintain the foundational systems and distributed infrastructure that power AI model post training, inference, and data pipelines. You will collaborate with engineering and research teams to ensure performance, scalability, and reliability of critical AI systems. What You’ll Do Design and implement large-scale, distributed AI infrastructure and services Optimize performance for GPU/xPU accelerators and cloud environments Build tools for observability, reliability, and scaling of AI workloads Partner with cross-functional teams to define AI infrastructure requirements and roadmap Contribute to architectural design and system longevity About You Have experience with GenAI infrastructure systems, distributed systems, cloud computing, and high-performance infrastructure Are proficient in programming languages like Python, Go, or similar Understand scaling challenges specific to AI workloads and accelerators Thrive in fast-paced, collaborative engineering environments The reasonably estimated base salary for this role ranges from $256,000.00 to $276,
About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work
About the team The AI Deployment Engineering team is responsible for helping developers and enterprises safely and effectively deploy OpenAI technologies in production. We act as trusted technical advisors and thought partners for customers, working side by side with their teams to identify high-value use cases, design practical architectures, and move from prototype to durable deployment. Cybersecurity is one of the most urgent domains where AI can help. Security teams are under pressure to reason across code, logs, infrastructure, tickets, alerts, and vulnerability data faster than ever. As frontier models become more capable, organizations need deep technical guidance on how to evaluate, validate, and safely deploy AI systems in security-critical workflows. About the role We are looking for a Cyber AI Deployment Engineer to partner with customers and help them apply OpenAI models, APIs, Codex, and agentic workflows to real cybersecurity use cases. You will work with CISOs, security executives, application security leaders, SOC teams, security engineering teams, and hands-on practitioners to identify where AI can create measurable security outcomes. This is a customer-facing technical role for someone who can move fluidly between executive strategy, practitioner-level cyber depth, and hands-on solution design. You will help customers evaluate and deploy workflows such as secure code review, vulnerability triage, threat modeling, remediation, SOC and incident response workflows, detection engineering, cloud security, GRC automation, and security validation. You will collaborate closely with Sales, Solutions Engineering, Product, Engineering, Research, and Security to turn customer needs into safe deployment patterns, reusable field assets, and product feedback. This role is based in our San Francisco HQ. We offer relocation support to new employees. In this role, you will: Deeply embed with strategic customers as the technical lead for AI-enabled cybersecurity work
From $151K/yr
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity We are seeking an experienced and vision-driven Lead Enterprise Systems Engineer to join our engineering team. In this role, you will bridge the gap between business objectives, solution architecture, and hands-on execution. The ideal candidate remains actively involved in coding (roughly 70–80% of the time) while serving as the primary technical point of contact for project stakeholders. What you'll do Technical Vision & Solution Architecture Lead the architectural design, development, and deployment of resilient, scalable solutions in our Salesforce Platform for both Sales & CPQ. Translate business and product requirements into clear, technical roadmaps and system specifications. Establish engineering best practices, design patterns, coding standards, and testing strategies. Hands-On Execution & Quality Assurance Write clean, maintainable, and highly efficient APEX code alongside the Salesforce development team. Conduct thorough code reviews to ensure quality, security, and performance. Manage technical debt, proactively balancing speed of delivery with long-term system health. Team Leadership & Mentorship Provide technical guidance, direct support, and actionable feedback to Salesforce engineers. Mentor team members to foster technical growth and career advancement. Lead agile ceremonies (sprint planning, daily stand-ups, technical grooming, post-mortems). Cross-Functional Collaboration Partner closely with Technical Managers, Enterpri
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s mission is to unlock financial freedom for everyone by making money movement and access to financial data simple and secure. As a Software Engineer, you will design and build the systems that power how millions of people connect to their finances. You will work across the stack, from reliable backend services and APIs to intuitive applications that bring those systems to life. You will collaborate with engineers, product managers, and designers to ship products that make financial services more accessible and transparent. At Plaid, engineers take ownership early, grow quickly, and see their work reach millions of users. Responsibilities: Design & Development: Build and maintain backend services with a focus on performance, reliability and scalability. Collaboration: Work closely with product managers and other stakeholders to define and implement new features that meet product and customer needs. Code Quality: Write clean, maintainable and efficient code. Testing & Debugging: Develop automated tests to ensure the quality and reliability of the codebase. Troubleshoot and resolve issues. Engage in hands-on coding and architectural design, setting and maintaining high technical standard
From £135K/yr
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Who are we? Cohere is at the forefront of AI innovation, building cutting-edge language models and AI systems. Our Finance team plays a critical role in supporting our rapid growth and ensuring operational excellence. We value technical expertise, collaborative problem-solving, and a commitment to accuracy and compliance. This role offers the opportunity to make a significant impact on our financial infrastructure while working at the intersection of finance and transformative technology. As we scale our business and navigate an evolving financial landscape, you'll play a pivotal role in driving operational excellence and serving as a key partner on accounting implications of major business decisions. In this role, you will: Own the full lifecycle management of our NetSuite ERP environment, including architecture design, configuration optimization, integration development, and strategic roadmap planning to support General Ledger, AP, AR, Procure-to-Pay, Fixed Assets, Revenue Recognition, Audit, and International Consolidations & Reporting Drive end-to-end process excellence across Order-to-Cash and Procure-to-Pay workflows,
About the Team The GTM Enablement team helps OpenAI’s customer-facing organizations turn rapidly evolving AI capabilities into consistent, high-quality customer outcomes. We build the onboarding, learning experiences, playbooks, and knowledge systems that help teams develop technical depth, stay current, and confidently guide customers through successful AI adoption. About the Role We’re hiring a Field Enablement Lead, Technical Success to design and scale enablement for our rapidly growing technical customer-facing teams. You will own programs spanning onboarding, continuous skill development, technical pitches and demos, and subject-matter-expert knowledge sharing. Working closely with Technical Success leaders, Product Enablement, and technical SMEs, you will turn complex product knowledge and field experience into practical systems that improve readiness, consistency, and the quality of customer interactions and deployments. In this role, you will: Design, implement, and scale comprehensive enablement programs aligned to Technical Success onboarding, role-based skill development, and ongoing readiness needs. Redefine and operate the Subject Matter Expert (SME) program, creating clear pathways for technical experts to share knowledge and raise technical depth and consistency across GTM. Own and proactively maintain a versioned repository of technical pitches, demos, playbooks, and launch-ready assets. Partner with Technical Success leadership, Product Enablement, and product SMEs to identify skill gaps and deliver targeted learning interventions. Capture, vet, organize, and make field- and SME-generated technical content easy to discover, trust, and reuse. You might thrive in this role if you have: 5+ years of experience in technical enablement, solutions engineering, solutions architecture, technical success, or a related role. A proven track record of designing and scaling technical enablement programs in high-growth SaaS or technology environments. Exceptional
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des
Other cities to consider
More places hiring for this role
Get new technical design architect jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime