ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early member of the Observability Team, you will be pivotal in building and shaping the observability experience for our internal and external customers. By joining this team, you’ll have a direct impact on the reliability and operational excellence of Basetens product systems. As Baseten scales its infrastructure across different cloud providers and diverse hardware, the volume and complexity of operational data is growing by orders of magnitude. This team is responsible for building high-throughput ingest pipelines, cost-efficient storage, and agentic diagnostic tools to ensure that we can detect, diagnose, and resolve issues in minutes rather than hours, even as the systems they operate become more complex. RESPONSIBILITIES Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces across Baseten’s multi-cloud infrastructure Own and evolve core observability platforms, driving migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth Build instrumentation libraries, SDKs, and integrations that make it easy for engineering teams to emit high-quality telemetry from their services Drive alerting and SLO infrastructure that enables teams to define, monitor, and respond to reliabi
Jobs in United States
Data Center Infrastructure Architect in San Francisco
442 active opportunities · Updated September 2026
Showing
15 jobs
Explore current data center infrastructure architect jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. At Baseten, we are building the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale, the network is the computer. We are looking for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building block in our infrastructure and unlocking the next generation of distributed inference optimizations. THE OPPORTUNITY Networking and compute are no longer separate disciplines; they are converging. The massive throughput of H100, B200, and NVL72 architectures enables and demands a new approach where communication is co-optimized alongside computation. We are entering an era where the network is an active accelerator, leveraging smart hardware offloads and direct interconnects to ensure that data movement operates at wire-speed. In this role, you will go beyond network configuration to architect the software fabric that unifies thousands of GPUs into a cohesive operating system. While you will leverage the best of the open-source ecosystem, you won't be limited by it. Where off-the-shelf solutions stop, you will build from scratch, engineering the primitives required to co-optimize communication and compute for Disaggregated Serving, Wide Expert Parallelism (WideEP), and lightening cold starts. WHAT YOU'LL DO Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack,
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE The largest, most demanding enterprises are starting to run on Baseten, and they arrive with a range of security, compliance, and procurement requirements. As a Senior Engineer on Baseten's enterprise engineering team, you'll build the capabilities that enable large organizations like Writer, HubSpot, and Notion to succeed on Baseten. Enterprise engineering authors the core building blocks, APIs, and user experiences powering the Baseten platform: identity and access management, billing, regional isolation, and self-hosted and single-tenant deployment options. This is deep product and systems work across the full stack, from designing authentication and authorization systems using standards like OAuth and OIDC to shipping the admin experiences enterprise IT teams use to manage their organization. EXAMPLE INITIATIVES Recent and upcoming work on the team: Fine-grained authorization for users, service accounts, and agentic workloads SSO and SCIM support, allowing customers to centralize and automate access to Baseten Expanding the billing platform to support evolving pricing models, advanced data exports, and controls to manage spend In-product management and enforcement of customer compliance requirements like data residency and HIPAA Securing network paths in and out of a customer's models with private connectivity and ingress and egress restrictions Allowing customers to run Baseten inside their own VPC, on-pr
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re hiring a People Business Partner to support our team during a critical phase of growth. This is a highly strategic, high-impact role for someone who has partnered closely with leadership teams and helped organizations scale with intention. You will be deeply embedded with managers, bringing clarity and rigor to how teams are structured, how leaders operate, and how talent is developed across the organization. You’ll help shape team effectiveness, identify critical talent gaps, drive talent and performance strategies that enable high-performing teams, and build the people practices and change management approaches that allow us to scale with both speed and discipline. This role requires strong business judgment and the ability to operate with deep context. You’ll partner closely with leaders to navigate complex organizational decisions, anticipate challenges before they surface, and bring a clear point of view on what great looks like at every level of the organization. RESPONSIBILITIES Strategic partnership to leadership Serve as the trusted people partner to leadership, maintaining deep business context and translating it into people priorities by proactively surfacing systemic issues, risks, and opportunities before they become urgent. Bring data-driven insights to advise management on org design, succession planning, performance, retention, and engagement. Build management capacity across the org, equ
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About North: North is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications. Why this role? This role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications. As a bridge between our core North product and our clients’ engineering teams, you’ll be at the forefront of solving complex problems and securely integrating AI into critical sectors such as finance, healthcare, and telecommunications. We’re looking for Software Engineers with Applied AI experience who can own the design, build, and deployment of agentic workflows powered by Large Language Models (LLMs), from early prototypes to production-grade AI agents, to deliver concrete business value in enterprise workflows. You’ll work closely with customers on real-world business problems, often building first-of-thei
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the Role Cohere is seeking an Indirect Tax Sr. Associate / Tax Manager to manage global indirect tax compliance, strategy, and reporting, focusing on VAT/GST and sales tax in Canada, the U.S., and other key markets. You’ll manage tax audits, optimize processes, collaborate cross-functionally, drive process and efficiency initiatives, and implement strategies to minimize liabilities while ensuring compliance. Key Responsibilities Manage end-to-end indirect tax compliance, including filings, registrations, and remittances for global VAT/GST, and US sales and use tax Streamline tax workflows, leverage tax automation tools and AI, and enhance data governance to improve efficiency Develop and implement indirect tax strategies aligned with Cohere’s business growth, product offerings (e.g., SaaS, AI services, Sovereign AI), and international expansion. Manage responses to tax audits and inquiries from authorities. Identify and mitigate tax risks, ensuring alignment with global tax regulations and internal policies Address digital services taxes and electronic services implications. Partner with Finance, Legal, and Product teams t
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About North: North is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications. Why this role? This role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications. As a bridge between our core North product and our clients’ engineering teams, you’ll be at the forefront of solving complex problems and securely integrating AI into critical sectors such as finance, healthcare, and telecommunications. We’re looking for Software Engineers with Applied AI experience who can own the design, build, and deployment of agentic workflows powered by Large Language Models (LLMs), from early prototypes to production-grade AI agents, to deliver concrete business value in enterprise workflows. You’ll work closely with customers on real-world business problems, often building first-of-thei
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role As a Senior Software Engineer on Sentry’s AI/ML team, you’ll be responsible for building the evaluation infrastructure that measures the accuracy, reliability, and real-world performance of our AI systems. This role is critical to ensuring that our debugging agents and AI-powered features behave correctly, safely, and predictably as they scale. You’ll design datasets, benchmarks, and test harnesses that turn ambiguous AI behavior into measurable signals, helping the team ship AI with confidence. In this role you will Design and build robust evaluation frameworks to measure accuracy, reliability, regressions, and edge cases in AI systems Create and curate high-quality datasets, golden test cases, and benchmarks grounded in real production data Build automated test harnesses and metrics pipelines to continuously evaluate models, prompts, and agentic workflows Partner closely with applied AI engineers and product leaders to define what “good” looks like and translate it into measurable criteria Own the evaluation lifecycle for major AI initiatives, from early experimentation through production monitoring You’ll love this job if you Care deeply about correctness, rigor, and measurement in AI systems Enjoy turning fuzzy product goals and model behavior into concrete tests and metrics Like building foundational infrastructure that unlocks faster iteration and higher confidence for the entire AI team Thrive in cross-functional environments and enjoy influencing model design through better evaluation Qualifications Minimum 5+ years of professional experience with a Bachelor’s degree in computer science, machine learni
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role At Sentry, Support is an engineering discipline. Our customers are the greatest technical minds in the world—developers at elite enterprises building the future of software—and they deserve answers that go deeper than a knowledge base link. We're looking for an APAC Technical Support Engineer based in San Francisco to join our global Support Engineering team. This role is designed to provide APAC coverage to our users; with the shift being Sunday through Thursday 4PM-12AM PST. We are architecting the Technical Support engine . We’re looking for an experienced engineer to help us redefine the standard of technical support by combining deep human expertise with autonomous agentic systems. You are a debugger of both code and systems. You will treat support volume as a data signal to build automated resolution paths, ensuring our human engineers only touch the most complex, high-impact architectural puzzles. Sentry Support Engineers aren't just clearing queues; they are Orchestrators . You will engage with our users across GitHub, Discord, and our internal systems, while acting as the Technical Lead for our Agentic Ops. You ensure that when a developer asks a complex question, our systems have the right context and a seamless "Human-in-the-Loop" path to you when deep, nuanced expertise is required. In this role you will Master the Sentry Ecosystem & Support Elite Developers Deep-Dive Debugging: Perform root-cause analysis on complex issues and distributed tracing gaps across polyglot environments. Support the Great Minds: Act as a strategic consultant for senior engineers at our largest enterprise customers, s
$200K – $240K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Platform org at Sentry is the engine that everything else runs on, spanning developer platform, core infrastructure underpinning the product, SRE, security, and IT. It's a broad, technically complex, and deeply consequential organization. We're looking for a Staff Technical Program Manager who can operate at the intersection of technical depth and strategic execution: someone who thrives in complexity, builds trust with senior engineering leaders, and has a gift for turning ambiguity into clarity and momentum. You'll report to the Head of Technical Program Management and work closely with the VP of Platform Engineering, their staff, and partner teams across Engineering, Product, and Design (EPD). This is a high-visibility role with real influence. You'll work directly with the CTO and senior leaders, shape how the Platform org operates, and help Sentry scale through one of its most important chapters. In this role, you will: Drive strategic execution. Partner with the VP of Platform Engineering and senior EPD leaders to translate priorities into clear, measurable plans. Own sequencing, milestones, and key decisions, and make sure leadership always has reliable visibility into progress and tradeoffs. Build data-driven delivery health. Establish the metrics and dashboards that reflect delivery confidence, risk, and engineering health across the Platform org. Keep planning and reporting high-signal and lightweight, focused on outcomes rather than activity. Manage capacity, dependencies, and risk. Create visibility into resourcing, cross-team dependencies, and constraints so leaders can align investment to the
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. About the role As a Senior Software Engineer, you will architect and scale the core systems that power Sentry's billing infrastructure, ensuring accuracy and reliability at massive scale. You will collaborate on building the next generation of Sentry’s usage tracking pipeline, processing hundreds of billions of events daily with low latency and financial-grade accuracy. You will help design flexible pricing primitives that support everything from per-event usage billing to complex enterprise contracts, enabling product and sales teams to experiment rapidly while maintaining revenue accuracy and reduced time-to-market for new products. You will contribute to technical decisions on data consistency challenges unique to billing—like handling event delays, retroactive pricing changes, and distributed count reconciliation across our infrastructure. You'll love this job if you Want to solve the "easy to explain, hard to build" problems—like ensuring a customer's bill matches their usage perfectly, even when processing hundreds of billions of events daily across distributed
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s Partnerships team unlocks “one to many” relationships and product partnerships with technology platforms to allow more customers and consumers to benefit from Plaid’s solutions. This role will focus specifically on partnerships that unlock new markets and scale new use cases alongside our Credit Team. Our Credit Team is building the future of lending by making cash flow data as ubiquitous and trusted as traditional credit data. Our mission is to expand access to more affordable credit by empowering lenders with real-time financial data that improves risk assessment and decision-making at scale. In this role, you will lead partnerships that make it easier for lenders to adopt cash flow underwriting. The potential partner ecosystem is broad and includes the software platforms, capital providers, and other companies involved throughout the lending lifecycle. You will assess Plaid’s partnership needs, identify and prioritize strategic partners, develop creative commercial structures, and negotiate complex agreements. You will work closely with internal and external stakeholders across product, engineering, legal, go-to-market, and other functions to build partnerships that create meaningful valu
What you’ll do Partner with medical image reconstruction scientists / engineers to build ML components that improve reconstruction quality, speed, robustness, or quantitative accuracy. Define training/evaluation pipelines, datasets, and metrics that map to user needs and design requirements. Productionize models: inference performance, reproducibility, monitoring for drift/regressions, and safe fallbacks. Collaborate on hybrid algorithms, incorporating physics and learned priors, denoisers, learned regularizers, and quality estimation. Help build tooling for rapid experimentation as well as rigorous verification of algorithm changes. What we’re looking for Strong applied ML experience plus comfort with signal processing / imaging or adjacent domains. Ability to move fluidly between research prototypes and production-quality systems. Strong evaluation discipline: metrics, ablations, data leakage avoidance, and reproducibility. A demonstrated track record of applying ML to physics-based or inverse problems (i.e., shipped projects, a portfolio, or publications.) Useful experience ML for imaging/inverse problems (or adjacent) with strong evaluation discipline and comfort with GPU performance constraints. Pragmatic production mindset: reproducible training/inference, regression testing, and safe deployment in high-stakes contexts. A background in computational physics or scientific computing. Leverage ML-based methods such as PiNNs and Neural Operators to solve partial differential equations arising in ultrasound simulation and imaging. Experience in Agentic-SciML is a plus. Hands-on experience with data curation for ML: building datasets from messy, real-world sources, defining ground truth, and managing labeling or simulation pipelines. Background in data assimilation: combining observations with physics-based models (Kalman filtering, variational methods, ensemble approaches, or learned variants).
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s Product team builds the network that powers the future of financial services. Our mission is to unlock financial freedom for everyone through open finance. Product Managers at Plaid are curious, customer-obsessed, and move quickly to deliver value. They take ownership, sweat the details, and make sound decisions with imperfect information. As a Product Manager on the AI Foundations team, you will drive Plaid’s AI strategy by building the data and intelligence layer that powers smarter financial experiences. You will work across engineering, data science, and research to develop scalable AI systems—from core embeddings and representation learning to applied model integrations that enhance developer and consumer outcomes. This role is for an experienced PM who thrives at the intersection of AI and platform products. You are technically fluent, strategic, and execution-oriented. You enjoy turning advanced machine-learning capabilities into reliable, trusted infrastructure that scales across Plaid’s ecosystem. Responsibilities AI Platform Vision: Define the strategy, roadmap, and success metrics for Plaid’s core AI and data foundation, enabling smarter, more adaptive financial products across th
Other cities to consider
More places hiring for this role
Get new data center infrastructure architect jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime