ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Platform team within the Infrastructure organization. As a senior member of Baseten's Platform Team, you will own the systems that let every engineer at Baseten prove their code works before it reaches production. Our product runs mission-critical AI inference for customers who measure downtime in dollars per second, which means our internal bar for correctness, performance, and failure tolerance has to be exceptional. Your focus is the full testing stack: fast and reliable unit test tooling, integration harnesses that spin up realistic environments on demand, load and performance testing for GPU-backed inference workloads, and resilience testing that deliberately breaks things so our customers never have to find out what happens when a node dies mid-request. This is a builder role with org-wide leverage. You won't be writing tests for other teams — you'll be building the frameworks, harnesses, and feedback loops that make writing good tests the path of least resistance, and you'll set the standards for what "well-tested" means at Baseten. RESPONSIBILITIES Own Baseten's testing strategy end to end — define the standards, the tiers, and the tooling that engineering teams build against. Build and maintain unit, integration, load and performance testing frameworks Design end to end test infrastructure that provisions realistic dependencies
Jobs in United States
Test Analyst 3 in San Francisco
94 active opportunities · Updated October 2026
Showing
15 jobs
Explore current test analyst 3 jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
$150K – $190K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Developer Experience team is growing in San Francisco! This is a hybrid role out of our main headquarters and must be based there. We’re looking for a hands-on builder with strong opinions on great developer docs. Are you someone who loves tinkering with the newest features in your favorite products? Do you enjoy taking the next random JavaScript framework for a spin and figuring out how to make them tick? Does it especially irk you when the code snippets are wrong, or the only docs for a feature are on X? Developer Experience at Sentry lives in the intersection of shipping really cool products and getting developers set up to use them. If you’re someone who is confident in partnering with Product and Engineering to test and ship the latest features, not afraid to jump in and go hands-on to solve problems for our biggest customers, and a good eye for what good docs looks like - this is a dream role. Developer Experience at Sentry is a team of builders who are constantly looking for ways to make it easier for every developer to use Sentry. We engage with the challenges facing technical communities, to support developers, gather feedback, and help our product and engineering teams ship new capabilities. An engineer in this role should be confident to “come with an answer," propose the product solutions, identify the content and/or execution plans and people to partner with, and go execute. In this role you will Attend and even host events and meetups in the developer community Have a point of view and aren’t shy about expressing it. You get your energy from both knowing the latest trends and having a POV on
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides tools that help developers find and fix issues in their applications. The Developer Infrastructure team owns the systems that make every engineer at Sentry more effective at shipping high quality software: developer environments, CI/CD for our open source and closed source codebases, the golden path for building new services, our SDK and library publishing tooling, and the metrics and dashboards that keep it all healthy. Our goal is straightforward: developers should spend their time thinking about the software they build for customers, not the software they use to build it. That means everything from local and cloud development environments, to the CI/CD pipeline that gets code out safely and quickly, to the tooling that catches flaky tests before they cost someone a day. As AI coding agents become part of how engineers work, we're also investing in making sure our environments, CI, and deployment systems support that shift. As a Software Engineer on Dev Infra, you'll help build and scale this infrastructure end to end. In this role, you will Build and maintain the tooling that powers local and cloud-based developer environments Improve CI for our codebases and CD for our deployment pipeline, keeping both fast and reliable as the org and its infrastructure grow Define and evolve the golden path for building new services and libraries at Sentry Contribute to our SDK and library publishing tooling and release processes Build the metrics, dashboards, and flaky test detection that give the org visibility into deployment health and CI reliability Build tooling that gives engineers, and the AI codin
What you’ll do Act as the in-house electrical lead for Midjourney Medical: own the electrical architecture of the scanner and the technical direction for all board-level design. Own complex board design end-to-end: architecture, schematic capture, layout (high-speed digital, analog/mixed-signal, power), DFM/DFT, fabrication and assembly vendor management, bring-up, and revision control. Write firmware for embedded targets (MCU/SoC): drivers, real-time control loops, safety-relevant logic, bootloaders, and field update paths. Audit and update HDL (FPGA) code for high-throughput data acquisition, timing/synchronization, triggering, and pre-processing of ultrasound and sensor data streams. Define electrical interfaces and data contracts with software, recon/ML and mechanical teams: timing budgets, clocking/sync, signal integrity, connectors/harnessing, and failure modes. Establish electrical engineering rigor: design reviews, schematic/layout review checklists, bring-up procedures, test fixtures, and documentation suitable for a regulated medical device program (DHF, traceability, change control). Mentor and grow the electrical function; select and manage external design partners where leverage is high. What we’re looking for Deep experience designing complex boards from blank page to stable revision, including high-speed digital and analog/mixed-signal domains. Strong schematic and layout skills (Altium/KiCad or equivalent) with real signal integrity, power integrity, grounding, and EMI/EMC instincts. Solid embedded firmware background in C/C++ (and Python for tooling): peripherals, DMA, interrupts, real-time constraints, and debugging on hardware. Practical HDL experience (VHDL/Verilog/SystemVerilog) for data acquisition, timing, and streaming interfaces. Track record of owning bring-up and debug on real hardware: scopes, logic analyzers, and disciplined root-cause analysis. Technical leadership: clear trade-offs, strong written documentation, and the ability to set
From $200K/yr
We are looking for a talented engineer to lead evaluation of startup acquisition opportunities in the AI, cloud and security space. You will drive product evaluations, prepare and manage technical architecture discussions with target groups in Product and Engineering and provide roadmap suggestions for M&A and investments for Datadog. You will be a key partner to Datadog’s C-level leadership and highly visible at the most senior levels of Datadog. The role is reporting into the Senior Director of Product Strategy and falls within the Product organization. We are looking for an innovative and strategic thinker who is passionate about the latest tech being developed by startups in the cloud, AI and security space. The ideal candidate enjoys researching and evaluating new technologies, works effectively with cross-functional teams, and communicates opinions concisely to our leadership team. Broad understanding of relevant Cloud Technologies and deep understanding of the full coverage of Datadogs current offerings is necessary. The Corporate Development team is small and values authentic, strong-willed individuals who think creatively and proactively. This role leads technical due diligence from a product and architecture perspective across our acquisition pipeline. You'll scope and stand up proof-of-concept and sandbox environments to stress-test candidate products, then give an honest, unvarnished view of their quality and depth - the kind of assessment that holds up regardless of deal momentum. You'll assess technical architecture, flag the risks and open questions that matter most early, and turn that into a clear post-acquisition integration path. Working closely with engineering, you'll keep the evaluation focused on what's actually decision-relevant, then translate the findings into strategic recommendations for leadership and help carry the integration through by partnering with the right people on the other side. What You’l
About the Role In this role, you’ll lead and evolve the design system foundations for ChatGPT. You’ll make nuanced design judgment explicit and usable across product teams, establishing the principles, patterns, and standards that enable coherent, high-quality experiences at scale. A paramount part of this role is raising OpenAI’s craft bar. This work extends beyond maintaining a component library: you’ll define what excellent product design looks and feels like, establish clear guidance for choosing and applying interface patterns, and ensure that new ChatGPT experiences strengthen a coherent system across platforms. Working closely with research, engineering, product, and design, you’ll create a continuous learning loop between design principles, product experiences, evaluation, and evolving technology—shaping a system that becomes more capable and expressive as ChatGPT grows. This role is based in our San Francisco HQ. We offer relocation assistance to new employees. In this role you will: Define an exceptionally high standard for craft and taste, translating that standard into principles, guidance, and examples that elevate the quality of ChatGPT. Own and evolve ChatGPT’s design foundations, including typography, sizing, spacing, layout, motion, accessibility, and responsive behavior. Establish a shared source of truth across platforms and improve the systems connecting tokens, components, patterns, widgets, learning experiences, and app-like product experiences. Build a comprehensive design systems framework with clear, opinionated guidance about when and how product teams should use different interface patterns. Translate design judgment into durable standards and use prototyping and evaluation workflows to test experiences, identify recurring issues, and improve the system. Identify promising patterns across emerging product experiences and turn them into reusable, documented foundations. Partner with research, engineering, product, and design to evolve the d
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking an Actuator Gear Design Engineer to lead the development of custom gears and gear stages for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, and reliability teams. You will partner with external suppliers and internal manufacturing to create full gearbox assemblies. This role focuses on the design, integration, and validation of precision gearing, including broader knowledge around motor electromagnetics, transmission types, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will: Lead the architecture, design, and integration of custom robotic actuator gearing. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, back drivability, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-lo
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a senior Actuator Electromagnetic Design Engineer to lead the development of custom electromechanical actuators for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, reliability, and manufacturing teams. This role focuses on the design, integration, and validation of precision electromechanical systems, including motors, transmissions, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will: Lead the architecture, design, and integration of custom robotic actuators, including the design, simulation, integration and sourcing of custom electromagnetic components. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-loop performance. Devel
About the Team OpenAI’s Governance, Risk, and Compliance team helps ensure security and privacy are grounded in how our products and systems actually operate. Assurance Operations partners with Security, Engineering, Infrastructure, Product, Privacy, and Legal to make controls provable, risk decisions explicit, and audit readiness a result of well-designed systems. About the Role We are hiring a technical, product-minded GRC builder who can own consequential audits while improving the control and evidence systems behind them. You will build a reusable common control framework, use Codex to automate assurance work, validate changing system scope, and turn repeated audit friction into measurable improvements. We are looking for someone who questions inherited assumptions, solves novel problems creatively, works closely with engineers, and makes the next audit easier by improving the underlying system. You’ll be responsible for: Lead external, internal, customer, and certification audit work from scoping through evidence review, fieldwork, remediation, and closeout. Build a common control framework linking risk, control intent, implementation, owner, system, environment, evidence, and applicable frameworks. Validate actual scope and ownership instead of assuming last year's controls, product boundaries, or evidence remain accurate. Use Codex to build and test evidence checks, control mappings, request triage, owner workflows, monitoring, and remediation reporting. Partner with engineers on cloud architecture, identity, logging, data flows, software changes, vulnerabilities, and control effectiveness. Design maintainable, permission-aware tools that preserve source provenance, human review, and evidence integrity. Reduce repeated requests and operational burden for control owners through measurable workflow improvements. Define roadmaps, decision rights, milestones, success metrics, and clear cross-functional escalations. We’re looking for someone with: Direct ownership
About the Team The Core Models team shapes how our models interact with people. We view the model as the product itself, aiming for intuitive experiences that exceed user expectations and feel like magic. About the Role As a Model Designer, you’ll have an outsized impact on how our models interact and resonate with users. You’ll strike a delicate balance between maximizing the model’s capabilities, reading in between the lines in user queries to understand how best to help, and upholding user trust. We’re looking for people who are passionate about the intersection of design, technology, and user experience — and are up for the challenge of defining new human-AI interaction paradigms. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you’ll join a small team in evolving and expanding the model design function. You will: Collaborate closely with researchers to understand, predict, and design model behavior. Partner with product managers and designers across the company to ensure a cohesive voice and approach. Proactively identify ways to improve our models based on product sense, user feedback, quantitative insights, and the research roadmap. Come up with creative strategies for collecting high-quality data. Do whatever needs to be done to make our models better. You might thrive in this role if you: Possess exceptional taste, creativity, and writing skills, allowing you to craft responses that delight users. Don’t mind ambiguity — you’re happy to throw yourself into a new, unfamiliar environment, build relationships, define a problem, and make progress. Love experimentation, and are willing to test and reject new ideas when the results don’t pan out. Exhibit high levels of empathy and self-awareness required to serve everyone in the world. Enjoy tackling profound and often philosophical questions while always driving towards clarity. Demonstrate technic
About the Team The B2B Marketing team is responsible for helping businesses understand, adopt, and get value from OpenAI’s products. B2B marketing is a major and growing priority for OpenAI as we scale our work with companies, developers, and institutions around the world. About the Role Within B2B Marketing, Demand Generation builds the integrated, full-funnel engine that connects audience insights, content, field and digital experiences, paid media, lifecycle, and sales follow-through to qualified pipeline. We partner closely with Sales, Partnerships, Product Marketing, Communications, Creative, Web, RevOps, Analytics, and regional teams to create a cohesive customer experience and scale what works. We’re looking for a Senior Mid-Market & SMB Strategist to own demand strategy for two high-velocity segments with distinct customer needs, buying journeys, and sales motions. You’ll translate segment goals into an integrated portfolio of programs, identify the highest-leverage opportunities to improve conversion and pipeline, and align channel teams and GTM partners around a shared audience, message, handoff, and measurement plan. This is a hands-on role for someone who combines strong strategic segmentation with a practical, test-and-learn operating mindset. In this role, you will: Own the Mid-Market and SMB demand strategy, including segment priorities, audience definitions, pipeline goals, program portfolio, investment recommendations, and quarterly roadmap. Translate segment insights into integrated campaigns and always-on journeys across web, email, paid media, webinars, content, partners, and sales-assisted follow-up. Partner with Product Marketing and Sales to define ideal customer profiles, buyer needs, priority use cases, value propositions, and offers for each segment. Work with Web, Lifecycle, RevOps, SDR, and Sales to improve high-intent conversion paths, qualification, routing, speed-to-lead, and follow-up quality. Build a disciplined testing agenda ac
About the Team The Personal AGI team is responsible for training and improving pre-trained models to be deployed into ChatGPT, the API, and potential future products. In the Model Experience team, we shape the default character and behavior of ChatGPT: how the model communicates, responds to users, uses its capabilities, and behaves across different contexts and languages. Our goal is to make every interaction with ChatGPT thoughtful, helpful, and trustworthy. We take an opinionated view of what good human–AI interaction should look like, then turn that vision into real model behavior through human data, evaluations, reward models, and post-training. Our work sits at the intersection of research, product, and model design. We partner closely with teams across OpenAI to conduct research and ensure our models are thoughtful, safe, reliable to serve millions of users. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models. Our team works in research areas combining reinforcement learning and products. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research and the quality of human-AI interaction. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve model capability and performance. Collaborate closely with the other research and product teams, allowing customers to optimize their own models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. You might thrive in this role if you: Have a deep understanding of machine learning and machine learning applications. Have good judgment about model behavior and can communicate this judgment effec
About Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About Role We are looking for an Operations Program Manager - Robotics Data Acquisition to own the day-to-day operating rhythm in our data collection facilities. You will work closely with operators, technicians, program managers, and engineers to keep rigs ready, campaigns moving, issues resolved, and performance improving. This is a hands-on operations role that requires you to be comfortable spending time on the floor, working through ambiguity, and using data to make the operation more reliable and efficient. This role is based in San Francisco, CA and requires in-person presence 5 days a week. In this role you will: Coordinate daily operations readiness across workstations, operators, materials. Track core operating metrics including utilization, cycle time, throughput, downtime, operator productivity, and data quality. Identify bottlenecks through workflow analysis, time studies, and capacity modeling, then drive practical fixes. Execute the rollout of new hardware, sensors, tools, and process changes with Engineering, Operations, Facilities, Supply Chain, and Safety. Identify equipment readiness issues and coordinate with technical support to keep workstations, and test equipment calibrated, configured, maintained, and ready for rollouts and evaluations. Lead root cause analysis for recurring operational issues and follow through on corrective actions. Provide operation input to create and maintain SOPs, work instructions, training materials, and process controls. Identify and flag resource constraints and manage issue escala
About the Team The Personalization-Memory team, within OpenAI's broader Personal AGI organization, is focused on developing agents that can learn from prior interactions in order to become more helpful and efficient over time. We build general-purpose memory and personalization capabilities that transfer across ChatGPT and other agentic products, and we collaborate with applied engineering on the product surfaces that allow users to interact with memory. About the Role As a Research Engineer / Research Scientist on the Personalization-Memory team, you will research and develop improvements to memory usage and personalization in OpenAI's frontier models. Our team works on reinforcement learning, dataset creation, evaluations, and other post-training methods. We partner closely with research and product teams across the company to realize the vision of a truly personalized ChatGPT. We're looking for individuals who have a background in frontier model post-training, are able to iterate quickly, and who are passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda for improving memory use and personalization in frontier models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. Collaborate closely with the research and product teams to influence the shape of technical solutions in the product. You might thrive in this role if you: Are passionate about personalization and building personalized assistants. Have experience working with user signals and human data to turn feedback into reliable signals for training and evaluation. Have a deep understanding of frontier model post-training and machine learning applications. Value principled approaches and research craftsmanship. Are comfortable diving into a lar
About the Team The Proactivity Research team, within OpenAI’s broader Personal AGI team, is focused on making our models in ChatGPT and future potential products proactive in ways that are truly useful. We're laying the technical foundations for AI that can anticipate what users need in real time, adapt as their goals and preferences shift, and build a deeper, evolving understanding of the person it's helping. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models’ personalization and agentic capabilities. Our team works on reinforcement learning, dataset creation, evaluations, and other post-training methods. We partner closely with research and product teams across the company to realize the vision of a highly personalized, collaborative, and proactive assistant. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve the proactivity and ability of our models to further user goals. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. Collaborate closely with the other research and product teams to influence the shape of technical solutions in the product You might thrive in this role if you: Have a deep understanding of machine learning and machine learning applications. Have a working knowledge of LLM post-training and evaluation approaches Are passionate about, or have experience thinking about, personalization and enabling users to achieve their goals Are comfortable diving into a large ML codebase to debug. Thrive in a dynamic and technically complex environment. About OpenAI
Other cities to consider
More places hiring for this role
Get new test analyst 3 jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime