Jobs in United States

Hardware Lead Engineer in San Francisco

184 active opportunities · Updated October 2026

Explore current hardware lead engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

S
📍 San Francisco, CA, United States· Full-time
✓ High-confidence listing

$960K – $1.4M/yr

Quick readStrong listing-quality and freshness signals

Position Overview As SingleStore’s IT Operations Engineer, you will help shape the IT toolset used by our end users. This is an active, hands-on position responsible for the planning, design, development, and Tier 1 support of several key technical areas at the SingleStore IT team, including end-user support, client engineering, executive support, and infrastructure application support. This is an incredible opportunity for someone to build upon their technical strengths and be a part of IT at SingleStore team . Roles and Responsibilities: Administering a wide variety of SaaS applications. Some main applications that need to be supported are OKTA (+ Workflows), Google Workspace, Slack, and Atlassian tools (JIRA + Confluence), MDM administration. Keep up to date with new features and new releases in these applications to identify opportunities for better automation or features that could be useful for our environment. Seize opportunities across the IT Operations team to eliminate manual work through tooling, integrations, and automation of IT workflows. Respond to tickets and execute new hire onboarding and user separation processes. Support members of the team with troubleshooting and resolution of complex issues. Design, architect, implement and maintain systems and solutions for various IT-related topics, including but not limited to staff computer hardware, operating systems, software applications, networking, videoconferencing, and printers. Partner and collaborate with all business units to help them evaluate hardware and software solutions. Able to communicate effectively and concisely with the entire company. Analyze existing processes, suggest and make improvements, and implement business processes where none exists. A desire to learn and expand your horizons; take on new challenges as the business scales Required Skills and Experience: Minimum 2 years of relevant experience Prior experience in implementing and administering Google Workspac

PythonSQLAWSGit
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.4%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a member of the Capacity Strategy & Operations team, you will sit at the intersection of supply intelligence, demand forecasting, and cross-functional execution, turning a complex, fast-moving hardware market into a predictable, reliable foundation for our customers and internal engineering teams. This is not a purely analytical role. You will own the end-to-end capacity planning process: from translating customer commitments and growth forecasts into concrete supply requirements, to coordinating fulfillment across vendors, finance, and the infrastructure team, to building the systems that make all of this repeatable and scalable. When supply is constrained and tradeoffs are unavoidable, you are the person in the room who can model the options, make a clear recommendation, and drive alignment fast. You are a strong fit if you have operated at the intersection of strategy and execution before — someone who is equally comfortable building a capacity model in a spreadsheet and running a cross-functional war room when a customer deployment is at risk. EXAMPLE INITIATIVES Demand-Supply Alignment Framework: Build and own the process that translates customer pipeline, signed commitments, and growth projections into a forward-looking GPU demand signal — so the team is never caught flat-footed when a customer scales faster than expected. Constrained Allocation Playbook: Define the decision framework for how Basete

Machine LearningAIGoExcel
C
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Coder is looking for an experienced and detail-oriented IT Generalist to join our growing team. This is a full-time role ideal for someone who thrives in a fast-paced, hands-on environment. You’ll be the go-to person for user support, SaaS administration, and IT infrastructure, playing a key role in ensuring our team stays productive, secure, and well-equipped. You'll work closely with the larger IT team to keep things running smoothly - from onboarding new hires to managing devices and licenses to supporting SOC 2 compliance initiatives. If you're a strong communicator with a passion for IT operations and solving real problems for real people, this role is for you. This position follows a hybrid work model. Candidates should be local and able to work from our local office on a regular weekly basis, with scheduling flexibility. What you’ll do here Act as the first line of IT support for employees via Slack and our internal ticketing system Manage user onboarding/offboarding, including account provisioning through Okta and license management via our SaaS management systems Administer and maintain core tools like Google Workspace, Slack, Jamf, 1Password, and other business-critical SaaS apps Set up and manage macOS hardware inventory, including procurement, configuration, and asset tracking Support SOC 2 and GDPR compliance by following processes for access controls, audits, and data security Help scale IT operations as we grow - identify gaps, recommend tooling, and streamline support workflows Document processes and contribute to internal knowledge bases to improve self-service and transparency What we’re looking for Have 2-4 years of experience in IT support or systems administration Are proficient with Okta, Google Workspace, Slack, Zoom, and Apple device management (Jamf preferred) Are comfortable working independently in a remote environment and prioritizing across a variety of IT tasks Have a security-first mindset and understand the importance of access contro

AWSRestAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. Product at Baseten Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the first people who will help define it. You'll work directly with our founders and with some of the best systems and infrastructure engineers in the world, and you'll set the standard for building great AI Infrastructure. PMs at Baseten don't sit above engineers - you earn ownership by being technical, finding the truth in front of customers, building great cross-functional relationships, and shipping great product experiences. The role Getting a model into production still takes real expertise — choosing a serving engine, sizing hardware, tuning it, wiring it into an app. We want a developer to go from "it runs on my laptop" to "it's serving production traffic" in minutes, on their own. You'll own the entire experience a developer touches to deploy and iterate: the CLI and SDKs, the console, onboarding, model discovery, deployment configuration, truss, and the increasingly agent-driven ways developers build. Your job is to make Baseten synonymous with Great DevEx and make it effortless to drive and self-serve deploy models on Baseten for far more developers than it is today. Impact and outcomes you'll drive You will collapse time-to-production — take a developer from first sign-up to a running, maint

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are looking for an IT Support / Operations Engineer to join Baseten as we continue to scale our IT team. In this role, you will play a critical part in bringing our technical support entirely in-house to provide a seamless, high-touch experience for all Baseten employees. As we continue to scale, you will be the primary point of contact for day-to-day technical issues, allowing you to have a direct impact on our team's productivity and overall office environment. This position is ideal for a hands-on problem solver who enjoys a mix of hardware and software troubleshooting, user lifecycle management, and maintaining the physical IT infrastructure of a modern office. While you will focus heavily on elevating our internal support standards, you will also assist with systems administration and workflow automation as our company evolves. This is a hybrid role based out of our San Francisco or New York office, following our standard policy of three days per week in-person to ensure our physical office and AV systems remain high-performing and reliable. RESPONSIBILITIES Serve as the escalation point for day-to-day technical support, diagnosing and resolving hardware and software issues across our Mac and Windows fleet Manage user lifecycle administration including provisioning, deprovisioning, and access management across all systems and services Own the IT onboarding experience for new employees — from laptop set

Machine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early member of the Observability Team, you will be pivotal in building and shaping the observability experience for our internal and external customers. By joining this team, you’ll have a direct impact on the reliability and operational excellence of Basetens product systems. As Baseten scales its infrastructure across different cloud providers and diverse hardware, the volume and complexity of operational data is growing by orders of magnitude. This team is responsible for building high-throughput ingest pipelines, cost-efficient storage, and agentic diagnostic tools to ensure that we can detect, diagnose, and resolve issues in minutes rather than hours, even as the systems they operate become more complex. RESPONSIBILITIES Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces across Baseten’s multi-cloud infrastructure Own and evolve core observability platforms, driving migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth Build instrumentation libraries, SDKs, and integrations that make it easy for engineering teams to emit high-quality telemetry from their services Drive alerting and SLO infrastructure that enables teams to define, monitor, and respond to reliabi

PythonRestMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Baseten’s Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost‑efficient. As part of this team, you’ll focus on Model APIs — the infrastructure powering our hosted API endpoints for the latest open‑source models. This work spans distributed systems, model serving, and developer experience. You’ll join a small, high‑impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale. RESPONSIBILITIES: Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations Productionize performa

KubernetesMachine LearningAIGo
M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Execute weekly system-level exploratory testing across the scanner and supporting software; log and triage issues with clear reproduction steps. Work with engineering to debug root cause and validate fixes. Help maintain the DHF and traceability between user needs, design requirements, tests, and results. Own practical test execution logistics (fixtures, test data, environments, calibration artifacts) and keep things repeatable. Help build the continuous testing strategy: automated tests where feasible, plus structured manual and system tests. Support V&V activities, including coordination with external partners as needed. What we’re looking for Strong hands-on testing instincts for complex electromechanical systems with substantial software. Ability to write clear bug reports and communicate risk/impact. Experience building and maintaining test plans/protocols; comfort operating lab equipment and debugging across layers. Useful experience Experience testing complex systems end-to-end (automation where it pays off, plus hands-on hardware/instrumentation). Medical device or other safety-critical environments and comfort translating risk into practical test coverage.

M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Be the generalist EE for the scanner system: integration, bring-up, debugging, and making the electrical side of the device reliable and serviceable. Own ultrasound experimentations that feeds the image reconstruction team Design and execute experiment setups for transducer characterization (element sensitivity, bandwidth, cross-talk mapping, beam profile measurements) and ex vivo / phantom clinical testing. Acquire, process, and analyze RF and baseband signals for data quality assessment and benchmarking. Design simple boards and adapters as needed (monitoring, power/safety, interface/conditioning), and take them from prototype through a stable revision. Prototype quickly, then harden what works: wiring/harnessing, grounding, safety interlocks, and reliable integration across subsystems. Own practical test setups and documentation (fixtures, scripts, procedures) that make experiments repeatable and results comparable over time. What we’re looking for Strong hands-on EE background with experience building, debugging, and iterating on real systems in the lab. Solid understanding of signal processing fundamentals — knows what to measure, how to condition and digitize it, and how to evaluate signal quality in the context of an imaging system (SNR, bandwidth, dynamic range, artifacts). Comfortable spanning system integration + occasional design work (schematics/layout reviews or light PCB design) in a fast-moving environment. Ability to work at the boundary between hardware and algorithms: measure reality, communicate constraints, and help close gaps vs simulation. High agency and practicality: able to set up experiments, get trustworthy data, and unblock others on a lean team. Useful experience Analog/mixed-signal, or high-speed data capture experience; strong instincts for instrumentation and noise/debugging. Ultrasound or acoustic sensor handling: hydrophone calibration and field mapping, transducer impedance characterization, element-level sensitivity

GitAIGoRust
M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Legal team plays a crucial role in advancing our mission by tackling innovative and fundamental legal issues in AI. The team includes professionals from diverse legal fields—technology, AI, infrastructure, privacy, IP, corporate, employment, tax, regulatory, and litigation—who collaborate closely with colleagues across the company. If you are passionate about being a technology lawyer working on cutting-edge challenges, you’ll thrive here. About the Role We’re seeking a senior lawyer to participate in commercial legal strategy and execution across OpenAI’s fast-growing infrastructure portfolio. This is a cross-functional role that will partner closely with procurement, supply chain, partnerships, finance, and product teams to structure, negotiate, and manage the transactions that will support OpenAI’s long-term infrastructure ambitions. We’re looking for an experienced infrastructure transactions lawyer who thrives in ambiguity and wants to help define the commercial playbook for infrastructure efforts in the AI era. This role reports to the Associate General Counsel for infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3-days in the office per week and offer relocation assistance to new employees. In this role, you will: Own commercial legal strategy and risk management for OpenAI infrastructure transactions. Draft, negotiate, and advise on complex agreements with infrastructure suppliers, manufacturers, distributors, and technology partners. Support strategic partnerships involving AI infrastructure and hardware supply chains. Develop frameworks for procurement, licensing, and collaboration across the infrastructure ecosystem. Partner with finance and operations teams to align contract terms with business and compliance requirements. Collaborate with policy and regulatory colleagues on issues impacting global supply chains, export controls, and manufacturing. Build scalable, efficient contracting process

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next). We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning. The team’s work spans four areas. First, execution and science: working with teams across OpenAI to decide what can go into the final model and how, using scientific experiments and evals that are representative of the final pipeline so issues can be recognized early. Second, RL scaling: executing the final large-scale reinforcement learning run, making sure GPUs are used efficiently and training stays healthy. Third, research: improving horizontal capabilities like instruction following, factuality, memory, and multi-agent behavior, where the team’s broad visibility helps identify cross-cutting improvements across teams and domains. Fourth, engineering: maintaining the infrastructure stack and internal tools to ensure that both the final run and all integrations go as smoothly as possible and that the systems are easy to work with. About the Role This role focuses on keeping our frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they emerge, from scaling and orchestration issues to inference bottlenecks, numerical problems, and hardware failures, as well as supporting large horizontal integrations in the big run, like multi-agent capabilities or memory. This is a role for a strong generalist who quickly learns anything needed for the task, has high attention to detail, debugs deeply, and is motivated by fixing the highest-impact problem in front of the team. In this role, you will: Keep large-scale async RL training runs moving by jumping into the most urgent engineering and infrastructure problems. Debug issues across training systems, inference, orchestration, scaling, and distributed infrastructure. Improve the reliability and efficiency of RL trai

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man

PythonSQLAWSGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need

AWSAzureRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

PythonAWSLinuxRest
🔔

Get new hardware lead engineer jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime