About the Team Tax and Trade at OpenAI shapes business strategy by embedding critical tax, export control, customs, and cross-border considerations into how the company builds, sources, scales, and operates in support of the mission. We combine deep expertise with practical systems thinking to look around corners, identify emerging risks and opportunities early, and help teams make smarter decisions at the point where strategy becomes execution. Across procurement, hardware operations, manufacturing, logistics, finance, legal, supplier onboarding, and operator workflows, we build robust, scalable support services leveraging cutting-edge technology—including governed AI and automation—to make complex regulated work more durable, more efficient, and easier to scale. About the Role We’re hiring a Senior Manager, Export Controls to lead OpenAI’s export controls strategy and operating model. This is a senior role with broad scope across advanced computing, semiconductors, software, hardware, manufacturing, and high technology partnerships. You will refine how OpenAI classifies controlled technology, software, and hardware, structures access-controlled environments, manages licensing and supplier commitments, and scales export-control operations in a way that supports the company’s pace of innovation. You will also shape how OpenAI applies AI and agentic workflows to policy-heavy operational work, building systems that make complex rules easier to navigate and easier to execute. In this role, you will: Refine the strategy and operating model for OpenAI’s export controls program across advanced computing, semiconductors, software, hardware, manufacturing, and high technology partnerships. Own export classification and licensing strategy for controlled technical data, software, hardware, and research environments. Lead the design and operation of compliant controlled environments and related governance processes. Partner with Research and Infrastructure to support efficient
Jobs in United States
Hardware Systems Planning Lead in San Francisco
153 active opportunities · Updated September 2026
Showing
15 jobs
Explore current hardware systems planning lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches. About the Role In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs. In this role, you will Build and refine performance models that translate microbenchmark results into cost-to-serve estimates. Analyze inference workloads end to end across applications, models, and fleet infrastructure. Enhance tooling to identify bottlenecks across layers for latency and throughput. Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference. You might thrive in this role if you: Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency. Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling. Have deep expertise with performance profiling, benchmarking, analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve o
What you’ll do Execute weekly system-level exploratory testing across the scanner and supporting software; log and triage issues with clear reproduction steps. Work with engineering to debug root cause and validate fixes. Help maintain the DHF and traceability between user needs, design requirements, tests, and results. Own practical test execution logistics (fixtures, test data, environments, calibration artifacts) and keep things repeatable. Help build the continuous testing strategy: automated tests where feasible, plus structured manual and system tests. Support V&V activities, including coordination with external partners as needed. What we’re looking for Strong hands-on testing instincts for complex electromechanical systems with substantial software. Ability to write clear bug reports and communicate risk/impact. Experience building and maintaining test plans/protocols; comfort operating lab equipment and debugging across layers. Useful experience Experience testing complex systems end-to-end (automation where it pays off, plus hands-on hardware/instrumentation). Medical device or other safety-critical environments and comfort translating risk into practical test coverage.
What you’ll do Act as the technical lead for large parts of the scanner platform: system architecture, codebase structure, and long-term maintainability. Own core runtime foundations: distributed control, state management, fault handling, and reliability. Drive engineering rigor: testability, code quality, review standards, performance regression prevention, and release processes. Build robust observability: logs, metrics, traces, and replayable diagnostics (with privacy constraints). Collaborate with hardware and recon/ML teams to define interfaces, data contracts, timing/synchronization, and failure modes. Lead complex refactors (e.g., message passing / RPC boundaries, modularization, concurrency model) without halting forward progress. What we’re looking for Deep software architecture experience for real-world systems: robotics, instrumentation, medical devices, or other complex distributed products. Strong Python and concurrency background (asyncio, multiprocessing, profiling, performance engineering). Track record of shipping systems that are observable, debuggable, and resilient. Strong technical leadership: clarity, pragmatic trade-offs, and mentoring. Useful experience Building but rock-solid systems: clear interfaces (gRPC/protobuf or equivalent), strong state modeling, and failure handling. High-leverage engineering habits on a lean team: good tests, CI, reproducible dev environments, and fast code review. Practical performance + concurrency work in Python (asyncio, profiling, multiprocessing) and comfort debugging distributed behavior. Security-minded device software: safe defaults, encrypted data paths, and disciplined handling of PII/PHI. Operational thinking: remote updates/management, excellent logging, and diagnostics that make real hardware debuggable.
What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.
About the Team At OpenAI, the User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We operate at the front line of real-world safety and risk management, translating user and operational signals into timely decisions, effective interventions, and improvements to our systems. This role sits on a team focused on building operational capacity for new, ambiguous, and fast-moving company priorities. The team defines what needs to be built, creates the operating model to support it, and works with partner teams to make the work scalable and durable over time. About the Role We are looking for a senior program manager to build the safety, quality, and risk operations supporting a new category of consumer devices. You will translate ambiguous product risks and evolving requirements into practical operating models, workflows, escalation paths, launch-readiness plans, and cross-functional decision-making. This is a foundational role: the systems you build will shape how OpenAI launches, monitors, and improves a new category of consumer devices safely at scale. You will help establish how potential safety incidents, product-quality concerns, sensitive customer escalations, privacy-sensitive issues, and other emerging device risks are identified, investigated, resolved, and incorporated into product and operational improvements. You will turn incomplete requirements into practical workflows, decision rights, launch plans, quality controls, measurement, and durable ownership. The role centers on program building, operational judgment, and execution. We welcome candidates from product safety, quality assurance, regulatory operations, technical program management, healthcare, medical devices, aerospace, consumer technology, and other environments involving complex products or regulated risks. Direct consumer-hardware experience is helpful but not required. The strongest candidates learn unfam
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking to build an investigative capability for Secure Manufacturing & Stealth programs. The risk surface for unreleased products, prototypes, confidential hardware, infrastructure, supply chain, manufacturing, and launch-readiness efforts spans employees, vendors, suppliers, logistics partners, physical movement of assets, procurement records, manufacturing workflows, access systems, device telemetry, and adversarial collection. This role is intended to build and run investigations across that specialized environment. In this role, you will: Lead complex SMS investigations to proactively identify and mitigate risks to unreleased products, prototypes, confidential hardware, secure manufacturing programs, and launch-readiness efforts. Investigate unauthorized disclosure, suspected leaks, insider risk, supplier compromise, vendor misconduct, theft, diversion, tampering, counterfeiting, surveillance, adversarial collection, and suspicious activity involving sensitive programs. Connect digital evidence, physical access activity, supply chain records, manufacturing data, vendor behavior, employee activity, collaboration metadata, procurement records, shipping data, and OSINT into clear findings and risk-reduction actions. Conduct proactive threat hunting to surface early indicators of compromise, collection, leakage, or insider activity affecting sensitive programs. Develop investigative playbooks, evidence-handling standards,
About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Infrastructure team is central to this mission, setting the core strategy and implementing the vision. From site selection to deployment to operations, this team sits at the intersection of commercial, technical, and operational domains, interacting with experts and executives inside and outside of OpenAI. We design and operate mission-critical facilities that support cutting-edge AI workloads at scale. About the Role We are seeking a Facilities Operations Lead to support the commissioning, deployment, and long-term operation of our next-generation AI data centers. This role bridges the interface between data center construction and hardware landing, ensuring seamless integration of mission-critical infrastructure with hardware deployment timelines. You will define and execute commissioning plans, support infrastructure bring-up, and take ownership of operations and maintenance for cutting-edge, large-scale, AI data centers. You will collaborate closely with design, construction, and hardware teams to define repeatable processes for new data center builds and lead hands-on operations to uphold the performance and reliability of our deployed infrastructure. Key Responsibilities Define and execute sequences of operations, commissioning steps, and bring-up processes for mission-critical data center facilities. Interface with the design and hardware teams to define deployment procedures tailored to each data center and hardware configuration. Oversee installation, commissioning, and operational readiness of large-scale data center campuses. Manage monitoring, maintenance, and quality control of the data center infrastructure, including high-performance liquid cooling systems. Develop on-site operations staffing strategy. Develop and enforce procedures for planed and unplanned downtime and SLAs for critical
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des
Coder is looking for an experienced and detail-oriented IT Generalist to join our growing team. This is a full-time role ideal for someone who thrives in a fast-paced, hands-on environment. You’ll be the go-to person for user support, SaaS administration, and IT infrastructure, playing a key role in ensuring our team stays productive, secure, and well-equipped. You'll work closely with the larger IT team to keep things running smoothly - from onboarding new hires to managing devices and licenses to supporting SOC 2 compliance initiatives. If you're a strong communicator with a passion for IT operations and solving real problems for real people, this role is for you. This position follows a hybrid work model. Candidates should be local and able to work from our local office on a regular weekly basis, with scheduling flexibility. What you’ll do here Act as the first line of IT support for employees via Slack and our internal ticketing system Manage user onboarding/offboarding, including account provisioning through Okta and license management via our SaaS management systems Administer and maintain core tools like Google Workspace, Slack, Jamf, 1Password, and other business-critical SaaS apps Set up and manage macOS hardware inventory, including procurement, configuration, and asset tracking Support SOC 2 and GDPR compliance by following processes for access controls, audits, and data security Help scale IT operations as we grow - identify gaps, recommend tooling, and streamline support workflows Document processes and contribute to internal knowledge bases to improve self-service and transparency What we’re looking for Have 2-4 years of experience in IT support or systems administration Are proficient with Okta, Google Workspace, Slack, Zoom, and Apple device management (Jamf preferred) Are comfortable working independently in a remote environment and prioritizing across a variety of IT tasks Have a security-first mindset and understand the importance of access contro
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. Product at Baseten Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the first people who will help define it. You'll work directly with our founders and with some of the best systems and infrastructure engineers in the world, and you'll set the standard for building great AI Infrastructure. PMs at Baseten don't sit above engineers - you earn ownership by being technical, finding the truth in front of customers, building great cross-functional relationships, and shipping great product experiences. The role Getting a model into production still takes real expertise — choosing a serving engine, sizing hardware, tuning it, wiring it into an app. We want a developer to go from "it runs on my laptop" to "it's serving production traffic" in minutes, on their own. You'll own the entire experience a developer touches to deploy and iterate: the CLI and SDKs, the console, onboarding, model discovery, deployment configuration, truss, and the increasingly agent-driven ways developers build. Your job is to make Baseten synonymous with Great DevEx and make it effortless to drive and self-serve deploy models on Baseten for far more developers than it is today. Impact and outcomes you'll drive You will collapse time-to-production — take a developer from first sign-up to a running, maint
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are looking for an IT Support / Operations Engineer to join Baseten as we continue to scale our IT team. In this role, you will play a critical part in bringing our technical support entirely in-house to provide a seamless, high-touch experience for all Baseten employees. As we continue to scale, you will be the primary point of contact for day-to-day technical issues, allowing you to have a direct impact on our team's productivity and overall office environment. This position is ideal for a hands-on problem solver who enjoys a mix of hardware and software troubleshooting, user lifecycle management, and maintaining the physical IT infrastructure of a modern office. While you will focus heavily on elevating our internal support standards, you will also assist with systems administration and workflow automation as our company evolves. This is a hybrid role based out of our San Francisco or New York office, following our standard policy of three days per week in-person to ensure our physical office and AV systems remain high-performing and reliable. RESPONSIBILITIES Serve as the escalation point for day-to-day technical support, diagnosing and resolving hardware and software issues across our Mac and Windows fleet Manage user lifecycle administration including provisioning, deprovisioning, and access management across all systems and services Own the IT onboarding experience for new employees — from laptop set
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early member of the Observability Team, you will be pivotal in building and shaping the observability experience for our internal and external customers. By joining this team, you’ll have a direct impact on the reliability and operational excellence of Basetens product systems. As Baseten scales its infrastructure across different cloud providers and diverse hardware, the volume and complexity of operational data is growing by orders of magnitude. This team is responsible for building high-throughput ingest pipelines, cost-efficient storage, and agentic diagnostic tools to ensure that we can detect, diagnose, and resolve issues in minutes rather than hours, even as the systems they operate become more complex. RESPONSIBILITIES Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces across Baseten’s multi-cloud infrastructure Own and evolve core observability platforms, driving migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth Build instrumentation libraries, SDKs, and integrations that make it easy for engineering teams to emit high-quality telemetry from their services Drive alerting and SLO infrastructure that enables teams to define, monitor, and respond to reliabi
Other cities to consider
More places hiring for this role
Get new hardware systems planning lead jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime