At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer, Capacity Engineering At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic, fast-moving environments and approach challenges with an experimental mindset, rapidly testing emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake’s infrastructure is expanding rapidly across AWS, Azure, and GCP. The Capacity team plays a pivotal role in provisioning the cloud resources essential for Snowflake's operations and ongoing growth. Capacity Engineering accurately models demand, forecasts requirements, and delivers optimal CPU and GPU capacity on schedule. We drive hardware cost-efficiency and price/performance while continually maximizing fleet utilization. To achieve this across all major cloud providers, the team is building a centralized, self-se
Jobs in United States
Hardware Lead Engineer in United States
400 active opportunities · Updated October 2026
Showing
15 jobs
Explore current hardware lead engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. Product at Baseten Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the first people who will help define it. You'll work directly with our founders and with some of the best systems and infrastructure engineers in the world, and you'll set the standard for building great AI Infrastructure. PMs at Baseten don't sit above engineers - you earn ownership by being technical, finding the truth in front of customers, building great cross-functional relationships, and shipping great product experiences. The role Getting a model into production still takes real expertise — choosing a serving engine, sizing hardware, tuning it, wiring it into an app. We want a developer to go from "it runs on my laptop" to "it's serving production traffic" in minutes, on their own. You'll own the entire experience a developer touches to deploy and iterate: the CLI and SDKs, the console, onboarding, model discovery, deployment configuration, truss, and the increasingly agent-driven ways developers build. Your job is to make Baseten synonymous with Great DevEx and make it effortless to drive and self-serve deploy models on Baseten for far more developers than it is today. Impact and outcomes you'll drive You will collapse time-to-production — take a developer from first sign-up to a running, maint
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are looking for an IT Support / Operations Engineer to join Baseten as we continue to scale our IT team. In this role, you will play a critical part in bringing our technical support entirely in-house to provide a seamless, high-touch experience for all Baseten employees. As we continue to scale, you will be the primary point of contact for day-to-day technical issues, allowing you to have a direct impact on our team's productivity and overall office environment. This position is ideal for a hands-on problem solver who enjoys a mix of hardware and software troubleshooting, user lifecycle management, and maintaining the physical IT infrastructure of a modern office. While you will focus heavily on elevating our internal support standards, you will also assist with systems administration and workflow automation as our company evolves. This is a hybrid role based out of our San Francisco or New York office, following our standard policy of three days per week in-person to ensure our physical office and AV systems remain high-performing and reliable. RESPONSIBILITIES Serve as the escalation point for day-to-day technical support, diagnosing and resolving hardware and software issues across our Mac and Windows fleet Manage user lifecycle administration including provisioning, deprovisioning, and access management across all systems and services Own the IT onboarding experience for new employees — from laptop set
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early member of the Observability Team, you will be pivotal in building and shaping the observability experience for our internal and external customers. By joining this team, you’ll have a direct impact on the reliability and operational excellence of Basetens product systems. As Baseten scales its infrastructure across different cloud providers and diverse hardware, the volume and complexity of operational data is growing by orders of magnitude. This team is responsible for building high-throughput ingest pipelines, cost-efficient storage, and agentic diagnostic tools to ensure that we can detect, diagnose, and resolve issues in minutes rather than hours, even as the systems they operate become more complex. RESPONSIBILITIES Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces across Baseten’s multi-cloud infrastructure Own and evolve core observability platforms, driving migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth Build instrumentation libraries, SDKs, and integrations that make it easy for engineering teams to emit high-quality telemetry from their services Drive alerting and SLO infrastructure that enables teams to define, monitor, and respond to reliabi
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Baseten’s Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost‑efficient. As part of this team, you’ll focus on Model APIs — the infrastructure powering our hosted API endpoints for the latest open‑source models. This work spans distributed systems, model serving, and developer experience. You’ll join a small, high‑impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale. RESPONSIBILITIES: Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations Productionize performa
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Large Language Models (LLMs) continue to push the boundaries of what AI systems can do — but inference is still the bottleneck. The Model Efficiency team is responsible for pushing the limits of LLM inference efficiency across our foundation models. We explore and ship breakthroughs across the model execution stack, including: model architecture and MoE routing optimization decoding and inference-time algorithm improvements software/hardware co-design for GPU acceleration performance optimization without compromising model quality Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. As a Staff Research Engineer, you will develop, prototype, and deploy techniques that materially improve how fast and efficiently our models run in production. You may be a good fit
What you’ll do Execute weekly system-level exploratory testing across the scanner and supporting software; log and triage issues with clear reproduction steps. Work with engineering to debug root cause and validate fixes. Help maintain the DHF and traceability between user needs, design requirements, tests, and results. Own practical test execution logistics (fixtures, test data, environments, calibration artifacts) and keep things repeatable. Help build the continuous testing strategy: automated tests where feasible, plus structured manual and system tests. Support V&V activities, including coordination with external partners as needed. What we’re looking for Strong hands-on testing instincts for complex electromechanical systems with substantial software. Ability to write clear bug reports and communicate risk/impact. Experience building and maintaining test plans/protocols; comfort operating lab equipment and debugging across layers. Useful experience Experience testing complex systems end-to-end (automation where it pays off, plus hands-on hardware/instrumentation). Medical device or other safety-critical environments and comfort translating risk into practical test coverage.
What you’ll do Be the generalist EE for the scanner system: integration, bring-up, debugging, and making the electrical side of the device reliable and serviceable. Own ultrasound experimentations that feeds the image reconstruction team Design and execute experiment setups for transducer characterization (element sensitivity, bandwidth, cross-talk mapping, beam profile measurements) and ex vivo / phantom clinical testing. Acquire, process, and analyze RF and baseband signals for data quality assessment and benchmarking. Design simple boards and adapters as needed (monitoring, power/safety, interface/conditioning), and take them from prototype through a stable revision. Prototype quickly, then harden what works: wiring/harnessing, grounding, safety interlocks, and reliable integration across subsystems. Own practical test setups and documentation (fixtures, scripts, procedures) that make experiments repeatable and results comparable over time. What we’re looking for Strong hands-on EE background with experience building, debugging, and iterating on real systems in the lab. Solid understanding of signal processing fundamentals — knows what to measure, how to condition and digitize it, and how to evaluate signal quality in the context of an imaging system (SNR, bandwidth, dynamic range, artifacts). Comfortable spanning system integration + occasional design work (schematics/layout reviews or light PCB design) in a fast-moving environment. Ability to work at the boundary between hardware and algorithms: measure reality, communicate constraints, and help close gaps vs simulation. High agency and practicality: able to set up experiments, get trustworthy data, and unblock others on a lean team. Useful experience Analog/mixed-signal, or high-speed data capture experience; strong instincts for instrumentation and noise/debugging. Ultrasound or acoustic sensor handling: hydrophone calibration and field mapping, transducer impedance characterization, element-level sensitivity
What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Account Security (AccSec) pod sits within the Accounts team (part of the broader Safety organization) and is at the heart of building a safe and civil platform for users of all ages. Account Security is especially challenging at Roblox because it is open to all ages. As the EM of AccSec pod (7+ ICs) you will work closely with peers throughout Roblox and the creator community to reduce the impact of account takeover (i.e. compromise) for Roblox players, creators, and the broader community. This is a role that requires entrepreneurship. We are looking for the EM of the AccSec pod to, in collaboration with product partners, to formulate a strategy for how to reduce ATO, reduce the impact of any remaining ATO and build community trust in the security of their accounts. Examples of past and current efforts include: Cryptographically binding user secrets to hardware backed secrets. Detecting client side tampering by Browser extensions. Integration with Passkeys. Post-login modeling for ATO detection. Reducing incentives for bad actors by making assets more recoverable. While this is a security team that focuses heavily on product and infrastructure changes to improve security, we also c
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Orchestration pod within Engine Productivity, you'll design and run the platform that executes large-scale end-to-end and integration tests, running the real, shipping client on real hardware, across Roblox's data centers, cloud, and our own device labs, so our engineering teams can ship the engine, clients, Studio, and more with speed and confidence. Every Roblox engine, client, and Studio change, along with the experiences built on top of them, should ship with confidence, and the Orchestration team is the layer that makes that possible. We build large-scale distributed services that turn thousands of test suites into a reliable, push-button pipeline: fanning work out across fleets of machines and real devices, moving artifacts to where they're needed, managing single- and multi-client test state, and giving test owners and maintainers a system to validate their own runs. It looks a lot like building a specialized cloud platform, with capacity-aware scheduling, isolation and sandboxing, and smart retry and backoff, plus the classic distributed systems problems (fairness, efficiency, failure handling, and reliability) at Roblox scale. You Will: Design a
About the Team OpenAI’s Legal team plays a crucial role in advancing our mission by tackling innovative and fundamental legal issues in AI. The team includes professionals from diverse legal fields—technology, AI, infrastructure, privacy, IP, corporate, employment, tax, regulatory, and litigation—who collaborate closely with colleagues across the company. If you are passionate about being a technology lawyer working on cutting-edge challenges, you’ll thrive here. About the Role We’re seeking a senior lawyer to participate in commercial legal strategy and execution across OpenAI’s fast-growing infrastructure portfolio. This is a cross-functional role that will partner closely with procurement, supply chain, partnerships, finance, and product teams to structure, negotiate, and manage the transactions that will support OpenAI’s long-term infrastructure ambitions. We’re looking for an experienced infrastructure transactions lawyer who thrives in ambiguity and wants to help define the commercial playbook for infrastructure efforts in the AI era. This role reports to the Associate General Counsel for infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3-days in the office per week and offer relocation assistance to new employees. In this role, you will: Own commercial legal strategy and risk management for OpenAI infrastructure transactions. Draft, negotiate, and advise on complex agreements with infrastructure suppliers, manufacturers, distributors, and technology partners. Support strategic partnerships involving AI infrastructure and hardware supply chains. Develop frameworks for procurement, licensing, and collaboration across the infrastructure ecosystem. Partner with finance and operations teams to align contract terms with business and compliance requirements. Collaborate with policy and regulatory colleagues on issues impacting global supply chains, export controls, and manufacturing. Build scalable, efficient contracting process
About the Team The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next). We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning. The team’s work spans four areas. First, execution and science: working with teams across OpenAI to decide what can go into the final model and how, using scientific experiments and evals that are representative of the final pipeline so issues can be recognized early. Second, RL scaling: executing the final large-scale reinforcement learning run, making sure GPUs are used efficiently and training stays healthy. Third, research: improving horizontal capabilities like instruction following, factuality, memory, and multi-agent behavior, where the team’s broad visibility helps identify cross-cutting improvements across teams and domains. Fourth, engineering: maintaining the infrastructure stack and internal tools to ensure that both the final run and all integrations go as smoothly as possible and that the systems are easy to work with. About the Role This role focuses on keeping our frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they emerge, from scaling and orchestration issues to inference bottlenecks, numerical problems, and hardware failures, as well as supporting large horizontal integrations in the big run, like multi-agent capabilities or memory. This is a role for a strong generalist who quickly learns anything needed for the task, has high attention to detail, debugs deeply, and is motivated by fixing the highest-impact problem in front of the team. In this role, you will: Keep large-scale async RL training runs moving by jumping into the most urgent engineering and infrastructure problems. Debug issues across training systems, inference, orchestration, scaling, and distributed infrastructure. Improve the reliability and efficiency of RL trai
About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man
About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need
Other cities to consider
More places hiring for this role
Get new hardware lead engineer jobs in United States by email
Daily job updates · Unsubscribe anytime