Jobs in United States

Hardware Operations Engineer in San Francisco

184 active opportunities · Updated October 2026

Explore current hardware operations engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

We're looking for a Senior Mechanical Engineer for Midjourney Medical — from precision electromechanical assemblies up to the large-scale structures and mechanisms that hold everything together and make it move. This role spans the full range of scale: one week you might be refining a compact transducer mount, the next you're architecting a structural frame or designing the motion system that positions hardware within it. This is a hands-on role on a small cross-functional team. You'll take problems from whiteboard sketch to working hardware yourself — designing in CAD, prototyping in the shop, testing, iterating, and integrating with research teams, electrical and software engineers along the way. We move fast and expect you to drive your own work: identifying what needs to happen next, making sound engineering calls without waiting for permission, and shipping hardware that works. What you'll do Own parts of the mechanical systems end to end — structures, mechanisms, and electromechanical integration Design large-scale structures: frames, weldments, enclosures, and support systems, with attention to stiffness, weight, manufacturability, and serviceability Design mechanisms: linkages, motion stages, actuation systems for precise, reliable positioning and articulation Integrate transducers, electronics, cabling, and thermal management into large physical systems Perform engineering analysis (tolerance stack-ups, structural/FEA, mechanism kinematics) to de-risk designs before committing to hardware Build and test relentlessly — 3D printing, rapid prototyping, and hands-on fabrication are core to how you'll validate designs Select parts and materials for system-level integration, balancing performance, cost, and lead time Drive projects independently from concept through validation, and collaborate closely with a small cross-disciplinary team to hit project goals What we're looking for Bachelor's degree in Mechanical Engineering or a related field 5+ years of experien

AIGoExcelSEM
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -86.4%

About the Team The Future of Computing Research team is an applied research team within OpenAI’s Consumer Devices group. We study how AI systems perceive people and their surroundings, and we turn that research into capabilities for future products. Our work spans machine learning, sensing, and hardware, with a focus on building systems that work beyond controlled environments. About the Role We’re looking for a machine learning engineer to help shape how future AI systems understand the physical world and the people in it. The role focuses on multimodal perception and authentication, bringing together signals from cameras, microphones, and other sensors. You’ll work with specialized perception models and larger multimodal models, and partner with hardware, firmware, software, and product teams to bring new research into real-world systems. This role is based in San Francisco. We work in the office three days per week and offer relocation assistance. In this role, you will: Research and develop multimodal perception and authentication methods across visual, audio, and other sensing signals. Explore how specialized perception models and larger multimodal models can work together. Design data, training, and evaluation approaches that improve performance in real-world conditions. Study model behavior, robustness, and failure modes across sensing, data, and deployment environments. Integrate and validate new capabilities in real-time or resource-constrained systems. Work with hardware, firmware, software, and product teams to turn research into working systems. You might thrive in this role if you: Have a strong background in computer vision, audio or speech machine learning, multimodal learning, or sensing. Have experience developing specialized machine learning models, larger multimodal models, or both. Have brought research ideas into practical systems, prototypes, or products. Know how to design experiments, build evaluations, and investigate model behavior. Have wo

PythonAWSRestMachine Learning
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -86.4%

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b

TypeScriptPythonAWSDocker
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -86.4%

About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation

AWSRestAIGo
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -86.4%

This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea

AWSAzureKubernetesCI/CD
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -86.4%

About the Team We are building general-purpose robotics. In the short term, we are focused on robots to support skilled workers to build our future infrastructure. In the long term, we imagine everyone having a personal robot doing anything they need. Progress is rapid, and based on a foundation of co-design between robotics hardware and ML research. About the Role As a Firmware Engineer, you will define and drive the architecture of embedded systems for next-generation hardware products. You will own foundational firmware decisions across real-time execution, device bring-up, hardware interfaces, fault handling, safety mechanisms, and production readiness. We’re looking for someone with deep experience building safety-critical or high-consequence systems, where failures can have meaningful consequences. You should be comfortable reasoning about risk, designing for diagnosability and graceful degradation, and creating engineering practices that raise the reliability bar for the entire team. You should also be unusually good at moving fast. Sometimes the right answer is a carefully reviewed architecture that will endure for years; sometimes it is getting a rough-but-useful prototype working by the end of the afternoon so the team can learn something concrete tomorrow. We value engineers who know the difference, make that call well, and can operate credibly in both modes. You will be both a technical leader and a hands-on builder: setting direction, reviewing critical designs, unblocking the hardest problems, and writing production firmware when it matters most. Our embedded stack uses a lot of Rust. Extensive experience in the language is a big help! This role is based in San Francisco, CA. This role will be expected to be in office 4 days per week and offer relocation assistance to new employees. In this role, you will: Rapidly bring up new hardware and set execution pace for the team. Lead firmware architecture for embedded systems spanning boot, RTOS/runtime behav

AWSRestAIC++
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -86.4%

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software is reliable, testable, and ready to ship. We design and maintain automated test frameworks, hardware-in-the-loop labs, and release pipelines that keep quality signals trustworthy and enable rapid, safe product launches. Our work spans developer tools, automation, systems integration, and cross-team collaboration to ensure every release meets the highest standards. About the Role As a Software Engineer, Quality and Developer Tools , you will build and own the systems that validate our device software—from test frameworks and regression infrastructure to hardware-in-the-loop labs and release gates. You’ll design the tooling and automation that keep quality signals trustworthy, integrate them into CI/CD, and make it easy for engineers and QA vendor technicians to execute reliable, repeatable workflows. We’re looking for engineers with deep experience in software quality, automation, developer tooling, and hardware-software integration who thrive on building scalable, reliable systems for validation and release readiness. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Test infrastructure & frameworks: Design, implement, and maintain a unified test framework for device software across unit, integration, system, and end-to-end testing, with reproducible runs and integrations with GitHub, Linear, and Slack. CI/CD integration & releases: Integrate test suites with Buildkite, enforce promotion criteria for staging and production, auto-file regressions, and publish traceable artifacts and release notes. Hardware-in-the-loop lab design & orchestration: Plan and bring up racks, power and networking systems, and orchestration for device testing; support automated flashing, provisioning

PythonAWSCI/CDGit
M
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -67.9%
Quick readStrong listing-quality and freshness signals

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll automate the integration of new capacity from a growing set of hardware providers; from auditing and benchmarking hosts and clusters, to maintaining our machine images, configuring GPUs, RDMA, networking, and storage, and getting machines into production. You'll build the automation that keeps the fleet healthy without human intervention: detecting bad GPUs, thermals, and disks. You'll dig into whatever is between the hardware and the software that runs on

PythonLinuxAIAuditing
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -86.4%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s mission is to ensure that general-purpose artificial intelligence benefits all of humanity. The Design Systems team creates the foundations that help teams across OpenAI build coherent, high-quality experiences at scale. Accessibility is fundamental to that work. We’re building the shared standards, components, tools, and practices that make inclusive experiences the default across OpenAI’s products and platforms. About the Role In this role, you’ll lead OpenAI’s company-wide accessibility program from within Design Foundations. You’ll pair strategy and organizational influence with hands-on product design—advocating for disabled people, removing barriers across complete customer journeys, and establishing durable ownership throughout the company. This is an opportunity to define accessibility for a new generation of products spanning web, mobile, desktop, voice, AI, and hardware. You’ll help teams consider the many ways people perceive, understand, navigate, and interact with technology, going beyond minimum compliance to create experiences that are genuinely useful and inclusive. Working across design, engineering, research, product, go-to-market, and legal, you’ll make accessibility an expected part of product quality—not a final review or the responsibility of a single expert. This role is based in our San Francisco HQ. We offer relocation assistance to new employees. In this role you will: Define OpenAI’s accessibility strategy, roadmap, governance, priorities, and measures of progress. Design and ship accessible experiences while identifying systemic barriers across complete customer journeys. Build accessibility into our design system through inclusive components, tokens, interaction patterns, Figma libraries, documentation, and engineering guardrails. Center disabled people through participatory research, usability studies, and assistive-technology testing rather than relying solely on assumptions, simulations, or automated tools. Des

AWSRestAIGo
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an

KubernetesGitAIGo
M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Be the generalist EE for the scanner system: integration, bring-up, debugging, and making the electrical side of the device reliable and serviceable. Own ultrasound experimentations that feeds the image reconstruction team Design and execute experiment setups for transducer characterization (element sensitivity, bandwidth, cross-talk mapping, beam profile measurements) and ex vivo / phantom clinical testing. Acquire, process, and analyze RF and baseband signals for data quality assessment and benchmarking. Design simple boards and adapters as needed (monitoring, power/safety, interface/conditioning), and take them from prototype through a stable revision. Prototype quickly, then harden what works: wiring/harnessing, grounding, safety interlocks, and reliable integration across subsystems. Own practical test setups and documentation (fixtures, scripts, procedures) that make experiments repeatable and results comparable over time. What we’re looking for Strong hands-on EE background with experience building, debugging, and iterating on real systems in the lab. Solid understanding of signal processing fundamentals — knows what to measure, how to condition and digitize it, and how to evaluate signal quality in the context of an imaging system (SNR, bandwidth, dynamic range, artifacts). Comfortable spanning system integration + occasional design work (schematics/layout reviews or light PCB design) in a fast-moving environment. Ability to work at the boundary between hardware and algorithms: measure reality, communicate constraints, and help close gaps vs simulation. High agency and practicality: able to set up experiments, get trustworthy data, and unblock others on a lean team. Useful experience Analog/mixed-signal, or high-speed data capture experience; strong instincts for instrumentation and noise/debugging. Ultrasound or acoustic sensor handling: hydrophone calibration and field mapping, transducer impedance characterization, element-level sensitivity

GitAIGoRust
M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Act as the in-house electrical lead for Midjourney Medical: own the electrical architecture of the scanner and the technical direction for all board-level design. Own complex board design end-to-end: architecture, schematic capture, layout (high-speed digital, analog/mixed-signal, power), DFM/DFT, fabrication and assembly vendor management, bring-up, and revision control. Write firmware for embedded targets (MCU/SoC): drivers, real-time control loops, safety-relevant logic, bootloaders, and field update paths. Audit and update HDL (FPGA) code for high-throughput data acquisition, timing/synchronization, triggering, and pre-processing of ultrasound and sensor data streams. Define electrical interfaces and data contracts with software, recon/ML and mechanical teams: timing budgets, clocking/sync, signal integrity, connectors/harnessing, and failure modes. Establish electrical engineering rigor: design reviews, schematic/layout review checklists, bring-up procedures, test fixtures, and documentation suitable for a regulated medical device program (DHF, traceability, change control). Mentor and grow the electrical function; select and manage external design partners where leverage is high. What we’re looking for Deep experience designing complex boards from blank page to stable revision, including high-speed digital and analog/mixed-signal domains. Strong schematic and layout skills (Altium/KiCad or equivalent) with real signal integrity, power integrity, grounding, and EMI/EMC instincts. Solid embedded firmware background in C/C++ (and Python for tooling): peripherals, DMA, interrupts, real-time constraints, and debugging on hardware. Practical HDL experience (VHDL/Verilog/SystemVerilog) for data acquisition, timing, and streaming interfaces. Track record of owning bring-up and debug on real hardware: scopes, logic analyzers, and disciplined root-cause analysis. Technical leadership: clear trade-offs, strong written documentation, and the ability to set

PythonGitAIC++
F
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -87.7%

From $350K/yr

Quick readStrong listing-quality and freshness signals

About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. Help Win New Business The Opportunity: We are scaling our dedicated Data Center practice, and we are looking for the person who will lead it. This is a founding commercial role. You will start as a team of one, owning the full sales motion end-to-end, and you will build the team around you as the practice grows. You will define how Flexport goes to market with hyperscalers, hardware OEMs, and the broader ecosystem. You will set the playbook, win the first marquee accounts, and hire the people who scale what you build. Reporting directly to the Regional General Manager, you will operate with a high degree of autonomy and direct access to executive leadership. This role features a 50/50 compensation model (Base + Uncapped bonus, with accelerators), with OTEs of $350k+, designed for strong leaders who take bets on themselves. Why This Role Is Different: The data center logistics market is not a standard freight problem. A single AI compute rack can cost more than $1 million and weigh up to 4,000 pounds. Racks contain Class 9 dangerous goods (lithium-ion batteries) and liquid cooling systems requiring specialized handling. Construction sequencing failures delayed 57% o

GitAIGoRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -86.4%

About the Team The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next). We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning. The team’s work spans four areas. First, execution and science: working with teams across OpenAI to decide what can go into the final model and how, using scientific experiments and evals that are representative of the final pipeline so issues can be recognized early. Second, RL scaling: executing the final large-scale reinforcement learning run, making sure GPUs are used efficiently and training stays healthy. Third, research: improving horizontal capabilities like instruction following, factuality, memory, and multi-agent behavior, where the team’s broad visibility helps identify cross-cutting improvements across teams and domains. Fourth, engineering: maintaining the infrastructure stack and internal tools to ensure that both the final run and all integrations go as smoothly as possible and that the systems are easy to work with. About the Role This role focuses on keeping our frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they emerge, from scaling and orchestration issues to inference bottlenecks, numerical problems, and hardware failures, as well as supporting large horizontal integrations in the big run, like multi-agent capabilities or memory. This is a role for a strong generalist who quickly learns anything needed for the task, has high attention to detail, debugs deeply, and is motivated by fixing the highest-impact problem in front of the team. In this role, you will: Keep large-scale async RL training runs moving by jumping into the most urgent engineering and infrastructure problems. Debug issues across training systems, inference, orchestration, scaling, and distributed infrastructure. Improve the reliability and efficiency of RL trai

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -86.4%

About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need

AWSAzureRestMachine Learning
🔔

Get new hardware operations engineer jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime