Jobiba hiring network

Hardware Lead Engineer Jobs

1,301 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current hardware lead engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About BlockTech BlockTech is a fast-paced algorithmic trading firm facilitating global cryptocurrency derivatives and spot trading while expanding into new markets. As we continue to grow rapidly, we are looking for a Software Engineer to join our Foundation team amid our exciting scale-up phase! You will Build & optimize: Design, develop, and maintain high-reliability, low-latency, and high-throughput foundational systems that enable our trading and technology teams to scale efficiently. Ingest & aggregate: Collect trading business data with minimal latency impact and ingest both public and private exchange information into our trading system. Store & stream: Develop and maintain infrastructure for real-time data aggregation and long-term storage, as well as our Kafka-based messaging systems. Collaborate & support: Work closely with multiple teams, assisting them in integrating with and making the most of our foundational systems. Innovate: Drive projects from concept to deployment with full ownership, and explore new tools, frameworks, and approaches to keep our infrastructure best-in-class. The Foundation team develops core software infrastructure (libraries, frameworks, and systems) for BlockTech, solving common problems and lending its expertise to enable other teams to stay focused on their respective domains. They own, develop, and configure a wide variety of critical, high-reliability software, ranging from low-level ultra-low-latency shared memory IPC libraries to high-throughput data buses and data aggregation systems including Kafka, NATS, PostgreSQL, and Iceberg. They work primarily in Rust, but also use Python and SQL. If you thrive on low-level problem-solving, building robust frameworks from scratch, and enabling others to move faster, this role is for you. What We're Looking For Essential: 5+ years of experience as a Software Engineer, with a strong focus on systems-level optimisation and awareness of hardware constraints Proficiency

pythonsqlpostgresql
View job →
D
Dillards
📍 Little Rock• Full-time
1mo ago

THE OPPORTUNITY We're looking for a Windows Server Administrator to join our team in Little Rock, Arkansas, to help plan, design, build, and support solutions for one of the nation’s largest retail companies. Our ideal candidate enjoys working in a fast-paced, hands-on environment supporting critical infrastructure. In this role, you’ll manage, secure, and optimize Windows-based systems across physical and virtual environments. You’ll have the opportunity to work with various technologies, contribute to modernization efforts, and help strengthen our disaster recovery and automation processes. The ideal candidate is proactive, dependable, and ready to take ownership of their work while collaborating with others. This role offers room to grow, whether through learning new tools, streamlining daily operations, or leading technical initiatives that make a real impact. THE TEAM The Windows Server Team is a close-knit, collaborative group responsible for managing and administering all company Windows servers and domains. We work across multiple IT areas to ensure stability, security, and efficiency. Our team thrives on helping each other, sharing knowledge, and finding more innovative ways to do the job. From hardware setup and server virtualization to Active Directory management and vulnerability remediation, you’ll take ownership of our entire Windows ecosystem. This is a highly collaborative, hands-on role where you will maintain and secure the full stack that keeps our operations running. WHAT YOU WILL DO Configure, deploy, manage, and secure complex Windows Server environments. Collaborate across teams to plan and maintain enterprise-level server architectures. Automate routine tasks to improve reliability and reduce manual work. Participate in disaster recovery planning and testing. Support physical and virtual infrastructure, including patching, upgrades, and migrations. Stay current with Microsoft technologies and industry trends. Docum

B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. Product at Baseten Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the first people who will help define it. You'll work directly with our founders and with some of the best systems and infrastructure engineers in the world, and you'll set the standard for building great AI Infrastructure. PMs at Baseten don't sit above engineers - you earn ownership by being technical, finding the truth in front of customers, building great cross-functional relationships, and shipping great product experiences. The role Getting a model into production still takes real expertise — choosing a serving engine, sizing hardware, tuning it, wiring it into an app. We want a developer to go from "it runs on my laptop" to "it's serving production traffic" in minutes, on their own. You'll own the entire experience a developer touches to deploy and iterate: the CLI and SDKs, the console, onboarding, model discovery, deployment configuration, truss, and the increasingly agent-driven ways developers build. Your job is to make Baseten synonymous with Great DevEx and make it effortless to drive and self-serve deploy models on Baseten for far more developers than it is today. Impact and outcomes you'll drive You will collapse time-to-production — take a developer from first sign-up to a running, maint

machine learningaigo
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are looking for an IT Support / Operations Engineer to join Baseten as we continue to scale our IT team. In this role, you will play a critical part in bringing our technical support entirely in-house to provide a seamless, high-touch experience for all Baseten employees. As we continue to scale, you will be the primary point of contact for day-to-day technical issues, allowing you to have a direct impact on our team's productivity and overall office environment. This position is ideal for a hands-on problem solver who enjoys a mix of hardware and software troubleshooting, user lifecycle management, and maintaining the physical IT infrastructure of a modern office. While you will focus heavily on elevating our internal support standards, you will also assist with systems administration and workflow automation as our company evolves. This is a hybrid role based out of our San Francisco or New York office, following our standard policy of three days per week in-person to ensure our physical office and AV systems remain high-performing and reliable. RESPONSIBILITIES Serve as the escalation point for day-to-day technical support, diagnosing and resolving hardware and software issues across our Mac and Windows fleet Manage user lifecycle administration including provisioning, deprovisioning, and access management across all systems and services Own the IT onboarding experience for new employees — from laptop set

machine learningaigo
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early member of the Observability Team, you will be pivotal in building and shaping the observability experience for our internal and external customers. By joining this team, you’ll have a direct impact on the reliability and operational excellence of Basetens product systems. As Baseten scales its infrastructure across different cloud providers and diverse hardware, the volume and complexity of operational data is growing by orders of magnitude. This team is responsible for building high-throughput ingest pipelines, cost-efficient storage, and agentic diagnostic tools to ensure that we can detect, diagnose, and resolve issues in minutes rather than hours, even as the systems they operate become more complex. RESPONSIBILITIES Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces across Baseten’s multi-cloud infrastructure Own and evolve core observability platforms, driving migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth Build instrumentation libraries, SDKs, and integrations that make it easy for engineering teams to emit high-quality telemetry from their services Drive alerting and SLO infrastructure that enables teams to define, monitor, and respond to reliabi

pythonrestmachine learning
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Baseten’s Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost‑efficient. As part of this team, you’ll focus on Model APIs — the infrastructure powering our hosted API endpoints for the latest open‑source models. This work spans distributed systems, model serving, and developer experience. You’ll join a small, high‑impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale. RESPONSIBILITIES: Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations Productionize performa

kubernetesmachine learningai
View job →
C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Large Language Models (LLMs) continue to push the boundaries of what AI systems can do — but inference is still the bottleneck. The Model Efficiency team is responsible for pushing the limits of LLM inference efficiency across our foundation models. We explore and ship breakthroughs across the model execution stack, including: model architecture and MoE routing optimization decoding and inference-time algorithm improvements software/hardware co-design for GPU acceleration performance optimization without compromising model quality Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. As a Staff Research Engineer, you will develop, prototype, and deploy techniques that materially improve how fast and efficiently our models run in production. You may be a good fit

gitrestmachine learning
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! We’re looking for a senior engineer to help build, maintain and evolve the training framework that powers our frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure. You will design and maintain the core components that enable fast, reliable, and scalable model training — and build the tooling that connects research ideas to thousands of GPUs. If you enjoy working across the full stack of ML systems, this role gives you the opportunity and autonomy to have massive impact. What You’ll Work On Build and own the training framework responsible for large-scale LLM training. Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing). Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100). Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics. Collaborate closely with infra teams to ensure our cluster, container environments, and hardware configurations support high-performance training. Investigate and res

dockerkubernetesgit
View job →
M
1mo ago

What you’ll do Execute weekly system-level exploratory testing across the scanner and supporting software; log and triage issues with clear reproduction steps. Work with engineering to debug root cause and validate fixes. Help maintain the DHF and traceability between user needs, design requirements, tests, and results. Own practical test execution logistics (fixtures, test data, environments, calibration artifacts) and keep things repeatable. Help build the continuous testing strategy: automated tests where feasible, plus structured manual and system tests. Support V&V activities, including coordination with external partners as needed. What we’re looking for Strong hands-on testing instincts for complex electromechanical systems with substantial software. Ability to write clear bug reports and communicate risk/impact. Experience building and maintaining test plans/protocols; comfort operating lab equipment and debugging across layers. Useful experience Experience testing complex systems end-to-end (automation where it pays off, plus hands-on hardware/instrumentation). Medical device or other safety-critical environments and comfort translating risk into practical test coverage.

What you’ll do Be the generalist EE for the scanner system: integration, bring-up, debugging, and making the electrical side of the device reliable and serviceable. Own ultrasound experimentations that feeds the image reconstruction team Design and execute experiment setups for transducer characterization (element sensitivity, bandwidth, cross-talk mapping, beam profile measurements) and ex vivo / phantom clinical testing. Acquire, process, and analyze RF and baseband signals for data quality assessment and benchmarking. Design simple boards and adapters as needed (monitoring, power/safety, interface/conditioning), and take them from prototype through a stable revision. Prototype quickly, then harden what works: wiring/harnessing, grounding, safety interlocks, and reliable integration across subsystems. Own practical test setups and documentation (fixtures, scripts, procedures) that make experiments repeatable and results comparable over time. What we’re looking for Strong hands-on EE background with experience building, debugging, and iterating on real systems in the lab. Solid understanding of signal processing fundamentals — knows what to measure, how to condition and digitize it, and how to evaluate signal quality in the context of an imaging system (SNR, bandwidth, dynamic range, artifacts). Comfortable spanning system integration + occasional design work (schematics/layout reviews or light PCB design) in a fast-moving environment. Ability to work at the boundary between hardware and algorithms: measure reality, communicate constraints, and help close gaps vs simulation. High agency and practicality: able to set up experiments, get trustworthy data, and unblock others on a lean team. Useful experience Analog/mixed-signal, or high-speed data capture experience; strong instincts for instrumentation and noise/debugging. Ultrasound or acoustic sensor handling: hydrophone calibration and field mapping, transducer impedance characterization, element-level sensitivity

What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.

PE
Private Employer
📍 United Kingdom• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role The Palantir platform is deployed in numerous critical mission environments including combat zones and classified networks—from the back of a Humvee to a command post to the cloud. This means operating in multiple cloud environments, on-prem air-gapped networks, and at the edge—at scale. We are looking for Edge Infrastructure Engineers to build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure. This role demands a deep focus on low-level systems, including the deployment and management of physical bare metal servers in both traditional data centers and edge environments. You will be responsible for physical network engineering and the development of robust infrastructure that ensures performance of the Palantir platform. In addition to ensuring performance and reliability, you will play a critical role in building and scaling new environments in a forward-deployed capacity, including onsite. Edge Infrastructure Engineers combine hardware-level engineering experience with the drive to improve existing systems and the creativity to develop novel solutions for evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate for and participate in sensible, scalable systems design, sharing responsibility for diagnosing, resolving, and preventing production issues across our most demanding deployments.

PE
Private Employer
📍 Washington• Full-time
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role As a Software Developer in Palantir’s Mixed Reality Group (MRG), you will join a Tier 1 team of researchers, 3D artists, immersive UX designers, and cross-discipline engineers to bring the most exquisite software of our time to the factory floor, the battlefield, and the briefing room. You will collaborate with customers, operators, MR designers, and Forward Deployed Engineers to deliver force-multiplying apps that exceed tomorrow’s requirements today. Further, you will work on high-impact, mission-critical flows in both 3D mixed reality and 2D end-user devices while fielding performance-sensitive infrastructure for bespoke next-generation mixed reality hardware, sensors, and data platforms.

PE
Private Employer
📍 Seattle• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Apollo is Palantir’s autonomous software management and deployment platform. It enables seamless, continuous delivery of mission-critical software (Foundry, Gotham, AIP) across a vast range of environments: on-prem, public cloud, disconnected (air-gapped) networks, and highly regulated settings (including IL-5 and FedRAMP). As a Software Engineer on the Apollo team, you’ll build and operate a large-scale distributed system to allow the remote operation and maintenance of Kubernetes clusters. Our mission is to extract the entire state of a cluster into a portable, high-performance artifact within minutes, enabling full and almost instant cluster reconstruction from the ground up—all while pushing the limits of speed, reliability, and scale. You’ll design and implement backup and restore solutions for Kubernetes, leveraging proprietary compression infrastructure tailored to Palantir’s unique deployment models. You’ll also build and optimize our container artifact store, which is based on the OCI (Open Container Initiative) distribution spec—the industry standard for storing and distributing container images and artifacts. You’ll own the backbone of every environment Apollo supports, from hyperscalers to Army trucks. If you’re excited by challenges at the intersection of container technologies like OCI and docker, storage, and distributed systems, you’ll find opportunities here to dive deep into storage formats and low-level optimizations, where milliseconds matter. As we increasingly automate cluster creation and management on diverse hardware, you’ll play a key role in scaling Palantir’s presence at the edge and solving tough distributed systems proble

dockerkubernetesrest
View job →
PE
Private Employer
📍 New York• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Apollo is Palantir’s autonomous software management and deployment platform. It enables seamless, continuous delivery of mission-critical software (Foundry, Gotham, AIP) across a vast range of environments: on-prem, public cloud, disconnected (air-gapped) networks, and highly regulated settings (including IL-5 and FedRAMP). As a Software Engineer on the Apollo team, you’ll build and operate a large-scale distributed system to allow the remote operation and maintenance of Kubernetes clusters. Our mission is to extract the entire state of a cluster into a portable, high-performance artifact within minutes, enabling full and almost instant cluster reconstruction from the ground up—all while pushing the limits of speed, reliability, and scale. You’ll design and implement backup and restore solutions for Kubernetes, leveraging proprietary compression infrastructure tailored to Palantir’s unique deployment models. You’ll also build and optimize our container artifact store, which is based on the OCI (Open Container Initiative) distribution spec—the industry standard for storing and distributing container images and artifacts. You’ll own the backbone of every environment Apollo supports, from hyperscalers to Army trucks. If you’re excited by challenges at the intersection of container technologies like OCI and docker, storage, and distributed systems, you’ll find opportunities here to dive deep into storage formats and low-level optimizations, where milliseconds matter. As we increasingly automate cluster creation and management on diverse hardware, you’ll play a key role in scaling Palantir’s presence at the edge and solving tough distributed systems proble

dockerkubernetesrest
View job →
🔔

Get new hardware lead engineer jobs by email

Daily job updates · Unsubscribe anytime