About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
Jobs in United States
Technical Lead in United States
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current technical lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Applied team brings OpenAI’s technology to the world through products used by hundreds of millions of people and by developers and businesses building on our APIs. We work across research, engineering, product, policy, safety, and operations to deploy frontier AI systems responsibly and safely. The Trust & Safety Data Engineering team builds the data foundations that help OpenAI understand, detect, investigate, and mitigate abuse and safety risks across our products. We partner with Integrity, Investigations, Safety Systems, Product Policy, Privacy, Data Science, Engineering, and Data Platform to create reliable, privacy-safe datasets and pipelines for fraud and abuse detection, enforcement workflows, safety measurement, ML feature generation, launch readiness, and transparency reporting. About the Role We are hiring a Technical Lead Manager to lead and grow the Trust & Safety Data Engineering team. This is a hands-on leadership role for someone who can set strategy, shape data architecture, align senior stakeholders, coach engineers, and drive execution on high-impact data systems. You will help turn fragmented launch and incident support into durable, reusable, privacy-safe data foundations that Trust & Safety teams can rely on. The systems your team builds will help OpenAI detect risk, investigate abuse, power operational workflows, develop and evaluate safety models, measure interventions, support product launches, and report accurately on platform integrity. In This Role, You Will Lead and grow a high-performing Trust & Safety Data Engineering team. Define the roadmap and technical strategy for Trust & Safety data systems. Build canonical, privacy-safe datasets and pipelines for abuse detection, fraud detection, risk signals, enforcement, scaled review, transparency reporting, and safety monitoring. Create reusable foundations for Trust & Safety model development, including features, labels, training data, backtesting,
About the Team Training Runtime builds the distributed systems that power OpenAI's largest model training runs - most recently GPT-5.5! The Data Movement area owns the infrastructure that keeps training jobs supplied with the right data at the right time, and keeps model state moving safely and efficiently across large clusters. Our work spans machine learning systems, distributed storage, high-throughput data loading, reliability engineering, and developer experience. Success means researchers can move quickly while training runs remain fast, reproducible, debuggable, and resilient at scale. About the Role We are looking for a deeply hands-on Technical Lead Manager to own datasets throughout our training infrastructure. This person will set the direction for how training jobs read data: the APIs, storage contracts, versioning model, benchmarks, debugging tools, and reliability guarantees that make data access consistent across current and future training frameworks. You will begin as the primary technical owner for dataset reads, working directly in the code while aligning researchers, training framework owners, storage teams, and infrastructure partners around a durable platform. The problem is deceptively hard at frontier scale: make enormous, heterogeneous datasets easy to consume, correct across distributed workers, observable when something goes wrong, and flexible enough to support pretraining, reinforcement learning, and multimodal training. In this role, you will Design and build a unified dataset read platform for multiple current and future training frameworks. Define dataset APIs, storage-format expectations, registration/versioning, and migration paths that make data access reproducible and maintainable. Build reliability into the read path, including stateful iteration, caching, fast restart, recovery, and clear operational contracts. Build terminal and web-based visualizers that let teams inspect text, multimodal, and reinforcement learning data late
About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by leading the design of high-performance, scalable and reliable machine learning systems? Do you want to set technical direction and help shape the next generation of AI platforms powering advanced NLP applications? We are looking for a Lead Member of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will provide technical leadership across multiple teams, driving the architecture and strategy for deploying optimized NLP models to production in low latency, high throughput, and high availability environments. You will serve as a key point of contact for customers, leading the design of customized deployments to meet their specific needs, and mentoring engineers to raise the technical bar across the team. You may be a good fit if you have: 8+ years of engineering experience running production infrastructure at a large scale, with a track record of technical leadership Demonstrated experience leading the architecture
About the Team The GTM Data Science team partners with Go-to-Market, Technical Success, Product, Engineering, RevOps, and Strategic Finance to build the shared intelligence layer for OpenAI's B2B business. The team turns product usage, customer behavior, revenue, field activity, and customer feedback into rigorous insight products that help leaders and field teams understand where customers are succeeding, where adoption is blocked, and what actions will accelerate durable growth. We are building systems that make customer intelligence proactive: surfacing risk, expansion potential, product gaps, and repeatable playbooks before they show up as escalations or missed opportunities. About the Role As the Applied Data Science & Insights Lead for GTM Intelligence Solutions and Technical Success, you will be a hands-on technical leader responsible for shaping how OpenAI measures, understands, and improves customer adoption across our B2B products. You will build AI/ML-powered intelligence products that connect account health, product usage, customer lifecycle, support tier, qualitative sentiment, commercial context, and field actions into a practical operating system for GTM and Technical Success. This role will build the data science foundation for Technical Success: defining the metrics, models, operating insights, and decision systems that help the team scale customer adoption and expansion with rigor. You will also be expected to build and lead a small mighty team over time: setting direction, hiring and developing talent, creating operating cadences, and holding a high bar for technical rigor and business impact. You will lead the development of models, metrics, and decision systems that recommend what GTM and Technical Success teams should do next, explain why, and measure whether those interventions worked. Your work will help customers move from pilots to production, deepen usage across products, identify high-value use cases, reduce churn risk, and create a f
From $154K/yr
We are Datadog’s in-house technical leaders. The Technical Account Management team drives Datadog’s continued global growth by ensuring our customers realize long-term value from our platform through successful adoption, expansion, and partnership. As a Manager 1 in Technical Account Management, you will lead and develop high-performing technical teams while influencing strategy, execution, and outcomes across customers, internal partners, and the broader organization. Manager 1 leaders at Datadog are people-first managers, trusted collaborators, and operational owners. You will coach and mentor individual contributors, drive execution against team and organizational goals, and serve as a strong voice for customer needs and technical excellence. At Datadog, we place value in our office culture; the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead and coach a team of up to 6 Technical Account Managers, providing regular 1:1s, team meetings, and bi-annual performance feedback Own and track team KPIs , including scheduling, utilization, productivity, and delivery outcomes Partner closely with Sales, Customer Success, Presales, Product Management, Support, and Marketing to align post-sales strategy and execution Lead and participate in customer-facing engagements when appropriate, including escalations, strategic reviews, and key account discussions Drive account strategy discussions focused on product adoption, expansion, and services delivery Actively participate in recruiting , hiring, and onboarding efforts across your team and the broader organization Gather and synthesize customer feedback to influence product direction, process improvements, and internal initiatives Lead multiple OKR initiatives annually , coordinating and delegating efforts across your team Demonstrate thought leadership by identifyi
Work Flexibility: Remote As a Senior Lead, Data Engineering, you will serve as a technical leader who helps shape the future of enterprise data solutions. In this role, you will drive complex data initiatives, influence technical strategy, and partner with teams across the organization to build scalable, high-impact data products. This is an opportunity to solve challenging business problems while mentoring fellow engineers and elevating data engineering best practices. What You Will Do Lead the architecture, development, and modernization of scalable enterprise data platforms that support global procurement analytics and business transformation. Define and help execute a multi-year data engineering strategy focused on platform scalability, reliability, automation, technical debt reduction, and long-term maintainability. Design, build, and optimize Azure-based data solutions using technologies such as Databricks, Delta Lake, Azure Data Factory, Azure DevOps, CI/CD pipelines, and infrastructure automation. Integrate and harmonize data across multiple ERP systems by standardizing supplier, purchasing, and master data into common enterprise data models. Partner with procurement analysts, architects, engineers, and business stakeholders to translate complex business needs into reusable, scalable data products and engineering solutions. Establish engineering standards, conduct architecture reviews, improve documentation, and mentor engineers to raise the overall technical capability of the team. Identify and implement AI-enabled approaches that accelerate development, improve data quality, automate documentation, support testing, and enhance analyst productivity. Evaluate and recommend tools, frameworks, patterns, and platform investments that improve performance, reliability, security, governance, and operational ef
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why This Role Is Different This is not a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers on problems that push LLMs to their limits. You’ll rapidly understand customer domains, design custom LLM solutions, and deliver production-ready models that solve high-value, real-world problems. Train and customize frontier models — not just use APIs. You’ll leverage Cohere’s full stack: CPT, post-training, retrieval + agent integrations, model evaluations, and SOTA modeling techniques. Influence the capabilities of Cohere’s foundation models. Techniques, datasets, evaluations, and insights you develop for customers will directly shape the next generation of Cohere’s frontier models. Operate with an early-startup level of ownership inside a frontier-model company. This role combines the breadth of an early-stage CTO with the infrastructure and scale of a deep-learning lab. Wear multiple hats, set a high technical bar, and define what Applied ML at Cohere becomes. Few roles in the industry combine application, research, customer-facing engineeri
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why This Role Is Different This is not a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers on problems that push LLMs to their limits. You’ll rapidly understand customer domains, design custom LLM solutions, and deliver production-ready models that solve high-value, real-world problems. Train and customize frontier models — not just use APIs. You’ll leverage Cohere’s full stack: CPT, post-training, retrieval + agent integrations, model evaluations, and SOTA modeling techniques. Influence the capabilities of Cohere’s foundation models. Techniques, datasets, evaluations, and insights you develop for customers will directly shape the next generation of Cohere’s frontier models. Operate with an early-startup level of ownership inside a frontier-model company. This role combines the breadth of an early-stage CTO with the infrastructure and scale of a deep-learning lab. Wear multiple hats, set a high technical bar, and define what Applied ML at Cohere becomes. Few roles in the industry combine application, research, customer-facing engineeri
From $295.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox’s daily active users growing at a record pace, we are building the next generation of advertising in an immersive metaverse. Our team sits at the core of the Roblox Ads ecosystem, serving billions of impressions while providing the critical infrastructure that connects global brands with our massive community of creators. We are seeking a Tech Lead Manager (TLM) to drive the technical strategy and lead the execution of the Ads Experience team. This team is responsible for the full spectrum of our monetization experience, including Advertiser Experience (Roblox Ads Manager), Brand Experience (Brand Ads & Integrations), and Publisher Experience (Creator Monetization & Commerce). This is a foundational leadership challenge beyond standard ad-tech. As a TLM, you will be responsible for both the technical vision of our high-scale monetization backbone and the growth and management of a high-performing engineering team. You will pioneer AI integration and balance a three-sided marketplace economy in real-time while fostering a culture of technical excellence. You Will Drive Technical Strategy: Act as the technical authority for the Ads Experience team. Translate ambitious busi
Become a part of our caring community You have shipped AI products before. You understand the difference between a demo and a production system. You have strong opinions about evaluation frameworks because you have experienced the consequences of operating without them. You are at your best when you own architecture decisions while continuing to build and deliver critical code yourself. We build the platform that transforms millions of clinical documents into trusted, actionable data. Our systems use large language models (LLMs) to read medical records, extract structured facts, answer complex questions with citations back to source documents, and route complex cases to human experts. The output of these systems supports healthcare decisions that impact real members. As a Lead AI Applied Engineer, you will provide technical leadership for AI-enabled products and platforms, define architectural direction, establish engineering standards, and personally design and build the most critical components of our systems. You will lead through both technical expertise and execution, helping the team deliver reliable, scalable, and auditable AI solutions in a highly regulated healthcare environment. Why Join Us Lead the architecture of production AI systems where LLMs are foundational to the product experience. Make key technical decisions regarding model selection, system boundaries, platform architecture, and build-versus-buy strategies. Own the highest-risk and highest-impact technical challenges involving reliability, explainability, and correctness. Influence engineering culture and establish standards that shape how the team builds and ships AI products. Work on systems operating at meaningful scale, processing millions of documents and supporting healthcare decisions across a large member population. Partner
What you’ll do Act as the in-house electrical lead for Midjourney Medical: own the electrical architecture of the scanner and the technical direction for all board-level design. Own complex board design end-to-end: architecture, schematic capture, layout (high-speed digital, analog/mixed-signal, power), DFM/DFT, fabrication and assembly vendor management, bring-up, and revision control. Write firmware for embedded targets (MCU/SoC): drivers, real-time control loops, safety-relevant logic, bootloaders, and field update paths. Audit and update HDL (FPGA) code for high-throughput data acquisition, timing/synchronization, triggering, and pre-processing of ultrasound and sensor data streams. Define electrical interfaces and data contracts with software, recon/ML and mechanical teams: timing budgets, clocking/sync, signal integrity, connectors/harnessing, and failure modes. Establish electrical engineering rigor: design reviews, schematic/layout review checklists, bring-up procedures, test fixtures, and documentation suitable for a regulated medical device program (DHF, traceability, change control). Mentor and grow the electrical function; select and manage external design partners where leverage is high. What we’re looking for Deep experience designing complex boards from blank page to stable revision, including high-speed digital and analog/mixed-signal domains. Strong schematic and layout skills (Altium/KiCad or equivalent) with real signal integrity, power integrity, grounding, and EMI/EMC instincts. Solid embedded firmware background in C/C++ (and Python for tooling): peripherals, DMA, interrupts, real-time constraints, and debugging on hardware. Practical HDL experience (VHDL/Verilog/SystemVerilog) for data acquisition, timing, and streaming interfaces. Track record of owning bring-up and debug on real hardware: scopes, logic analyzers, and disciplined root-cause analysis. Technical leadership: clear trade-offs, strong written documentation, and the ability to set
About the Team The Finance & Supply Chain Engineering organization includes two complementary teams. Software Engineering builds internal full-stack applications, durable agentic workflows, plugins, MCPs, and measurable AI-enabled engineering practices. Data Engineering builds trusted analytics data assets for Finance and Supply Chain. The teams have distinct charters, with important shared dependencies and broad cross team partnerships across Engineering, Applications, Finance, and Supply Chain. About the Role We are looking for a hands-on senior technical leader who will report alongside the Software Engineering and Data Engineering managers. This is an individual-contributor role with no immediate people-management responsibility. The Tech Lead will raise the technical bar across both teams, participate in important cross-team or high-risk design decisions, and directly own and ship high-impact work. The role should improve team judgment and autonomy rather than act as a floating architect or universal approval gate. In this role, you will: Partner with the Software Engineering and Data Engineering managers as a peer technical leader; managers retain accountability for people, staffing, priorities, performance, and delivery commitments. Directly own the architecture, implementation, launch, and operation of one or more high-impact initiatives, remaining accountable for real outcomes rather than advisory output alone. Guide important design decisions that are cross-team, difficult to reverse, or material to security, financial controls, reliability, data quality, or long-term cost of ownership. Establish pragmatic engineering standards across architecture, APIs and data contracts, testing, security, reliability, observability, lineage, data quality, and operational ownership. Advance engineering standards for building with AI, including agentic workflows, evaluation, telemetry, adoption, and outcome measurement. Work across backend, full-stack, data, and big-d
$342K – $445K/yr
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Technical Lead to lead deployment and operations for OpenAI’s Silicon & Systems team. This person will become the Directly-Responsible Individual responsible for bringing OpenAI’s custom silicon and associated systems into data center environments, ensuring successful deployment, bring-up, validation, operational readiness, and ongoing reliability at scale. This role sits at the intersection of silicon, systems, infrastructure, data center operations, and software. You will lead a team focused on taking new hardware platforms from lab validation into production data center deployment. You will be responsible for building the operational processes, technical workflows, tooling, and cross-functional alignment required to deploy and operate custom AI hardware reliably in OpenAI’s supercomputing infrastructure. The ideal candidate is both a strong leader and a deeply technical operator. You should be comfortable staying close to the technical details of hardware bring-up, fleet deployment, debugging, system validation, data center integration, and production operations. This role requires strong execution, excellent cross-functional judgment, and the ability to drive clarity in ambiguous, fast-moving environments. In this role, you will: Lead a team responsible for deployment and operations of OpenAI’s custom silicon and systems in data center environments Own the path from hardware bring-up and validation through production deployment, operati
Other cities to consider
More places hiring for this role
Get new technical lead jobs in United States by email
Daily job updates · Unsubscribe anytime