Jobs in United States

Lead Senior Engineering Manager 2c Platform Reliability in United States

2,434 active opportunities · Updated October 2026

Explore current lead senior engineering manager 2c platform reliability jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

Hiring demand

44/100

watch · 138 related jobs

Hiring trend

-54.7%

Job postings compared with the previous 30 days

Remote options

20.3%

Share of matching jobs listed as remote

Typical salary

$190K – $190K/yr

Based on 9 salary observations

MT
📍 San Jose, California, Canada
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Our vision is to transform how the world uses information to enrich life for all. Join an inclusive team passionate about one thing: using their expertise in the relentless pursuit of innovation for customers and partners. The solutions we build help make everything from virtual reality experiences to breakthroughs in neural networks possible. We do it all while committing to integrity, sustainability, and giving back to our communities. Because doing so can fuel the very innovation we are pursuing. Job Summary The Senior Director, Semiconductor Sourcing and NPI Procurement leads the global strategy, sourcing, and supplier ecosystem for wafer foundry services, IP and EDA services, discrete actives semiconductors, and 3rd party NAND flash controller categories as well as the New Product Introduction Procurement function. This role is accountable for delivering cost competitiveness, supply assurance, and long-term resilience across a complex semiconductor value chain. Partnering closely with engineering, manufacturing, and business leaders, the role drives category strategies, supplier performance, and risk mitigation at category and product level while enabling Micron’s growth and innovation priorities. Success in this role requires deep semiconductor domain expertise, strong commercial acumen, and the ability to lead large, global teams to deliver enterprise-level outcomes. Main Responsibilities Define and execute global procurement strategies for wafer foundry services, IP and EDA ser

AISupply ChainProcurementRecruitment
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Role We’re seeking an Associate General Counsel to lead commercial legal strategy and execution for OpenAI’s silicon initiatives and the semiconductor ecosystem that supports our AI infrastructure. This is a senior, highly cross-functional role for a lawyer who can advise business and technical leaders across the semiconductor development, manufacturing, and supply lifecycle. The role will partner closely with silicon engineering, infrastructure, strategic sourcing, supply chain, manufacturing, finance, partnerships, IP, policy, regulatory, and trade compliance teams to structure and negotiate strategic arrangements with semiconductor ecosystem partners, technology providers, and critical suppliers. This person will help establish the commercial frameworks needed to protect OpenAI’s technology and support resilient, scalable commercial arrangements. We’re looking for an experienced technology transactions lawyer with meaningful semiconductor industry experience who can translate highly technical and operational issues into practical commercial structures. The ideal candidate combines strong judgment, commercial creativity, and the ability to lead complex, high-value transactions in a fast-moving and complex global ecosystem. This role is based in San Francisco, CA. We use a hybrid work model of 3-days in the office per week and offer relocation assistance to new employees. This role will: Lead commercial legal strategy and risk management for OpenAI’s silicon and related technology initiatives. Advise senior business, engineering, sourcing, and operational stakeholders on strategic relationships, commercial priorities, and complex transactions. Draft, negotiate, and advise on sophisticated development, manufacturing, supply, licensing, services, and strategic partnership agreements. Structure commercial arrangements that support long-term business and operational requirements. Advise on the commercial and operational issues that arise across the developmen

AWSRestAIGo
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Principal Engineer at Micron’s HBM New Product Validation team in Boise, ID you will be a senior technical guide responsible for validation strategy and readiness for future-generation High Bandwidth Memory (HBM) products. You will apply deep expertise in complex RAM and high-bandwidth memory architecture and building to establish architecture validation strategies. You will impact decisions that improve validation observability and debug efficiency . You will lead investigations of complex architecture, building , and silicon issues from pre-silicon simulation through post-silicon characterization. You will partner across Design Engineering, Design Validation, Product Engineering, Systems teams, and will develop and promote AI-enabled validation methodologies that improve efficiency, quality, and scalability. Responsibilities: Define architecture-aware validation strategies and coverage direction for future-generation HBM products. Drive validation readiness for next-generation products by preparing strategies, test environments, and execution plans before first silicon, reducing risk, accelerating issue resolution, and preventing late-stage surprises. Evaluate new DRAM/HBM architectural features and define validation requirements and testability needs before DBR. <span style="color:#

PythonGitLinuxAI
C
📍 Dallas 8000 Frankford Road, United States
✓ Quality checkedCompany trend +340.2%

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking a highly experienced Principal / Director-level Full-Stack Software Development Engineer to join our Digital Caremark organization and lead the architecture, design, and delivery of next-generation digital solutions. This role spans AI-enabled applications, scalable digital platforms, and enterprise integrations that power critical healthcare and pharmacy experiences. This is a senior technical leadership role for a hands-on engineer who can operate across the full stack—from intuitive front-end applications to resilient backend services—while setting architectural direction, influencing engineering standards, and mentoring teams. The ideal candidate combines deep technical expertise, platform thinking, and strong collaboration skills to build secure, scalable, API-first solutions in a highly regulated environment. Key Responsibilities: Architecture & System Design • Define and drive architecture for large-scale distributed systems and digital platforms • Lead design reviews and set architecture standards and best practices • Champion API-first, microservices, and event-driven architecture patterns • Ensure systems meet scalability, reliability, security, and compliance requirements • Balance performance, cost, and speed in technical decision-making Fu

TypeScriptPythonJavaReact
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a Senior Mechanical Engineer to lead the design, integration, and sustaining engineering of mechanical subsystems in robotic platforms. You will work closely with experienced engineers and cross-functional partners to set functional requirements, develop and iterate on hardware to meet program expectations. This role is aimed at candidates with strong fundamentals in mechanical system design — including tolerance, alignment, load paths, wear, and failure modes — and robotics. You will contribute to real hardware programs moving from prototype through early production. You will lead both new subsystem development and ongoing improvements to existing systems based on testing, field performance, and manufacturing feedback. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will Lead the design and iteration of mechanical subsystems, including structures, mechanisms, and actuators. Create and maintain CAD models, assemblies, and drawings with appropriate tolerancing and documentation. Build and test prototypes, supporting debugging of mechanical issues such as fit, alignment, friction, and wear. Assist in developing test methods and executing validation to evaluate performance, durability, and failure modes. Work with cross-functional teams to integrate mechanical components with sensors, actuators, and control systems. Support transition of designs from prototype to manufacturable assemblies, incorporating DFM and DFA considerations. Collaborate with manufacturing partners

AWSRestAIGo
O
📍 Washington, District of Columbia, United States· Full-time
✓ Quality checkedCompany trend -82%

About the team The AI Deployment Engineering team is responsible for ensuring the safe and effective deployment of Generative AI applications. We act as a trusted advisor and thought partner for our customers, working to build an effective backlog of GenAI use cases for their industry and drive them to production through strong technical guidance. As an AI Deployment Engineer (ADE) in the OpenAI for Government team, you’ll help government agencies transform their organization through solutions such as automated content generation, contextual search, and novel applications that make use of our newest, most exciting models and technology. About the Role We are looking for a driven solutions leader with a product mindset to partner with our public sector customers and ensure they achieve tangible value with GenAI. You will pair with government agencies (federal, state, and local), policymakers, and other public institutions to establish a GenAI strategy and identify the highest value applications. You’ll then partner with their technical teams, subject matter experts, systems integrators, and implementation partners to move from prototype through production. You’ll take a holistic view of their needs and design an architecture using the OpenAI API and other services to maximize customer value. You will collaborate closely with Sales, Solutions Engineering, Global Affairs, Applied Research, and Product teams. This role is based in Washington, DC. We offer relocation support to new employees. In this role, you will: Deeply embed with our most sophisticated public sector customers as the technical lead, serving as their technical thought partner to ideate and build novel applications on our API and other OpenAI products. Work with senior customer stakeholders to identify the best applications of AI in their industry and to build/qualify a comprehensive backlog to support their AI roadmap. Intervene directly to accelerate customer time to value through building hands-on pr

JavaScriptPythonJavaAWS
MT
📍 San Jose, California, Canada
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron's Global Supplier Quality organization is seeking a Controller Quality Principal Engineer to lead the quality strategy, qualification, and continuous improvement of storage and memory controller (ASIC/SoC) manufacturers supporting Micron's SSD, embedded, and storage solutions portfolio. This is a senior technical leadership role responsible for driving controller supplier quality performance from design qualification through mass production and field support The successful candidate will collaborate across functions with ASIC Development Team, Compose Engineering, Product Engineering, Dependability, Manufacturing, and Commodity Management, as well as directly with controller IC vendors, foundries, third party reliability labs and OSAT (outsourced assembly and test) partners, to ensure controller quality, reliability, and supply continuity meet Micron's standards. This role can be based in Taiwan, Hyderabad, or San Jose and will work extensively across time zones with global partners and suppliers! Develops, evaluates, revises, and applies technical quality assurance protocols/methods to inspect and test in-process raw materials, production equipment, and finished products. Ensures activities and items are in compliance with both company quality assurance standards and applicable government regulations. Performs analysis and identifies trends in the inspection of finished products, in-process materials and bulk raw materials, and recommends corrective actions when vital. Ensures that established manufacturing inspection, sampling and statisti

AISupply ChainRecruitment
O
📍 Washington, District of Columbia, United States· Full-time· Remote
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team The mission of the Applied AI Engineering team is to enable the secure and impactful implementation of GenAI solutions. We serve as technical thought partners and trusted advisors to our clients, ideating high-value use cases and providing the hands-on guidance necessary to drive projects into production. In this role within the Government team, you will empower agencies to evolve their operations through automated content synthesis, advanced search capabilities, and bespoke applications leveraging our latest foundational technologies and models. About the Role We are looking for a driven solutions leader with a product mindset to partner with our public sector customers and ensure they achieve tangible value with GenAI. You will pair with government agencies (federal, state, and local), policymakers, and other public institutions to establish a GenAI strategy and identify the highest value applications. You’ll then partner with their technical teams, subject matter experts, systems integrators, and implementation partners to move from prototype through production. You’ll take a holistic view of their needs and design an architecture using the OpenAI API and other services to maximize customer value. You will collaborate closely with Sales, Solutions Engineering, Global Affairs, Applied Research, and Product teams. This role is based in Washington, DC. We offer relocation support to new employees. In this role, you will: Deeply embed with our most sophisticated public sector customers as the technical lead, serving as their technical thought partner to ideate and build novel applications on our API and other OpenAI foundational technologies like Codex. Work with senior customer stakeholders to identify the best applications of GenAI in their industry and to build/qualify a comprehensive backlog to support their AI roadmap. Intervene directly to accelerate customer time to value through building hands-on prototypes and/or by delivering impactful strate

TypeScriptPythonAWSRest
S
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.6%

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role At Sentry, Support is an engineering discipline. Our customers are the greatest technical minds in the world—developers at elite enterprises building the future of software—and they deserve answers that go deeper than a knowledge base link. We're looking for an APAC Technical Support Engineer based in San Francisco to join our global Support Engineering team. This role is designed to provide APAC coverage to our users; with the shift being Sunday through Thursday 4PM-12AM PST. We are architecting the Technical Support engine . We’re looking for an experienced engineer to help us redefine the standard of technical support by combining deep human expertise with autonomous agentic systems. You are a debugger of both code and systems. You will treat support volume as a data signal to build automated resolution paths, ensuring our human engineers only touch the most complex, high-impact architectural puzzles. Sentry Support Engineers aren't just clearing queues; they are Orchestrators . You will engage with our users across GitHub, Discord, and our internal systems, while acting as the Technical Lead for our Agentic Ops. You ensure that when a developer asks a complex question, our systems have the right context and a seamless "Human-in-the-Loop" path to you when deep, nuanced expertise is required. In this role you will Master the Sentry Ecosystem & Support Elite Developers Deep-Dive Debugging: Perform root-cause analysis on complex issues and distributed tracing gaps across polyglot environments. Support the Great Minds: Act as a strategic consultant for senior engineers at our largest enterprise customers, s

JavaScriptPythonJavaReact
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Controls Network Engineer, you will design, validate, and scale the controls and OT network architectures that support high-density AI data centers. You will work across controls systems, OT infrastructure, telemetry, commissioning, deployment, and operations, partnering with mechanical, electrical, IT/networking, security, and external delivery teams. About the Role We are seeking a mid to senior OT Network Engineer with a strong controls systems background to lead the design and operation of resilient, secure, and scalable OT network architectures for high-density AI data centers. This role translates compute, power, cooling, and operational requirements into practical OT network designs, evaluates vendor solutions, and drives technical decisions across controls infrastructure, telemetry, commissioning, and operations. The ideal candidate has strong hands-on experience in mission-critical OT environments, including industrial networking, virtualized infrastructure, and OT network operations, with expertise in routing, switching, segmentation, firewall policy, time synchronization, monitoring, and network lifecycle support. Key Responsibilities Define controls, automation, and OT network requirements for AI data center campuses. Develop reference architectures, engineering standards, and reusable design templates. Review and develop basis-of-design and functional design documents, including OT network diagrams, IP/VLAN schemes, telemetry architectures, data flow diagrams, and commissioning requirements. Design OT and infrastructure network architectures, including physical topology, logical topology, IP addressing, subnetting, VLA

PythonSQLPostgreSQLMySQL
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

AWSLinuxRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a senior Actuator Electromagnetic Design Engineer to lead the development of custom electromechanical actuators for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, reliability, and manufacturing teams. This role focuses on the design, integration, and validation of precision electromechanical systems, including motors, transmissions, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will: Lead the architecture, design, and integration of custom robotic actuators, including the design, simulation, integration and sourcing of custom electromagnetic components. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-loop performance. Devel

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a Senior Actuator Design and Integration Engineer to lead the development of custom electromechanical actuators for advanced robotic systems. You will own actuator development from early architecture and concept generation through prototype validation and system integration, partnering closely with mechanical, electrical, controls, firmware, reliability, and manufacturing teams. This role focuses on the design, integration, and validation of precision electromechanical systems, including motors, transmissions, sensing, structural components, and thermal architectures. You will help drive actuator development across the full engineering lifecycle while establishing scalable design, test, and integration practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will Lead the architecture, design, and integration of custom robotic actuators, including motors, transmissions, sensing, thermal systems, structural components, and packaging. Define actuator requirements and system-level trade studies around torque density, bandwidth, efficiency, thermal performance, backdrivability, inertia, reliability, manufacturability, and cost. Design precision electromechanical assemblies with strong attention to tolerances, alignment, load paths, thermal expansion, sealing, wear, and serviceability. Drive actuator integration into robotic systems, partnering closely with controls, firmware, electrical, and robotics software teams to optimize closed-loop pe

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des

AWSRestAIGo

Related career options

Similar roles with stronger pay

Client Service Associate

Demand 46/100 · 8 jobs

$840K – $840K/yr

Salary →

$840K – $840K/yr

Salary →
Director of Product

Demand 43/100 · 6 jobs

$382.5K – $382.5K/yr

Salary →
Physical Design Engineer

Demand 43/100 · 8 jobs

$300K – $300K/yr

Salary →
Sr. Engineer

Demand 42/100 · 7 jobs

$300K – $300K/yr

Salary →
Senior Director

Demand 43/100 · 22 jobs

$278.9K – $278.9K/yr

Salary →
🔔

Get new lead senior engineering manager 2c platform reliability jobs in United States by email

Daily job updates · Unsubscribe anytime