What we’re doing isn’t easy, but nothing worth doing ever is. Diligent builds helpful robots that work safely and autonomously in real world environments. We move quickly, solve messy problems, and care deeply about reliability at scale. As a Fleet Engineer, you'll own the reliability and continuous improvement of our deployed robotic fleet — leading hands-on investigations into how and why robots fail in the field, across the mobile base, charging/docking, motion and power, connectivity (modem), and sensor hardware. You'll combine remote data analysis with bench/lab failure analysis at our Austin HQ, turning field-technician reports and fleet data into clear problem statements, validated root causes, and corrective actions driven to closure with engineering, operations, manufacturing, and vendors. We are hiring a Lead Engineer, Issue Management & Triage to lead the systems, tooling, and team at the intersection of our Customers, Remote Operations Center (ROC), and Engineering. This is a highly technical, hands-on role focused on building the infrastructure that powers how we detect, triage, diagnose, and resolve issues across a deployed robotic fleet. You will work deeply with Engineering teams to design classification frameworks, build internal tools, and develop automation pipelines that improve reliability at scale. Location: Austin preferred, Remote possible (U.S.) Travel: if remote up to ~50% travel to Austin, TX (especially in the your first 90 days) What You’ll Do: Own Issue Management & Triage Systems Design and own end-to-end systems for issue intake, triage, and escalation. Define severity frameworks, SLAs, and ensure issues are consistently structured for engineering prioritization. Build Tools & Automation (Hands-On) Develop automation and pipelines to ingest, process, and classify operational data, reducing manual triage effort. Contribute directly to codebases (Python, backend services) and partner with Engineer
Jobs in United States
Lead Engineer 2c Process Development Core Engineering Module Manager Manager Manager in United States
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current lead engineer 2c process development core engineering module manager manager manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
From $125K/yr
About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Lead Engineer on the Product Catalog Manager Team, you will help set the technical direction for the systems that power Stitch Fix’s product data ecosystem. You will work on the tools, workflows, and data models that support the full product lifecycle, from new style creation and catalog enrichment to product readiness, validation, and downstream product experiences. You will own complex problem spaces from discovery through delivery, translate business and merchandising needs into scalable technical solutions, and lead execution across ambiguous, cross-functional initiatives. This role requires strong technical judgment, deep ownership, clear communication, and the ability to influence partners across Engineering, Product, Merchandising, Data Science, and Operations. Your work will directly impact product data quality, catalog accuracy, merchandising efficiency, product readiness, and the client experience. Responsibilities: Own and evolve critical catalog systems, including product onboarding, attribute management, data enrichment, validation workflows, and product readiness tooling. Design and operate scalable services and data models that ensure product information is accurate, complete, consistent, and available to downstream systems. Drive discovery with Product, Merchandising, Data Science, and Operations partners to identify high-impact problems, evaluate tradeoffs, and define clear technical roadmaps. Independently lead initiatives from concept through production rollout, including tec
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. The Team The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Responsibilities and Duties Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. Own execution against program achievements and report validation progress and risks. Triage test failures, collect debug data, and collaborate on root cause analysis. Track
From $244K/yr
About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Opportunity: Datadog’s Staff Engineers are our technical leaders operating at the forefront of technology, building solutions that take us through at least our next five years of growth. They do this in three major ways: As individual contributors, they bring world class technical abilities to deliver industry leading systems in areas such as data visualization, virtual runtime profiling, and planet scale streaming. As technical leaders they bring experienced technical breadth and communication skills to tackling design and architectural problems spanning the organization, charting the right course, then leading delivery. In both roles they participate in the staff engineering community and help us learn from what the industry is doing and what we've built before, and so improve company wide standards around software and systems engineering. Some examples of projects a staff engineer may own include designing and building a new data storage engine handling hundreds of millions of records per second, being the lead engineer building a new product like synthetics or profiling, or rebuilding a critical service to handle the next two orders of magnitude of scale. What You'll Do: Be the technical owner of multiple pieces of critical architecture in your area of the business Own delivery of the systems you architect from beginning-to-end, doing what it takes to get things shipped and at full scale in production Dive deep into performance of systems; inventing new approaches that bring efficiency at scale Who You Are: You have a BS/MS/P
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a Lead Robotics Safety Engineer to define and lead the product safety strategy for our robotics program. You will work closely with leadership, engineering, research, legal, and policy teams to ensure safety is integrated into our products from the earliest stages of development through deployment. This role combines product safety, regulatory strategy, systems engineering, and risk management. You will help shape how we think about safety across robotic platforms, influence product architecture and development decisions, and ensure our systems are designed to meet both current and emerging regulatory expectations. You will serve as a subject matter expert on robotics safety standards, regulatory frameworks, and safety engineering practices while helping establish the long-term safety strategy for our products and organization. This role is based in San Francisco, CA. This role will be expected to be in office 4 days per week and offer relocation assistance to new employees. In this role, you will Define and lead the product safety strategy for robotic systems across the development lifecycle. Monitor and influence emerging regulatory frameworks, industry standards, and policy developments relevant to robotics and autonomous systems. Partner with engineering teams to translate safety requirements into product architecture, design decisions, and engineering requirements. Develop and maintain system-level hazard analyses, risk assessments, and safety cases for robotic products. Establish safety requirements, verification strat
We are looking for a disciplined and dynamic, Lead System Engineer – compute blade and rack Validation to join our growing compute rack validation team. As a diligent leader in Systems Engineering, you will drive multiple aspects of validation throughout the life cycle of the program. In this high visibility position, you will be part of a leading team to innovate and improve system bring-up and enablement abilities, as well as silicon and system validation to deliver the highest quality, industry leading technologies to market. Your technical leadership skills, validation and debug expertise will be necessary towards product development, definition, root cause and resolution. Your agility and collaborative approach will be essential to work within System Validation & other engineering teams (System Architects, SoC and Rack FW etc). The technical leader will be driving keys areas of system validation including leading first silicon & system bring-up (nodes and rack level systems) - rack level systems and blades will be based of ARM server architecture. Candidate will be immersed in challenging system enablement work, system validation (end-to-end) methodology, tests development and execution as well as triage/debug of critical issues to meet critical program milestones at POR quality. The candidate will also be a key contributor to state-of-the-art HW and lab capabilities for Grapchore’s system engineering. The candidate should be able to work in a global environment while maintaining a synergetic culture. Primary Responsibilities: Lead the systemenablement (including first silicon and other FW components) to ensure system capabilities are brought up as per plan of record and system architecture spec. Drive organization wide methodology for Firmware integration and best known configuration (HW/FW/SW) usage model by leading the release of deployment ready solutions. Develop key methodologies, lab HW and system SW capabilit
From $82.1K/yr
About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Lead Integration Engineer on the Business Systems Engineering team, you will collaborate with cross functional teams to understand requirements and synthesize them to build & enhance globally scalable solutions. This role requires deep technical expertise in integration architecture and is responsible for identifying issues, building & testing integrations, and administering Workday and other People platforms. You will wear multiple hats; modifying business processes, complex reporting, and providing end-user support. Build the future of People Tech: Build scalable, AI-ready solutions that automate end-to-end HR processes and unlock new capabilities. Establish and evolve automation and Gen AI frameworks purpose-built for the employee experience. Deliver technical excellence: Define system boundaries, data flows, and integration standards and design reviews. Establish integration patterns, reusable code libraries, and architectural standards; implement monitoring, logging, and alerting for all HRIS systems. Build scalable integrations: Design event-driven integration patterns, RESTful/SOAP APIs, and data pipelines connecting HRIS Systems across our technology ecosystem; leveraging your skills in Studio, Cloud Connect, EIB, iLoad, Conversion, Core Connectors, XSLT, RaaS, and Web Services. Ensure security & compliance: Build secure architectures using OAuth 2.0, encryption, and maintain SOX compliance across all integrations. Drive impact & Collaboration: Partner
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab
From $154K/yr
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the world, delivering the object, block, and file storage platforms that power GoDaddy's hosting infrastructure, internal services, OpenStack environments, and next-generation AI/HPC workloads. If you're passionate about distributed systems, storage architecture, and solving failure scenarios at massive scale, this is an opportunity to work on infrastructure few engineers will experience in their careers. Ceph is a strategic platform at GoDaddy — not an ancillary service. Our global footprint includes 80+ production clusters, 20,000+ OSDs, 1,830 storage nodes, 300 PB of raw capacity, and 69 billion objects spanning five datacenters across three continents. The platform supports RBD, RGW (S3/Swift), and CephFS workloads through more than 1,550 pools, 574,000 placement groups, and 900+ MDS daemons, creating engineering challenges that demand deep expertise in storage architecture, data durability, performance optimization, automation, and observability. As a Lead Senior Site Reliability Engineer, you'll serve as one of the principal technical leaders for GoDaddy's Ceph platform. You'll design the next generation of storage clusters, lead major platform upgrades, drive capacity and hardware strategy, and establish the standards that govern how the platform scales. You'll be the engineer the team turns to for the most complex s
From $151K/yr
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity We are seeking an experienced and vision-driven Lead Enterprise Systems Engineer to join our engineering team. In this role, you will bridge the gap between business objectives, solution architecture, and hands-on execution. The ideal candidate remains actively involved in coding (roughly 70–80% of the time) while serving as the primary technical point of contact for project stakeholders. What you'll do Technical Vision & Solution Architecture Lead the architectural design, development, and deployment of resilient, scalable solutions in our Salesforce Platform for both Sales & CPQ. Translate business and product requirements into clear, technical roadmaps and system specifications. Establish engineering best practices, design patterns, coding standards, and testing strategies. Hands-On Execution & Quality Assurance Write clean, maintainable, and highly efficient APEX code alongside the Salesforce development team. Conduct thorough code reviews to ensure quality, security, and performance. Manage technical debt, proactively balancing speed of delivery with long-term system health. Team Leadership & Mentorship Provide technical guidance, direct support, and actionable feedback to Salesforce engineers. Mentor team members to foster technical growth and career advancement. Lead agile ceremonies (sprint planning, daily stand-ups, technical grooming, post-mortems). Cross-Functional Collaboration Partner closely with Technical Managers, Enterpri
About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve people's lives. About the Role We are seeking a lead thermal simulation engineer to help accelerate the design and development of next-generation robotic systems through modeling, simulation, and analysis. You will work closely with mechanical, electrical, controls, and robotics engineers to evaluate designs before hardware is built, identify risks early, and guide critical architecture decisions. This role spans structural and thermal analysis and design across robotic subsystems including actuators, mechanisms, structures, electronics, and integrated systems. You will develop simulation workflows that improve engineering velocity, increase confidence in design decisions, and help us build more capable, reliable, and manufacturable robotic platforms. This role is based in San Francisco, CA. This role will be expected to be in office 4 days per week and offer relocation assistance to new employees. In this role, you will: Perform thermal simulations to assess heat generation, cooling strategies, thermal interfaces, and system-level thermal performance Partner with mechanical, electrical, and controls engineers to influence design decisions early in development Build simulation models to evaluate robotic actuators, transmissions, mechanisms, structures, soft goods, and integrated assemblies Correlate simulation results with physical testing and develop methodologies to improve model accuracy Support architecture trade studies by evaluating design concepts before hardware is built Develop simulation workflows, standards, and best practices that scale across the robotics o
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. As a Lead Migrations Forward Deployed Engineer (Migration FDE), you will lead scoping, architecture and direction as the FDE team embeds with customers to lead the modernization of their data and application ecosystem using the Snowflake AI Data Cloud. You will partner closely with clients to understand their business challenges to design, build, and implement cutting-edge data solutions on Snowflake. Your role is crucial in enabling customers to harness the full potential of their data, driving insights, and powering AI initiatives. You'll act as a trusted technical advisor, ensuring customers achieve significant business value and success with Snowflake. WHAT YOU WILL BE RESPONSIBLE FOR: Migration Project Execution : You will work with multiple FDE teams that each focus on a specific customer. You’ll set architecture and direction for the team as it works to design, build, and deploy robust software systems, tooling, and automation required to support customer solutions on Snowflake. Software Development: All Migration FDE members, leads and otherwise, contribute to our product. We build prototypes, contribute to feature development, fix problems that we find. Doing so helps drive impact on the roadmap and improves individual depth in the product itself. Client Collaborat
What you’ll do Act as the in-house electrical lead for Midjourney Medical: own the electrical architecture of the scanner and the technical direction for all board-level design. Own complex board design end-to-end: architecture, schematic capture, layout (high-speed digital, analog/mixed-signal, power), DFM/DFT, fabrication and assembly vendor management, bring-up, and revision control. Write firmware for embedded targets (MCU/SoC): drivers, real-time control loops, safety-relevant logic, bootloaders, and field update paths. Audit and update HDL (FPGA) code for high-throughput data acquisition, timing/synchronization, triggering, and pre-processing of ultrasound and sensor data streams. Define electrical interfaces and data contracts with software, recon/ML and mechanical teams: timing budgets, clocking/sync, signal integrity, connectors/harnessing, and failure modes. Establish electrical engineering rigor: design reviews, schematic/layout review checklists, bring-up procedures, test fixtures, and documentation suitable for a regulated medical device program (DHF, traceability, change control). Mentor and grow the electrical function; select and manage external design partners where leverage is high. What we’re looking for Deep experience designing complex boards from blank page to stable revision, including high-speed digital and analog/mixed-signal domains. Strong schematic and layout skills (Altium/KiCad or equivalent) with real signal integrity, power integrity, grounding, and EMI/EMC instincts. Solid embedded firmware background in C/C++ (and Python for tooling): peripherals, DMA, interrupts, real-time constraints, and debugging on hardware. Practical HDL experience (VHDL/Verilog/SystemVerilog) for data acquisition, timing, and streaming interfaces. Track record of owning bring-up and debug on real hardware: scopes, logic analyzers, and disciplined root-cause analysis. Technical leadership: clear trade-offs, strong written documentation, and the ability to set
We are investing in agentic AI and need a Senior AI Engineer to lead the design and delivery of these systems. This is a foundational hire: you will own both the agent-facing workstreams — pipelines, orchestration, conversational interfaces — and the underlying context layer that makes them reliable, including memory management, knowledge graph integration, and retrieval infrastructure. You will work closely with data engineers, project leads, and client stakeholders, and play a key role in shaping how Lynx builds and ships AI solutions at scale. What This Involves: Lead the architecture and delivery of agentic AI systems end-to-end: agents, orchestration, tool use, and multi-step reasoning workflows. Own the context layer: design and implement memory architectures (episodic, semantic, working memory) and integrate GraphRAG and knowledge graph retrieval into agentic pipelines. Build robust RAG systems — including vector retrieval, graph traversal, and hybrid search — and ensure retrieval quality through evaluation frameworks. Translate client requirements into technical designs, presenting approaches and trade-offs to both technical and non-technical stakeholders. Define standards and reusable patterns for agentic AI development that other engineers at Lynx can build on. Set up observability, evaluation, and monitoring pipelines to ensure AI systems perform correctly in production. Requirements: 5–8 years of software or ML engineering experience, with at least 2–3 years building LLM-based or agentic AI systems in production. Deep hands-on experience with agentic frameworks (LangChain, LlamaIndex, AutoGen, CrewAI, or similar) and LLM APIs (OpenAI, Anthropic, etc.). Strong understanding of agent design patterns: ReAct, planning loops, tool use, multi-agent coordination, and memory architectures. Practical experience with GraphRAG or knowledge graph-based retrieval (e.g., Neo4j, Microsoft GraphRAG) and vector databases (Pinecone, Weaviate, Qdrant, etc.). Proficiency in
From $200K/yr
We are looking for a talented engineer to lead evaluation of startup acquisition opportunities in the AI, cloud and security space. You will drive product evaluations, prepare and manage technical architecture discussions with target groups in Product and Engineering and provide roadmap suggestions for M&A and investments for Datadog. You will be a key partner to Datadog’s C-level leadership and highly visible at the most senior levels of Datadog. The role is reporting into the Senior Director of Product Strategy and falls within the Product organization. We are looking for an innovative and strategic thinker who is passionate about the latest tech being developed by startups in the cloud, AI and security space. The ideal candidate enjoys researching and evaluating new technologies, works effectively with cross-functional teams, and communicates opinions concisely to our leadership team. Broad understanding of relevant Cloud Technologies and deep understanding of the full coverage of Datadogs current offerings is necessary. The Corporate Development team is small and values authentic, strong-willed individuals who think creatively and proactively. This role leads technical due diligence from a product and architecture perspective across our acquisition pipeline. You'll scope and stand up proof-of-concept and sandbox environments to stress-test candidate products, then give an honest, unvarnished view of their quality and depth - the kind of assessment that holds up regardless of deal momentum. You'll assess technical architecture, flag the risks and open questions that matter most early, and turn that into a clear post-acquisition integration path. Working closely with engineering, you'll keep the evaluation focused on what's actually decision-relevant, then translate the findings into strategic recommendations for leadership and help carry the integration through by partnering with the right people on the other side. What You’l
Other cities to consider
More places hiring for this role
Get new lead engineer 2c process development core engineering module manager manager manager jobs in United States by email
Daily job updates · Unsubscribe anytime