Overview The Associate Director of Service Engineering leads the reliability, availability, and operational excellence of Natera’s lab-facing platforms. This role ensures that clinical systems, laboratory equipment workflows, and data pipelines operate with high reliability, scalability, and compliance in a regulated healthcare environment. You will lead a team responsible for production stability, incident response, and service health, partnering closely with Production Engineering, Lab Operations, Bioinformatics, Infrastructure, Facilities, and Compliance to support mission-critical genetic testing and diagnostics. Key Responsibilities Leadership & Team Development Lead, mentor, and scale a team of Service Engineers / SREs supporting clinical production systems Establish clear expectations around ownership, on-call readiness, and operational excellence Drive hiring, onboarding, performance management, and career growth Foster a blameless, learning-oriented culture focused on patient impact and reliability Service Reliability & Production Operations Own reliability and availability for production services supporting laboratory operations, reporting, and customer delivery Define and manage SLAs, SLOs, and operational KPIs aligned with clinical and business priorities Lead major incident response, ensuring rapid triage, clear communication, and thorough post-incident reviews Oversee on-call rotations, escalation paths, and operational playbooks Ensure operational readiness and go-live support for new assays, pipelines, and platform capabilities Technical Strategy & Execution Partner with Engineering and Development teams to design resilient, fault-tolerant systems Drive best practices for monitoring, alerting, logging, and observability across lab and cloud platforms Reduce operational toil through automation, tooling, and process improvements Advocate for reliability, performance, and scalability requirements early in t
Jobs in United States
Infrastructure Engineer in United States
1,475 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Job Details: Job Description: Intel is shaping the future of technology to help create a better future for the entire world. Our work in pushing forward fields like AI, analytics, and cloud-to-edge technology is at the heart of countless innovations. With a career at Intel, you'll have the opportunity to use technology to power major breakthroughs and create enhancements that improve our everyday quality of life. Join us and help make the future more wonderful for everyone. Want to learn more? Visit our YouTube Channel or the link below. Life at Intel The Role and Impact As a System Software Architect, you will play a pivotal role in designing and developing innovative software solutions that drive efficiency and scalability across Intel's manufacturing processes. You will be responsible for architecting robust systems and frameworks that enhance operational performance while ensuring seamless integration with existing technologies. Your contributions will directly impact Intel's ability to deliver world-class semiconductor products efficiently and reliably. Business Group Intel Foundry Services is at the forefront of manufacturing excellence, delivering cutting-edge solutions to meet the complex demands of the semiconductor industry. The group focuses on pioneering advancements in manufacturing technologies, infrastructure optimization, and secure operations, enabling Intel to stay ahead in a fast-evolving landscape. By joining this team, you will contribute to Intel's broader mission of solving technological challenges and empowering innovation worldwide. Key Responsibilities: - Architect scalable and efficient system sof
Datadog's Software Engineers with Systems depth leverage their experience with systems and tooling to build software that ensures Datadog remains reliable, performant, and secure. For this track, their Software Engineering experience may resemble the Distributed Systems track, but is typically applied in combination with their systems experience to build and run internal platforms and tools that our products are built on. These people typically have deep experience building and managing large cloud infrastructure deployments, or leading reliability efforts for orgs similar to ours, or building release machinery to allow hundreds or thousands of devs to do their jobs without stepping on each others' toes. The systems and tooling where they may have experience depth may include (but not limited to): bazel, build tooling, cassandra, CDN, chef, configuration management, container orchestration, consul, docker, elasticsearch envoy, haproxy, kafka, kubernetes, load balancing, network architecture, postgres, redis, release management, RPC frameworks, service discovery, spinnaker, terraform, zookeeper. Bonus: You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. #LI-KM5 Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan. Th
Yext (NYSE: YEXT) is the enterprise agentic marketing platform. Built on the world's most comprehensive structured data platform for local businesses, Yext gives brands and their partners the visibility intelligence to win every moment of discovery – across AI and traditional search. Yext's API-first architecture connects structured data to APIs, MCP servers, and generative interfaces, so partners and developers can build purpose-built experiences on the same infrastructure powering Yext's own products. Thousands of brands and digital marketing partners in financial services, healthcare, retail, hospitality, and food rely on Yext to manage, measure, and optimize visibility at scale. For more information, visit yext.com . At Yext, Product Engineering builds and evolves the technology behind our products and services. We’re looking for software engineers who want to solve meaningful technical problems, contribute to systems at scale, and help shape what we build next. We work in an agile environment with two-week sprints and regular demos that keep teams aligned and give engineers clear visibility into the impact of their work. From day one, you’ll contribute directly to the codebase and collaborate with experienced engineers from a wide range of leading universities and technology companies. We are looking for an engineer to join Team Watson , which owns and develops the systems that power Yext Search and Yext Chat. The team builds the indexing, retrieval, and serving technology that enables brands to deliver fast, relevant answers across their websites and digital experiences. Yext Search handles more than 50 million requests each month, serving users around the world in over a dozen languages. Watson also brings these search and retrieval capabilities to Yext Chat, helping conversational experiences generate useful answers grounded in trusted customer content. Because Watson’s systems serve real-time, customer-facing experiences at a global scale, engineers o
About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world’s most transformative technologies. Our new AI Engineering Campus in Austin will play a central role in building the future of AI computing. The Opportunity We are seeking a system validation engineering intern to help drive server blade and rack validation efforts for next-generation AI infrastructure hardware systems. This role focuses on post-silicon system validation across the full lifecycle of server hardware systems, ensuring functional and performance meets product objectives. You will help drive end-to-end blade and rack validation including development, execution, and debug while collaborating across silicon, firmware, systems, and platform teams. The Blade and Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Type: 12-week summer internship Timing: May - August (exact dates to be confirmed) Commitment: Full-time What You’ll Do Help drive and execute post-silicon validation goals of AI compute blades and racks including testcase planning, development, and automation Help drive validation testcase execution and system debug against program achievements and report validation progress and risks. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness Triage test failures, collect debug data, and collaborate on root cause analysis. Track validation coverage and continuously improve test processes and infrastructure. What You’ll Bring Working towards a Bachelor's
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. The Release Engineer is responsible for implementing and supporting automated software delivery processes that enable reliable, secure, and repeatable deployments across development, test, non-production, and production environments. This role serves as a technical bridge between Software Engineering, Quality Engineering, Infrastructure, and Operations teams to improve deployment speed, stability, and operational efficiency. The Release Engineer is responsible for implementing and supporting automated software delivery processes that enable reliable, secure, and repeatable deployments across development, test, non-production, and production environments. This role serves as a technical bridge between Software Engineering, Quality Engineering, Infrastructure, and Operations teams to improve deployment speed, stability, and operational efficiency. Key Responsibilities Develop, support, and maintain automated CI/CD pipelines to streamline application delivery. Standardize build, release, and deployment processes across applications and platforms. Establish repeatable, auditable, and traceable release management practices to ensure deployment consistency and compliance. Manage and support Git-based source control, branching strategies, and release workflows. Integrate automated testing, quality gates, security controls, and compliance checks into deploym
Experienced Engineering Technical Specialist Company: The Boeing Company Boeing Defense, Space & Security (BDS) is seeking Experienced Engineering Technical Specialists (Level 3) to join the team in Berkeley, MO. This position supports Risk Managed Framework compliance across multiple avionics test labs as well as the administration of embedded PCs and host PCs on avionics test benches. The Lab Infrastructure team is responsible for the design, modification, and troubleshooting of all test benches used in the avionics' lab to support systems integration testing on the F-15. Position Responsibilities Build and configure embedded PCs and host PCs on test benches Support RHEL 9 Installation, configuration, patching and maintenance of Linux server environments Support containerization Provide active directory management and file server management Troubleshoot user problems related to hardware, software, or processes Conduct monthly OS patching and anti-virus updates Create user accounts and reset passwords Run backups on workstations and servers Maintain compliance with Risk Managed Framework and NISPOM rules Keep facility signage up to date Maintain lab access list Travel may be required up to 10% of the time; domestically and/or internationally depending on business needs. This position requires an active U.S. Secret Security Clearance (U.S. Citizenship Required). (A U.S. Security Clearance that has been active in the past 24 months is considered active) Basic Qualifications (Required Skills/Experience) <ul
The DFP Engineer – Manufacturing role defines and implements the validation and screening of new silicon features within high‑volume manufacturing flows. You will translate product requirements into executable test methodologies, infrastructure, and detailed manufacturing test plans that ensure quality, yield, and efficiency at scale. This role sits at the intersection of multiple multi-functional teams to make manufacturing test an outstanding part of the overall codesign and DFP lifecycle. What you will be doing: Own end-to-end manufacturing test methodology across all test stages. Translate system specs and product POR into DFP requirements, test content, coverage, and flows. Define and maintain the DFP roadmap, including infrastructure and turning point planning. Partner multi-functionally to implement test content, debug hooks, and coverage improvements. Drive alignment on manufacturability, test time, binning strategies, and cost vs. coverage trade-offs. Embed testability requirements into design to enable robust screening and debug. Define data and analytics frameworks to support yield analysis and continuous improvement. Lead DFP documentation as the single source of truth and feed findings into future methodologies. What we need to see: MS in Electrical Engineering, Computer Engineering, or related field (or equivalent experience) 6+ years in silicon post‑silicon validation and/or high‑volume manufacturing test for complex SoCs, GPUs, CPUs, or similar. Hands‑on experience with test content bring‑up, limit setting, correlation to characterization, and yield/coverage optimization. Proficiency with scripting and data analysis (e.g., Python, MATLAB, R, SQL) for test data analytics, limit tuning, and yield/debug analysis. <
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Software Engineer (Federal) Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Okta Identity Governance Team Okta Identity Governance (OIG) is Okta's Identity Governance and Administration solution, responsible for some of the most critical and visible workflows in enterprise identity: how people request access, how approvers grant it, how entitlements are enforced, and how organizations prove to auditors that only the right people have the right access. OIG is the leading contributor to Okta's new product bookings, and it runs inside some of the largest enterprises in the world — the kind of accounts where a single tenant carries hundreds of thousands of identities across thousands of applications. OIG currently supports high-compliance public sector workloads and plans to expand our footprint to support the most secure, highly classified environments in government. The Senior Software Engineer (Federal) Opportunity As a Senior Fullstack Engineer on the OIG (Federal) team, your mission
Join the Team at Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore brings together deep expertise to solve complex problems and deliver meaningful progress in AI compute. Ready to raise the standard for supplier quality in advanced AI manufacturing? Apply now to be part of the journey. Job Summary We are seeking a highly motivated individual with experience in managing supply chain quality in a high-tech manufacturing organisation. The ideal candidate will be hands-on, comfortable with ambiguity and able to manage multiple projects and stakeholders simultaneously. Working across our entire supply chain in a technically challenging, fast paced environment, you will ensure the readiness of our global supply chain to ramp successfully as we develop and manufacture the world’s most advanced AI systems and services. The Team The Graphcore Quality Team delivers outstanding customer experience and champions sustainable excellence throughout Graphcore. We are responsible for customer, supply chain and product quality, as well as organisational compliance, governance and assurance. You will be joining a diverse
Manufacturing Test Engineering Manager Position Summary We are seeking an experienced Manufacturing Test Engineering Manager to lead the development and execution of the end-to-end manufacturing test strategy for next-generation AI server platforms and datacenter infrastructure. This role is responsible for defining and driving the manufacturing test architecture from L6 board assembly through L11 rack-level integration and final system validation , ensuring world-class product quality, manufacturability, and production scalability. This leader will manage a team of 3–5 Manufacturing Test Engineers while partnering closely with Hardware, Firmware, Platform, Validation, Quality, Operations, and Joint Design Manufacturing (JDM) partners. The role owns the manufacturing test strategy, test coverage, factory test infrastructure, manufacturing capacity planning, and continuous improvement of manufacturing quality. Key Responsibilities Manufacturing Test Strategy Define and own the end-to-end manufacturing test strategy from L6 board assembly through L11 rack integration . Develop standardized manufacturing test methodologies that optimize quality, throughput, cost of test, and scalability across multiple products and JDM sites. Establish manufacturing test standards, best practices, and engineering processes that support high-volume server manufacturing. Technical Leadership & People Management Lead, mentor, and develop a team of 3–5 Manufacturing Test Engineers supporting multiple hardware programs. Establish team priorities, allocate resources, and ensure successful execution of manufacturing test deliverables. Foster a culture of technical excellence, accountability, collaboration, and continuous improvement. Serve as the primary technical escalation point for manufacturing test and production issues. Cross-Functional Engineering Collaboration Partner with Hardware, Platform, Firmware, Validation, Reliability, Quality, and Operations teams to ensure manufac
We are investing in agentic AI and need a Senior AI Engineer to lead the design and delivery of these systems. This is a foundational hire: you will own both the agent-facing workstreams — pipelines, orchestration, conversational interfaces — and the underlying context layer that makes them reliable, including memory management, knowledge graph integration, and retrieval infrastructure. You will work closely with data engineers, project leads, and client stakeholders, and play a key role in shaping how Lynx builds and ships AI solutions at scale. What This Involves: Lead the architecture and delivery of agentic AI systems end-to-end: agents, orchestration, tool use, and multi-step reasoning workflows. Own the context layer: design and implement memory architectures (episodic, semantic, working memory) and integrate GraphRAG and knowledge graph retrieval into agentic pipelines. Build robust RAG systems — including vector retrieval, graph traversal, and hybrid search — and ensure retrieval quality through evaluation frameworks. Translate client requirements into technical designs, presenting approaches and trade-offs to both technical and non-technical stakeholders. Define standards and reusable patterns for agentic AI development that other engineers at Lynx can build on. Set up observability, evaluation, and monitoring pipelines to ensure AI systems perform correctly in production. Requirements: 5–8 years of software or ML engineering experience, with at least 2–3 years building LLM-based or agentic AI systems in production. Deep hands-on experience with agentic frameworks (LangChain, LlamaIndex, AutoGen, CrewAI, or similar) and LLM APIs (OpenAI, Anthropic, etc.). Strong understanding of agent design patterns: ReAct, planning loops, tool use, multi-agent coordination, and memory architectures. Practical experience with GraphRAG or knowledge graph-based retrieval (e.g., Neo4j, Microsoft GraphRAG) and vector databases (Pinecone, Weaviate, Qdrant, etc.). Proficiency in
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We are looking for an experienced Mechanical Engineer with 7+ years of experience in design of IT hardware from chip/package to system levels. You’ll work alongside experts in thermal, mechanical, electrical, software, and systems engineering to support the design, analysis, and validation of mechanical and thermal systems that ensure the reliability, efficiency, and longevity of mission-critical hardware. This position requires strong analytical skills, hands-on testing experience, and the ability to work in a fast-paced, cross-disciplinary environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead mechanical design for AI supercomputer product in the data center application Collaborate with the cross functional team to design and optimize thermal solutions for data center hardware, including chips, power modules, and system-level cooling architectures Collaborate with cross-functional teams to integrate thermal management strategies into hardware design, from concept to mass production Design and validate mechanical systems, including chassis, enclosures, cooling systems, and high-power connections, ensuring alignment with performance and reliability standards. Perform 3D modeling, FEA, tolerance analysis, and prototyping, ensuring manufacturability and a
About the Team OpenAI’s Healthcare team is working to ensure that advances in AI meaningfully improve health for everyone. We build AI systems that support patients, clinicians, and healthcare organizations while setting a high standard for deploying AI responsibly in a high-stakes domain. On the Healthcare team, we are building products and infrastructure that bring OpenAI’s capabilities into healthcare organizations and clinical workflows. We serve health care systems, insurance companies, and the broad range of healthcare IT companies. Our products include connecting ChatGPT with healthcare data and systems and building reliable enterprise experiences that can operate within complex security, privacy, compliance, and interoperability requirements. We work closely with product, research, design, security, privacy, GTM, and our healthcare customers to turn powerful AI capabilities into dependable products that clinicians and healthcare organizations use in their everyday work. About the Role We are looking for backend and full-stack software engineers to build the systems behind OpenAI’s healthcare products. You will work across backend services, APIs, data systems, integrations, and product surfaces to connect AI with healthcare workflows and enterprise data. You’ll tackle the technical challenges involved in making these systems secure, reliable, observable, and scalable enough for real-world healthcare environments. This is a product-engineering role for someone who can move between customer problems and complex system architecture, and who is comfortable owning ambiguous problems end-to-end. In this role, you will: Design and build backend and full-stack systems powering OpenAI’s healthcare products. Build services, APIs, and data pipelines that connect OpenAI products with healthcare systems and enterprise data. Develop integrations with electronic health records and other healthcare data sources. Design systems that meet demanding requirements around privacy,
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: As a New Grad Software Engineer, you'll join a team of exceptional builders working on products that are reshaping how the world creates software. You'll have the opportunity to work on everything from our AI-powered development platform to the distributed systems that enable real-time collaboration for millions of developers. This is a chance to define your career while defining the future of software development. You'll work on problems that matter, with the autonomy to drive solutions and the support to grow into a technical leader. What you will build: Product features that delight users and make it possible for anybody to create software AI coding agent that understands intent and generates production-ready applications Cloud infrastructure that provides instant, powerful development environments at global scale Platform features that enable one click deployments and scale to millions of users Required skills and experience: Recent graduate (2027) with a degree in Computer Science, Computer Engineering, or related field Strong programming skills in a modern language (JavaScript/TypeScript, Python, Go, Rust) Full-stack capabilities with experience in React, Node.js, and database technologies Growth orientation - eager to learn new technologies and take on increasing responsibility Collaborative spirit - you work well in cross-functional teams and value diverse perspectives What we value : Problem-solving mindset: Ability to approach complex operational challenges systematically and devise effective solutions Self-directed and autonomous: Capable of working independently while collaborating effectively with cross-functional teams Strong communication skills: Ability to explain complex technical conce
Other cities to consider
More places hiring for this role
Get new infrastructure engineer jobs in United States by email
Daily job updates · Unsubscribe anytime