Jobs in United States

Excellence Center Systems Engineer in United States

3,690 active opportunities · Updated October 2026

Explore current excellence center systems engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As an AI Reimagination Engineer in Micron's Generative AI Center of Excellence (GenAI COE), you have an outstanding chance to transform how businesses function. You will partner with different business areas to break down large, manual, multi-step processes and redesign them into effective, autonomous systems. This position combines process reimagination with AI systems engineering, making you a key part of the transformation journey. You will collaborate closely with the GenAI COE, IT architecture, security, and project teams. You will lead projects from the initial redesign to the final build, delivering solutions that business teams can adopt and scale. Responsibilities: Decompose end-to-end processes: Map current-state flows, quantify effort and risk, and lead eliminate/simplify/agentify analysis before automation. Architect the agentic solution: Build future-state flows and AI architecture, including task and agent decomposition, orchestration patterns, tool and data access, memory and context strategy, and human-in-the-loop controls. Translate inventions into buildable solutions by developing agent workflows, composing prompt and context strategies, MCP/connector and integration requirements, and evaluation criteria. Follow Micron's “Secure by Design”

PythonSQLAIC#
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As an AI Reimagination Engineer in Micron's Generative AI Center of Excellence (GenAI COE), you have an outstanding chance to transform how businesses function. You will partner with different business areas to break down large, manual, multi-step processes and redesign them into effective, autonomous systems. This position combines process reimagination with AI systems engineering, making you a key part of the transformation journey. You will collaborate closely with the GenAI COE, IT architecture, security, and project teams. You will lead projects from the initial redesign to the final build, delivering solutions that business teams can adopt and scale. Responsibilities: Decompose end-to-end processes: Map current-state flows, quantify effort and risk, and lead eliminate/simplify/agentify analysis before automation. Architect the agentic solution: Build future-state flows and AI architecture, including task and agent decomposition, orchestration patterns, tool and data access, memory and context strategy, and human-in-the-loop controls. Translate inventions into buildable solutions by developing agent workflows, composing prompt and context strategies, MCP/connector and integration requirements, and evaluation criteria. Follow Micron's “Secure by Design”

PythonSQLAIC#
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As an AI Reimagination Engineer in Micron's Generative AI Center of Excellence (GenAI COE), you have an outstanding chance to transform how businesses function. You will partner with different business areas to break down large, manual, multi-step processes and redesign them into effective, autonomous systems. This position combines process reimagination with AI systems engineering, making you a key part of the transformation journey. You will collaborate closely with the GenAI COE, IT architecture, security, and project teams. You will lead projects from the initial redesign to the final build, delivering solutions that business teams can adopt and scale. Responsibilities: Decompose end-to-end processes: Map current-state flows, quantify effort and risk, and lead eliminate/simplify/agentify analysis before automation. Architect the agentic solution: Build future-state flows and AI architecture, including task and agent decomposition, orchestration patterns, tool and data access, memory and context strategy, and human-in-the-loop controls. Translate inventions into buildable solutions by developing agent workflows, composing prompt and context strategies, MCP/connector and integration requirements, and evaluation criteria. Follow Micron's “Secure by Design”

PythonSQLAIC#
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

PythonAWSAzureGit
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten's GTM org is in hyper growth. As it grows and matures, the needs of the GTM stack get more sophisticated with scale — and this team exists to stay ahead of those needs. This role owns a core set of the platforms the field runs on. You'll administer, configure, and continuously improve them, build AI and automation on top of the CRM to keep it clean and current, and run the adoption programs that make sure the field actually uses what we put in front of them. We're an AI-native company scaling fast, and we want to add sophistication without adding drag — you'll be the person deciding how much system is enough. This isn't a classic Salesforce admin role. You know the ecosystem, but you come at it as a builder and an orchestrator, not a ticket-taker. What You'll Walk Into Central RevOps is the systems and operations center of excellence for Baseten's GTM org. The stack includes Salesforce, Pylon, Outreach, Sales Navigator, Clay, and a growing set of AI-native GTM tools. The field is scaling fast, and the systems need to mature with it without slowing anyone down. You'll co-own Salesforce with our other GTM Systems Manager and work alongside a GTM Engineering team who builds internal AI products for GTM productivity — plus partner closely with sales leadership and the field. Responsibilities Own adminship, configuration, and enablement for Pylon, Gong, Sales Navigator, Outreach, and Prospect. These platfor

Machine LearningAIGoRust
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world's most transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the future of AI computing. The Opportunity As Technical Services Director, you will lead the teams that operate and evolve Graphcore's engineering labs, high-performance computing (HPC) platforms, and data center environments globally. You will be accountable for reliable, secure, cost-effective infrastructure that supports demanding engineering, AI, silicon-development, and validation workloads. This role combines people leadership, infrastructure strategy, operational excellence, capacity and financial planning, procurement, and program delivery. You will partner with Engineering, Information Technology, Security, Finance, Facilities, Supply Chain, customers, and external suppliers. The position is based onsite in Austin and requires travel to company facilities, data centers, and supplier locations, including international travel. What You'll Do Lead, recruit, mentor, and develop the systems administration, lab operations, and technical services teams responsible for the facility supporting global Engineering and Research and Development. Own the reliability, efficiency, protection, safety, supportability, and continuous improvement of engineering labs, HPC systems, and infrastructure facilities. Establish service levels, operating standards, escalation paths, performance measures, monitoring, observability, automation, ticketing, and configuration-management practices. Translate engineering and customer requirements into infrastructure roadmaps, capacity p

LinuxAIGoExcel
S
📍 United States· Full-time
✓ Quality checkedCompany trend -81%

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. Remote (US East Coast preferred, for timezone coverage) About the team Cloud Infrastructure owns the platform every Synthesia product runs on — AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them. We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for. The role We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure. Today, too much critical operational knowledge — vendor relationships, cost management, and incident response — lives with one or two people. Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run. This is not simply a ticket-queue or keep-the-lights-on role. You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring . We deliberately pair operational and engineering work so the role grows rather than narrows. What you'll own Incident management & operational excellence — take custody of the incident process: on-call quality, resp

PythonMongoDBAWSKubernetes
H
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Manager, Electrical Engineering is responsible for the electrical systems of the KARNO generator, including high-voltage power electronics, battery systems, low- and high-voltage architecture, wiring harnesses, and the hardware that converts linear motion into electrical output. This is a working manager role: the position leads and develops a team of electrical engineers while remaining directly involved in technical execution, including circuit architecture, schematic review, and hardware bring-up in the lab. The Manager is accountable for the technical excellence, safety, and reliability of the electrical engineering function, and for establishing the design standards and review practices the team works to. The position plans team capacity, owns hiring and development for the electrical engineering staff, and partners with mechanical, controls, supply chain, and program management on system integration. KARNO systems are deployed in data center, military, and industrial applications. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Own critical electrical designs personally, including regular time at the bench and in the test cell, while leading the team as a practicing engineer. Lead the electrical engineering team in the design and development of KARNO generator electrical systems, including high-voltage power electronics, battery systems, linear generator power stages, and low-voltage controls hardwar

AIGoExcelSEM
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $280K/yr

Quick readStrong listing-quality and freshness signals

The Detection Platform organization is responsible for helping customers identify, understand, and act on issues across their environments through alerting, event intelligence, and autonomous detection capabilities. As Director, Detection Platform, you will lead a group of engineering managers and teams responsible for foundational alerting infrastructure, event management, monitor creation experiences, and AI-powered detection systems. This role sits at the center of Datadog’s efforts to evolve how customers detect, investigate, and respond to operational issues at massive scale. You will partner closely with Product Management, Applied Science, Design, and Engineering leaders to shape the future of detection and observability experiences for Datadog customers while leading a growing organization of engineers. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead a multi-team engineering organization responsible for alerting, event management, monitor creation experiences, and autonomous detection capabilities. Define and execute the technical and organizational strategy for the Detection Platform while aligning stakeholders across Engineering, Product, Design, and Applied Science. Drive innovation in AI-powered detection, anomaly identification, and signal generation that helps customers proactively identify and resolve issues. Scale highly available platform systems that process hundreds of millions of evaluations while maintaining reliability, performance, and operational excellence. Develop and mentor engineering managers and technical leaders, fostering a culture of execution, collaboration, and technical rigor. Champion customer-centric product thinking by balancing platform investments with intuitive user experiences and measurable customer

Machine LearningAIGoRust
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Observability team builds the infrastructure that empowers engineers to understand, operate, and improve the Roblox platform and ecosystem. Our team owns the end-to-end observability stack across telemetry, distributed tracing, logging, profiling, storage systems, and developer-facing visualization tools. We are looking for an Engineering Manager to lead the next generation of AI-powered observability platforms. In this role, you will help build intelligent systems that leverage AI to revolutionize CI/CD, testing, and DevOps workflows — enabling engineers to move faster, improve reliability, and operate large-scale distributed systems with greater efficiency and confidence. This is a highly impactful leadership role at the center of Roblox infrastructure. Your work will directly improve developer productivity, platform reliability, and operational excellence across the company. You will partner closely with infrastructure, product engineering, and AI platform teams to shape the future of developer tooling and autonomous operations at scale. You Have 3+ years of engineering management experience with a proven track record of hiring, mentoring, and growing high-performing teams. Strong ex

AWSCI/CDGitAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and

AWSRestAIRust
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore fosters continuous learning and innovation. Job Summary Reporting into the Systems Engineering organisation, the Distinguished Engineer, End-to-End Security Architect will define and lead the security architecture for Graphcore’s inference service platform. This role is responsible for establishing a comprehensive security strategy spanning platform, infrastructure, networking, service operations, customer assurance, and compliance readiness. Working across multiple engineering and operational functions, the successful candidate will provide technical leadership, drive security requirements, and ensure the platform delivers robust protection, resilience, and trust for customers. The Team You will work closely with teams across security architecture, infrastructure engineering, networking, site reliability engineering, platform software, firmware, data centre operations, compliance, legal, customer engineering, and customer security. The team collaborates across the business to deliver secure, reliable, and scalable AI infrastructure and services while supporting customer assurance, regulatory requirements, and operational excellence. Responsibilities and Duties Own the end-to-end security a

AIRustExcelRecruitment
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent engineering manager to own and deliver an end to end manageability stack for Data Center Systems. We are seeking an experienced manager who is deeply technical, hands-on, and has a wide system view. You will manage a team of experts, design & build OpenBMC based manageability software stack for NVIDIA’s next generation Data Center Compute Systems. We want to grow our teams with the smartest people in the world. If you're creative and autonomous, we want to hear from you! What you’ll be doing: Own and deliver OpenBMC based manageability stack for next generation Data Center Compute Systems. Own firmware delivered to data centers in terms of quality, reliability and telemetry performance. Manage and lead a distributed team of software engineers to deliver firmware stack with high quality. Work with data center architects and cloud customers for correct requirements and scope implementation to ensure speed of light product development. Work closely with cross functional teams to ensure scalable manageability architecture for all data centers products Drive efficiency, reliability and optimization in firmware architecture from a data center view point. Work closely with customers and internal teams to resolve issues at Speed of Light. What we need to see: BS, MS, or PhD in EE/CS or related field o

PythonGitAIProject Management
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as &#34;the AI computing company.&#34; We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent Senior Engineering Manager to lead a large firmware engineering organization delivering end-to-end manageability firmware for NVIDIA's next generation Data Center Compute Systems. This role owns HGX product line and OpenBMC-based management firmware and MCU firmware components in data center platforms, including architecture, execution, quality, reliability, telemetry, and customer readiness. We are seeking an experienced senior leader with strong technical depth, broad system perspective, and a proven ability to lead large teams through complex product cycles. This role is onsite in Santa Clara, CA, USA. If you're creative and autonomous, we want to hear from you! What you'll be doing: Lead a large firmware engineering organization delivering OpenBMC based firmware and MCU firmware for next-generation Data Center Compute Systems. Own HGX platform as a lead for Firmware and System software readiness working across the organization. Define and drive the long-term firmware roadmap, balancing architectural innovation with product execution and delivery milestones. Drive architecture strategy across BMC, MCU, platform software, manageability, health management, and data center firmware interfaces. <spa

PythonGitLinuxArtificial Intelligence
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. In this role you will: As a Hardware Test Engineer, you will work on Machine Learning/AI hardware system projects to craft the solutions for current and future data center deployments. You will bring a strong understanding of hardware system testing, excellent project management skills, and the ability to collaborate across multiple teams to ensure efficient lab operations. You will be responsible for designing, implementing, and executing comprehensive test plans that ensure the reliability, performance, and scalability of our supercomputing hardware systems. You will develop detailed test plans and methodologies tailored to hardware components, including processors, memory modules, custom accelerators and interconnects. You will collaborate with hardware design, manufacturing, firmware teams and vendors to identify, analyze, and resolve issues affecting hardware, power, thermal and high-speed interconnects. You will perform in-depth debugging on the hardware system Excellent analytical skills to diagnose hardware issues, troubleshoot problems, and propose solutions. Ability to interpret complex test data, identify trends, and draw meaningful conclusions. High-speed links, with a focus on SerDes (Serializer/Deserializer) technology to assess signal integrity, error rates, and overall link performance. You will collaborate with the lab manager to maintain the equipment and hardware systems, including oscilloscopes, thermal test chambers, liquid cooling systems, and other mea

PythonAWSRestMachine Learning
🔔

Get new excellence center systems engineer jobs in United States by email

Daily job updates · Unsubscribe anytime