Jobs in United States

Ai Systems Engineer in United States

5,082 active opportunities · Updated October 2026

Explore current ai systems engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the hardware platforms that support the next generation of AI systems. The Opportunity We are looking for a recent graduate or early-career engineer to join the Mech/Thermal Platform team as a Graduate Mechanical Engineer. You will contribute to the mechanical design, thermal validation, and verification of advanced AI hardware and data center systems. You will work with experienced mechanical, thermal, hardware, and systems engineers throughout the development lifecycle. The role combines 3D computer-aided design, prototype evaluation, lab testing, data analysis, troubleshooting, and clear engineering documentation. What You Will Do Create and update mechanical parts, assemblies, and drawings using Creo or comparable 3D CAD software. Take ownership of defined mechanical design tasks from requirements and concepts through detailed design, review, release, and verification. Apply mechanical design, materials, manufacturing, heat transfer, and thermodynamics principles to engineering decisions. Develop mechanical and thermal test plans and procedures for prototypes and development systems. Set up and operate laboratory equipment, collect accurate data, analyze results, and document conclusions. Evaluate prototypes and identify mechanical, thermal, assembly, or manufacturability issues. Support design verification, troubleshooting, root-cause analysis, and implementation of verified design improvements. Maintain accurate engineering documentation, including design notes, drawings, test procedures, results, and change r

Artificial IntelligenceAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

PythonAIExcelHR
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

PythonAIExcelHR
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

PythonArtificial IntelligenceAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team builds next-generation AI-native silicon and systems while working closely with software, research, and manufacturing partners to co-design hardware tightly integrated with AI models. In addition to delivering systems for OpenAI’s supercomputing infrastructure, the team develops the tools, methodologies, and strategic partnerships needed to accelerate hardware innovation. About the Role We’re seeking an experienced Hardware Strategic Sourcing Manager to own sourcing strategy and supplier partnerships for fiber and optical interconnect components across OpenAI’s next-generation AI infrastructure. Reporting to the Head of Partnerships & Strategic Sourcing, you will lead sourcing across fiber cable assemblies, internal optical harnesses, fiber shuffles, optical backplane assemblies, connectorized and standalone passive optical assemblies, fiber-array units (FAUs), fiber-to-chip and coupling interfaces, detachable connectors, optical routing, and assigned optical packaging, assembly, and test services. You will work closely with electrical engineering, optical engineering, systems engineering, mechanical and packaging engineering, quality, rack integration, data-center deployment,manufacturing, supply chain, finance, legal, and program management teams to translate demanding bandwidth, signal integrity, reliability, and scale requirements into resilient supplier partnerships and scalable commercial strategies. Your work will directly support the performance, reliability, manufacturability, and scale of the high-speed optical connectivity required for OpenAI’s next-generation AI systems. In this role, you will: Develop and execute a comprehensive sourcing strategy for fiber and optical interconnect components supporting high-bandwidth AI systems and infrastructure. Own sourcing across optical fiber cable assembli

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions. Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency. About the Role We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments. This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment. The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries. Key Responsibilities Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure. Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios. Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize

PythonAWSGitRest
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the hardware platforms that support the next generation of AI systems. The Opportunity As a Systems Engineering Intern, you will contribute to projects that combine hardware, firmware, and software engineering for advanced AI compute platforms. You will work with experienced engineers on subsystem design, laboratory testing, system validation, automation, and performance analysis. The internship provides hands-on experience with modern hardware development and system-level engineering. You will own clearly defined technical tasks with guidance from the team and document your methods, results, and conclusions. What You Will Do Support the design and testing of CPU and high-speed input and output subsystems for advanced compute platforms. Run laboratory tests and measurements to help evaluate performance, power, signal behavior, and reliability. Contribute to system-level validation by creating scripts and tools that streamline testing, data collection, and analysis. Explore emerging input and output technologies, including PCIe 6.0 and 800G Ethernet, and learn how they support advanced computing workloads. Assist with investigations into platform power, cooling, and energy efficiency, including liquid-cooling systems for high-performance processors. Use power meters, oscilloscopes, logic analyzers, or comparable lab equipment under appropriate supervision. Analyze test results, identify unexpected behavior, and work with engineers to reproduce and investigate issues. Collaborate across hardware, firmware, software, m

PythonArtificial IntelligenceAIC++
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Graphcore Director-Post Silicon Validation (Functional) Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Austin, Texas which will play a central role in Graphcore's work building the future of AI computing. We are developing the next generation of AI compute, a large-scale system-on-chip (SoC) designed to power future high-performance AI systems. Job Summary We have an exciting opportunity to be part of a collaborative, cross-functional development team validating cutting-edge, high-performance AI chips and platforms. You will play a key role in supporting new product introductions and validation. You will lead a team delivering post-silicon validation across the full AI SoC, working across silicon, firmware, and platform levels. The role requires a deep technical understanding, strong hands-on debug experience, and the ability to collaborate effectively with hardware, software, and systems engineering teams. Working within the Validation team, you will be involved with bringing first silicon to life, functionally validating it and working closely with many other teams to help it become a fully characterised and working product, reporting project status/progress to program management on a regular basis. You will have the opportunity to provide technical guidance to other engineering team members. In this role, you can leverage our experience and industry knowledge to architect and drive implementation of continuous improvements to test infrastructure and processes. Th

PythonGitLinuxAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for an experienced systems software engineer to help define and build the host software stack for our custom next-generation AI systems. You will work close to the hardware on performance-critical software, including Linux kernel drivers, high-throughput I/O paths, and system-scale networking and RDMA. This role spans architecture, implementation, platform bring-up, debugging, and performance optimization. You will work across hardware and software boundaries to make new systems usable end to end, from low-level device interfaces through userspace tooling and production validation. In this role you will: Design, implement, and debug host-side systems software for AI infrastructure, including Linux kernel drivers and supporting userspace components. Build and optimize software paths for high-throughput, low-latency communication, including RDMA and related networking functionality. Develop software around PCIe, DMA, NICs, accelerators, memory movement, and device interaction. Bring up new hardware platforms and diagnose complex issues across kernel, firmware, networking, and hardware boundaries. Build tooling for integration, testing, diagnostics, observability, qualification, and performance characterization. Collaborate with hardware, networking, and platform teams to define interfaces and integrate new capabilities. Work with external vendors where needed to integrate technologies and drive issues to resolution. Contribute across the systems sof

PythonAWSLinuxRest
P
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of Postman’s AI-powered API and agentic systems in production. This role focuses on monitoring, availability, incident response, and automation to support AI services and tools trusted by millions of developers globally. What You’ll Do Develop and manage reliability metrics (SLOs) for AI-driven API services and agentic AI platform features Implement comprehensive observability and monitoring systems for real-time performance and fault detection Design and drive automated failover, recovery, and incident response strategies for high-availability AI infrastructure Optimize resource utilization, particularly GPU/accelerator efficiency, ensuring cost-effective AI system operation Collaborate closely with engineering, platform, and product teams to align reliability efforts with broader organizational goals Lead efforts to build internal tooling and automation focused on AI system stability and operational excellence Drive continuo

AIGoRustDevOps
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a Product Manufacturing Engineer to drive manufacturing strategy and execution for next-generation AI hardware within the Silicon & Systems team, with a particular focus on PCBA manufacturing, assembly, process development, and production readiness. You will work closely with design engineering, systems engineering, operations, TPMs, contract manufacturers, suppliers, and other external partners to ensure new hardware products successfully transition from concept and prototype builds through NPI and high-volume production. You’ll own critical manufacturing initiatives, identify and resolve production risks, and help establish the processes and controls required to deliver complex hardware at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Drive manufacturing and quality initiatives to ensure product success from concept and early development through NPI, production launch, and scale. Lead manufacturing process development for next-generation AI hardware systems, partnering closely with design engineering, systems engineering, operations, TPMs, suppliers, and manufacturing partners. Own PCBA manufacturing development and production readiness, including assembly processes, process validation, manufacturing test, rework, yield improvement, and successful transition into volume production. Establish NPI manufact

AWSRestAIRust
SF
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -100%

From $103.1K/yr

Quick readStrong listing-quality and freshness signals

About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As an HRIS Solutions Architect on the Business Systems Engineering team, you will serve as the architect and platform owner for our People Technology ecosystem. You will design, build, and operate enterprise-grade systems that power talent management, workforce analytics, and HR operations. This role requires deep technical expertise in integration architecture, data modeling, API design, and software engineering practices. You will collaborate with People & Culture, Finance, Platform Engineering, Data Engineering, and Application Development teams to build resilient, scalable, and AI-enabled systems. Responsibilities: Architect the future of People Tech: Design scalable, AI-ready architectures that automate end-to-end HR processes and unlock new capabilities. Innovate with Agentic AI: Champion the exploration and implementation of AI agents and machine learning to revolutionize workflows and generate actionable insights. Build impactful applications: Build impactful applications using Workday Extend and develop complex integrations in Workday Studio using Java, Python, or other supported languages. Lead technical strategy: Partner with Enterprise Architecture to define system boundaries, data flows, and integration standards; conduct design reviews and provide technical mentorship. Build scalable integrations: Design event-driven integration patterns, RESTful/SOAP APIs, and data pipelines connecting HRIS Systems across our technology ecosystem. Drive technical excellence: Esta

PythonJavaCI/CDGit
S
📍 Bellevue, Washington, United States· Full-time· Remote
✓ High-confidence listingCompany trend -92.9%
Quick readStrong listing-quality and freshness signals

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a rock star leader for our AI DevEx team. This team plays a crucial part in transforming how the Snowflake product is developed, evaluated and supported. The job is no less than building the ecosystem that powers our AI-pilled engineers and our agentic organizations and reinvents the SDLC for Snowflake. As the manager for this growing team, you will lead the group of engineers working to reinvent how code is written, reviewed and runs in production to usher scale and velocity for the AI era. With context as the new source code, you will make it possible for every engineer at Snowflake to deliver through intent and direction. In this role, you will: Lead, coach, and grow the ES AI DevEx team while creating a high-energy, cohesive environment with strong planning, ownership, and career development. Deliver the vision and roadmap for transforming the SDLC for all Snowflake engineers. Drive measurable improvements in code review latency, ensure Snowflake developers can use the latest harnesses and models, and create the golden paths that agentic codebases are built upon. Act as an agent of clarity in a space that is moving fast, making high quality decisions quickly by keeping up with the latest developments in the industry. Own operational excellence for the area

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a Security Engineer to join our First-Party Hardware team. In this role, you will own the end-to-end security foundation for OpenAI's first-party AI hardware systems, working across hardware security, embedded security, system security, and practical deployment at data center scale. You will partner with silicon, hardware, firmware, infrastructure, manufacturing, operations, and security teams to define and deliver system-level device trust. This includes boot integrity, device identity, provisioning, attestation, management-plane security, storage encryption, debug controls, firmware update and recovery, RMA, and decommissioning. You will be accountable for turning threat models into requirements, requirements into implementation, and implementation into validation evidence that can support launch decisions. Location: San Francisco, CA (Hybrid: 3 days/week onsite) Relocation assistance available. In this role, you will: Own security requirements, threat models, validation strategy, and launch-readiness evidence for first-party hardware platforms from early design through production deployment. Design and review secure boot, measured boot, roots of trust, platform firmware resilience, firmware signing, recovery, and anti-rollback strategies across heterogeneous devices. Own device identity, provisioning, enrollment, attestation, certificate lifecycle, and key-management requirements across manufacturing and data center bring-up. Harden management

AWSRestAIC++
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a System Software Engineer to join our First-Party Hardware team. In this role, you will design, build, integrate, and validate low-level system software for the manageability and health of OpenAI's first-party AI hardware systems. You will work across BMC, Linux, firmware interfaces, automation infra, boot and recovery, hardware diagnostics, telemetry, host and platform drivers, network software interfaces, and manufacturing and fleet readiness. A major part of this role is owning the acceptance path for partner-delivered system software: defining requirements, reviewing code and artifacts, reproducing builds, building tests, pushing fixes, and producing the evidence needed for launch decisions. This role is hands-on and high-ownership. You will write and review low-level software, debug issues across hardware and software boundaries, build infra and automation to test and manage devices in lab, guide partner deliverables, build validation evidence, and help carry platforms from bring-up through production deployment. Location: San Francisco, CA (Hybrid: 3 days/week onsite) Relocation assistance available. In this role, you will: Design, develop, and maintain low-level firmware and system software for first-party AI hardware manageability, including BMC software, Redfish services, gNMI telemetry, firmware update and recovery flows, BIOS/UEFI interactions, platform drivers, and hardware diagnostics. Own integration and acceptance of partner and ve

AWSLinuxRestAI
🔔

Get new ai systems engineer jobs in United States by email

Daily job updates · Unsubscribe anytime