About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe
Jobs in United States
Senior It Systems Engineer in United States
2,033 active opportunities · Updated October 2026
Showing
15 jobs
Explore current senior it systems engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
From $131K/yr
Help shape the technology that enables a global organisation to do its best work. As Senior Manager, Platform Engineering, you’ll lead the team responsible for Diligent’s Atlassian and Microsoft platforms while setting the architectural direction for the wider internal IT estate. You’ll combine people leadership, enterprise platform strategy and hands-on technical judgement to create secure, reliable and scalable experiences for employees worldwide. From modernising service management and automating joiner, mover and leaver processes to enabling AI safely through Microsoft Copilot and Atlassian Rovo, your work will reduce friction, strengthen governance and deliver measurable business impact. Working across IT, Security, HR, Finance, Legal, Compliance and business teams, you’ll turn complex requirements into well-governed platforms that are easy to use, resilient and ready for the future. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead, coach and grow a global team of platform engineers and systems administrators, building a high-performing and inclusive culture. Own the strategy, architecture, governance and roadmap for Atlassian Cloud, including Jira, Jira Service Management, Confluence, Atlassian Guard and Rovo. Set the direction for Diligent’s Microsoft 365 E5 estate, including Teams, SharePoint, Exchange Online, Intune, Defender, Purview, Power Platform and Copilot. Design scalable integration and automation patterns across identity, HRIS, ITSM and business systems using APIs, event-driven automation, Okta Workflows, Power Platform and scripting. Partner with IT Support to improve self-service, automate repetitive work and reduce ticket volume, escalation effort and time to resolution. Establish strong standards for security, access governance, AI adoption, reliability, compliance and business continuity across the internal technology estate. These are the essentials you’ll need to get an interview Significant experience in i
$225K – $300K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Fullstack Software Engineer on CLEAR’s Healthcare team, you will build and scale secure, interoperable identity and data solutions that connect patients, providers, and partners. You’ll operate at the intersection of modern web platforms, healthcare interoperability standards, and high-assurance identity systems powering frictionless, trusted healthcare experiences nationwide. A brief highlight of our tech stack: Python / React / Typescript AWS cloud What you’ll do: Design and deliver secure, scalable fullstack solutions that integrate with enterprise EHR systems and national health information exchange frameworks Build and maintain healthcare data integrations leveraging FHIR (RESTful APIs/JSON) and HL7 v2 messaging to enable compliant, real-time data exchange Develop identity resolution and patient matching capabilities using identifiers such as MRNs and NPIs to ensure integrity across disparate clinical systems Partner with Engineering, Security, Product, and Health Information Management teams to implement compliant, audit-ready workflows for regulated healthcare processes Collaborate with external vendors (e.g., Epic Technical Services) to troubleshoot integration issues, manage deployments across TST/PRD environments, and ensure production reliability How you’ll measure success: Successful delivery and stability of FHIR/HL7 integrations across healthcare partners Reduction in data integrity issues related to patient matching and identity resolution High system uptime and successful production deployments across tiered
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry's issue platform processes billions of events every day to help millions of developers find and fix bugs. The hard part is deciding which events point to a problem, which belong together, and what a developer needs to know to investigate. As a Senior Software Engineer on the Issue Detection team, you'll design, build, and operate the systems that make those decisions. You'll work on real-time processing pipelines, monitors, and analysis systems that detect problems and turn them into issues. The work combines distributed systems with product engineering. Choices about detection accuracy and processing latency affect which problems developers see and how soon they can act. You'll help shape how developers monitor their applications and how Sentry groups related events into issues. You'll also build the context developers and AI agents need to investigate what went wrong. Keeping these systems reliable and fast as Sentry grows is part of the job, alongside making the issues they produce more useful. In this role you will Build and scale features on a product surface handling billions of events daily, where both query latency and correctness are immediately visible to users. Own the design and delivery of substantial projects end to end, scoping alongside product and design, making the technical calls within your scope, shipping, and instrumenting what you ship so the team can measure it. You will contribute to meaningful technical product decisions : grouping quality, search performance, migrations and backfills against enormous datasets, and making the surface work well for both humans and agents. Champi
Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug compl
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. The Team The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Responsibilities and Duties Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. Own execution against program achievements and report validation progress and risks. Triage test failures, collect debug data, and collaborate on root cause analysis. Track
Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning. The Team The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle. Responsibilities and Duties: · Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams · Apply ad
From $295.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Why Safety AI Systems? As Senior Engineering Manager for Safety AI Systems at Roblox, you'll lead technical efforts and manage a team of experienced engineers to develop innovative AI solutions for multimodal content safety. You’ll oversee machine learning systems, constructing multimodal model architectures, improving data quality, training pipelines, and model performance to address challenges like real-time multi-verse content understanding and advanced moderation with large vision language models, spanning avatars, images, videos, audios, text, code / data models, and their composites. In close collaboration with product, policy, and Trust & Safety teams, you'll design large-scale systems to detect and mitigate abusive behavior before it harms the community. You'll own critical services at massive scale, balancing user freedom with platform civility to protect and empower our users. Your leadership will help ensure Roblox remains a safe, inclusive space for self-expression and shared experiences. You Will Own the vision, technical direction, and execution of machine learning solutions for the Multimodal Safety AI system, ensuring these systems effectively detect and prevent ha
From $242.1K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Engine DataModel team, you will own and innovate on the foundational components that form the backbone of the Roblox platform. In the Roblox Engine, the DataModel is a tree-like structure that is analogous to a scenegraph in other 3D engines. This role will report to the engineering manager and will be based out of our HQ in San Mateo, CA in a hybrid model 3 days a week (Tuesdays to Thursdays). Our team owns: The core structures and systems are used to build the DataModel and interact with it. The C++ reflection bindings that form the Engine’s Luau API surface and let creators interact with the DataModel. We’ve built custom codegen tooling to generate the C++ for these reflection bindings and other related structures. DataModel serialization … and much more! You will: Develop engine code that performs well for all user-created games on the Roblox platform. Build the core systems and data structures used in the Roblox engine, working with other teams to find universal solutions. Take ownership of projects throughout their full lifecycles. Execute code that performs well on all the devices Roblox supports—from desktop clients to mobile phone clients to con
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Luau App Foundations team is responsible for the core infrastructure of the Roblox application. They bridge the gap between app and the performance-heavy Roblox game engine. This team is building the unified stack that powers mission-critical surfaces like Home, Avatar, and Social for millions of concurrent users. They are the ones who make it possible for "Web-style" development efficiency to exist within a high-performance C++ game engine. Why is this role exciting: Technical Pioneer: You will be writing libraries and modules using C++ inside a world class Roblox proprietary Game Engine. Systematic Impact: This is a "Force Multiplier" role. The frameworks and components you build will be used by dozens of other engineering teams to ship their features. Complex Problem Solving: You aren’t just building an app; you’re managing smooth data flow through a client that has to perform perfectly on a $100 Android phone and a $3,000 Gaming PC simultaneously. 0 to 1 Transitions: You will lead the charge in shaping some of the most crucial components of the App written in C++ inside the Game Engine. Key Challenges: Bridging Tech Stacks: The libraries you write sit between a modern UI (written in
From $242.1K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Engine DataModel team, you will own and innovate on the foundational components that form the backbone of the Roblox platform. In the Roblox Engine, the DataModel is a tree-like structure that is analogous to a scenegraph in other 3D engines. This role will report to the engineering manager and will be based out of our HQ in San Mateo, CA in a hybrid model 3 days a week (Tuesdays to Thursdays). Our team owns: The core structures and systems are used to build the DataModel and interact with it. The C++ reflection bindings that form the Engine’s Luau API surface and let creators interact with the DataModel. We’ve built custom codegen tooling to generate the C++ for these reflection bindings and other related structures. DataModel serialization … and much more! You will: Develop engine code that performs well for all user-created games on the Roblox platform. Build the core systems and data structures used in the Roblox engine, working with other teams to find universal solutions. Take ownership of projects throughout their full lifecycles. Execute code that performs well on all the devices Roblox supports—from desktop clients to mobile phone clients to con
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op
From $10K/yr
About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Ramp is in a critical phase of growth. We grew immensely last year and are building out a talented business systems team to ensure we maintain this trajectory for years to come. You’ll work directly with our Sales, Account Management, Partnerships, and Product teams to execute mission-critical business systems projects across the organization. This is a key role where you will be uniquely positioned to impact the full picture of Ramp’s growth efforts through systems development. What You’ll Do Work alongside Sales Operations to administer key go-to-market business systems, including Salesforce, Outreach, Qualified, Zendesk, Hubspot, Looker, Gong.io Build and deploy automation (flows), validations, and applications in Salesforce Implement new systems and integrations as needed Analyze key business requirements and systems capabilities to write specifications for systems build and run end to end implementation Create key reports and dashboards to track systems performance and data accuracy Write and maintain clear documentation on syste
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G
Other cities to consider
More places hiring for this role
Get new senior it systems engineer jobs in United States by email
Daily job updates · Unsubscribe anytime