About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About this role At Sentry, Finance plays a critical role in helping the company understand where we're investing, what we're getting in return, and where we can operate more effectively. We are looking for an FP&A Manager who will help build Sentry's FP&A function while focusing on establishing best-in-class FinOps practices. You'll partner across the business on forecasting and financial analysis, with a significant focus on cloud infrastructure, AI costs, vendor spend, and other areas where better financial visibility can translate directly into better decisions. You will be expected to understand the economics behind our infrastructure, build the financial models and reporting needed to manage it, identify meaningful opportunities, and work with Engineering and other partners to turn those insights into action. You'll also operate as a generalist FP&A partner and play an important role in building the processes, models, and operating rhythms that Finance will use as Sentry grows. You will report to the Director of FP&A. In this role you will Own forecasting, reporting, and financial analysis for significant areas of Sentry's operations, with particular emphasis on cloud infrastructure, AI-related spend, software, vendors, and other major cost categories. Partner closely with Engineering and infrastructure leaders to understand the drivers of cloud and AI spend, translate technical consumption into financial forecasts, and identify opportunities to improve unit economics and gross margin. Develop reporting that makes cloud and infrastructure costs understandable and actionable, including trends, cost alloca
Jobs in United States
Infrastructure Engineer in United States
1,475 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. You will build the operating system for EPD — the systems, agents, and practices that make engineering, product, and design unreasonably effective at delivery, with Jira as the substrate. The EPD Systems team is building the infrastructure Vanta's engineering, product, and design organization depends on to understand itself and operate well. We're constructing three interlocking layers: sources of truth at the foundation, a shared interpretive layer that reasons across all of them, and the operating system that makes the work itself legible. This role owns the operating system. This is not a traditional PMO or status-reporting role. You build — agents, automation, workflow design, and the practices that make teams want to use the system rather than route around it. What you’ll do as a Senior Systems Designer at Vanta: Design and own workflow and hierarchy across engineering, product, and design in Jira Make sure Jira reflects real practice, and real practice reflects what needs to show up in Jira — in both directions Build the agents and automation that run the system yourself, with AI as part of how you work Drive adoption: go to the teams whose practice doesn't match the substrate today and change that — through conversation, well-built artifacts, or clear instruction that works without you in the room Understand delivery breakdowns at the mechanism level and build fixes that address the root cause, not the symptom Build alongside teammates who are growing into more technical work — raise their ceiling, not just your own output How to be successful in this role: You build and ship, recently, with AI as part of how you work. "
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. You will own the data Vanta's EPD organization actually runs on — bringing new signal sources online, standardizing them at the point of ingestion, and building the systems that make the underlying data trustworthy rather than merely stored. The EPD Systems team is building the infrastructure Vanta's engineering, product, and design organization depends on to understand itself. Not by asking teams to be more diligent but by going to the source, processing it, standardizing it, and pushing value back out so each source of truth earns its own adoption. What you’ll do as a Operations Manager, Signal Systems at Vanta: Bring new signal sources online end to end — from discovery and scoping through ingestion, standardization, and live operation Identify the specific failure modes in each information source and build systems that mitigate them at the point of ingestion Go to the teams that produce and consume a signal, understand what the information actually means and what they need back from it, and build accordingly Replace human-diligence dependencies with engineering solutions: derive fields, pull from source systems, validate at write time Build alongside teammates who are growing into building — raise their technical ceiling, not just your own output Partner with the inference and systems layers to ensure what you produce is queryable, trustworthy, and ready to build on How to be successful in this role: You look at a data source and see its failure modes before its contents: where it lies, where it goes stale, where it's duplicated, where the schema won't hold at 10x Your first move on an adherence problem is an engineering an
About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t
About the Team The Agent Infrastructure team at OpenAI is responsible for building systems that enable training and deployment of highly useful AI agents, both internally and for the world. We work hand-in-hand with researchers to design and scale the environment in which agentic models are trained – providing a workspace for AI models to execute code, debug issues, and develop software just as human SWEs do. Our training environment for agentic models operates at an extremely high scale and has the flexibility to emulate any environment in which an agent might work. At the same time, our team builds and maintains OpenAI’s core platform for the deployment and execution of agents in production. Our systems power products such as Codex, Operator, tool use in ChatGPT, and future agentic products. Some of the most challenging technical problems in scaling the capabilities and utility of agents and agentic models lie in the infrastructure layer – and our team is focused on building the research and production systems that enable OpenAI to train the most capable models in the world, and maximize the utility of our agentic products for users around the world. About the Role As a Software Engineer on the Agent Infrastructure team, you will have the opportunity to work closely with both research and product at OpenAI - building and scaling systems to train highly capable agentic models, and building the platform and integrations to launch new agents to hundreds of millions of users worldwide. Your work will consist of both building new capabilities - standing up the infrastructure and integrations needed to train more complex agentic models - and rapidly scaling these new capabilities to some of the largest compute clusters in the world. At the same time, you’ll be instrumental to the launch of agentic products at OpenAI - building, maintaining, and scaling the production platform on which all agents run. We’re looking for people with deep experience building AI infrastructu
This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea
About the Team The GPT Infrastructure team builds systems that turn advances in model inference and optimization into reliable production capabilities. We enable OpenAI workloads to be qualified and optimized across new accelerator platforms without requiring a one-off port and tuning effort for every hardware target. Our work spans distributed systems, model execution, compilers and runtimes, performance engineering, secure partner integrations, evaluation systems, and developer tooling. We build the infrastructure that makes optimization workflows automated, reproducible, and trustworthy. About the Role We are seeking a software engineer to help build the platform that qualifies and optimizes inference workloads across heterogeneous compute environments. You will develop both OpenAI-hosted services and secure partner-side software for running long-lived optimization workflows. These workflows generate candidate kernels, runtime configurations, and serving-stack changes; compile and execute them on target hardware; verify their correctness; measure their performance; and use the results to guide further optimization. You will work across model architecture, distributed execution, compilers, runtimes, networking, and accelerator systems. A central part of the role is turning research prototypes and one-off hardware bring-up efforts into reliable, reusable infrastructure with clear contracts, reproducible results, strong observability, and well-defined security boundaries. Key Responsibilities Design, build, and operate APIs and control-plane services for long-running workload qualification and optimization campaigns, including scheduling, retries, checkpointing, resource budgets, and observability. Build secure partner-side execution and evaluation software that can compile, run, verify, profile, and benchmark candidate artifacts on accelerator hardware. Integrate model workloads, hardware profiles, compiler toolchains, runtimes, serving engines, and distributed-exe
NVIDIA’s EDA Infrastructure organization builds and operates the systems that support chip development. We are looking for an engineering manager to lead the team responsible for operational processes and platforms across incident management, maintenance, on-call, issue management, and customer-serving readiness. You will own the roadmap and delivery, from defining how teams work to building the tools they use. You will partner with infrastructure and service owners to improve reliability, reduce manual work, and ensure services are ready to support customers. Your team will use automation, AI, and lessons from operational events to drive improvements. What you’ll be doing: Lead a team and own the roadmap for operational processes and platforms, from requirements and delivery through adoption and results. Set technical direction, prioritize work, and guide execution across engineering and operational disciplines. Partner with infrastructure, product, and security teams to establish consistent practices for incident response, maintenance, on-call, issue management, and customer-serving readiness. Hire and develop engineers and technical leads, building a team with clear ownership and accountability. Align priorities across teams, communicate progress and risks, and provide technical leadership during major incidents. What we need to see: <span style="co
From $105K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for an early career Software Engineer to join our Infrastructure team to accelerate building and scaling our innovative systems that support our growing identity platform. In this role, you will build the next-generation infrastructure that underpins all systems at CLEAR. The ideal candidate for this role will approach challenges with an eye toward reliability, simplicity, and scalability. What You'll Do: Develop and maintain a streamlined process for engineers to effortlessly build and deploy scalable and reliable software-defined networking solutions on AWS. Enhance our compute platform (Kubernetes) with new functionalities and features, focusing on AWS networking services and concepts such as VPCs, Route Tables, Security Groups (SGs), ALBs/ELBs, and Route53, optimize service communication and management. Collaborate across engineering teams to advocate for and implement best practices in observability, utilizing tools like Splunk or Datadog to ensure robust network monitoring. Act as a product owner for our infrastructure, collecting feedback and requirements from engineering teams to address pain points and develop solutions, particularly in the realm of AWS networking and cloud-native design principles. What you're great at: 0-2 years of experience in infrastructure and platform development and AWS cloud services. Proficient in Python, with understanding of Kubernetes and container orchestration tools like EKS and ECS. Understand AWS networking services, including VPC design, SGs, NATGWs, ALBs/ELBs, Rout
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a Software Engineer to join our First-Party Hardware team. In this role, you will design, build, integrate, and validate the software used to manufacture, qualify, and deliver our hardware from the factory. You will work across the stack to create the infrastructure that runs internally and externally to coordinate all aspects of the production process. You will create the critical tools and procedures to execute, capture, process, and present the data resulting from the end to end assembly and validation of our hardware across multiple vendors and sites. This role is hands-on and high-ownership. You will work closely across teams both internal and external to define the standards that will be used across our products to ensure the velocity and quality of our 1P hardware. You will own the implementation, deployment, and output of these systems as well their continued maintenance and SLAs. Location: San Francisco, CA (Hybrid: 3 days/week onsite). Relocation assistance available. In this role, you will: Design, develop, and maintain the software infrastructure for manufacturing process execution and data export. Own integration across internal customers and vendor systems and processes. Build and maintain the CI, release, and delivery pipeline of tooling to external partners. Build and maintain internal systems to ingest, process, deliver, and visualize critical data for internal teams and systems. Build system health monitoring, telemetry, remote d
From $299K/yr
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: This role will be based in San Francisco. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. You’ll build the systems that let any knowledge worker leverage fast, scalable databases without having to become a DBA. You’ll be a hands-on technical leader, helping set the architectural direction and roadmap as we take a new product from early alpha to general availability. You’ll work directly with early customers to shape foundational technical and product decisions. Then, together, we’ll work on scaling as adoption and workloads grow. This is a backend-leaning role with work spanning the stack. You might build the systems that provision and manage a fleet of customer Postgres instances, design how user-defined schemas evolve safely, or make complex queries execute efficiently. You’ll also follow those systems into the product: improving how a table loads, how users understand a slow operation, or how an agent safely works with their data. You’ll
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? The Data Infrastructure team at Cohere is responsible for the storage and data movement layer underlying every model training run. We're building the unified storage layer that feeds our training workloads. It needs to serve petabytes of training data and model checkpoints fast enough to keep thousands of GPUs busy across several training clusters. In this role, you’d have an opportunity to build this system from the ground up. You’d be a key contributor, working on a problem few teams have had to solve at this scale. In this role, you will: Design, build, and operate the distributed storage system that feeds model training and evaluation. Run this system multiple on Kubernetes clusters at petabyte scale. Work with researchers and training-infra teams on how jobs actually read and write data, and turn that into throughput, latency, and durability requirements Work through the networking, I/O, and consistency problems of moving large datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as the measures of success You may be a good fit if you have: Strong storage fundamentals,
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
About the Team The ChatGPT Search Product Infrastructure team builds the foundational systems that power search experiences across ChatGPT. We develop the product infrastructure that connects models with search systems and other sources of real-time information, enabling ChatGPT to deliver timely, relevant, and trustworthy answers to users around the world. Our work sits at the intersection of product engineering, AI, and large-scale infrastructure. We build shared platforms and abstractions that enable product teams to independently develop, evaluate, and launch new search-powered experiences. These platforms provide the guardrails, testing capabilities, observability, and rollout controls needed to prevent reliability, scalability, quality, and latency regressions while supporting rapid product iteration. The team partners closely with: Post-Training on model launches, experimentation, and prompt optimization Search product verticals on new user experiences Inference on GPU efficiencies Indexing and Retrieval on the systems that identify and deliver relevant information Capacity/Fleet team to ensure optimal regionalized provisioning of GPUs and CPUs About the Role We are looking for an Engineering Manager to lead the team responsible for ChatGPT’s Search Product Infrastructure. You will set the technical and organizational direction for the systems that bring search capabilities into ChatGPT. You will guide architectural decisions across search orchestration, model and prompt integration, serving infrastructure, experimentation, observability, evaluation, and product integrations. You will balance immediate launch and product needs with the long-term reliability, scalability, latency, and maintainability of the platform. A central responsibility of this role is creating leverage for Search product verticals. You will lead the development of extensible platforms that allow those teams to independently build, test, and launch features without requiring ongoing invol
We are hiring a Security Software Engineer to design and implement the hardware-backed security foundations used across OpenAI’s device ecosystem. A central focus of this role is hardening the boundary between our policy systems and the HSMs that protect sensitive cryptographic keys. This boundary determines which operations may be performed, what may be signed, which policies must be satisfied, and how changes to trusted software and policy are authorized. You will develop security-critical software and firmware within, or immediately adjacent to, an HSM trust boundary. Depending on your background, this may include HSM trusted applications, firmware services, cryptographic mechanisms, device drivers, PKCS#11 components, secure-provisioning protocols, or signing-policy enforcement systems. This is a hands-on software-engineering role. You will be expected to design systems, write and review production code, debug across hardware and software boundaries, and carry projects from initial requirements through deployment. It is not an HSM administration, PKI operations, compliance, or architecture-only position. In This Role, You Will Design and implement security-critical software and firmware for HSMs, secure elements, trusted execution environments, and hardware roots of trust. Build and harden the policy-to-HSM boundary responsible for authorizing certificate issuance and cryptographic signing operations. Develop HSM trusted applications, firmware components, host interfaces, device drivers, SDKs, or cryptographic service integrations. Implement or extend cryptographic interfaces such as PKCS#11, OpenSSL providers or engines, platform key-storage APIs, or comparable hardware-security interfaces. Build firmware and software that cryptographically enforces key generation, provisioning, usage, rotation, recovery, and destruction policies. Design and implement HSM-backed certificate authority, code-signing, key-management, and device-identity systems. Develop end-to-end
Other cities to consider
More places hiring for this role
Get new infrastructure engineer jobs in United States by email
Daily job updates · Unsubscribe anytime