Jobs in United States

Software Engineer Infrastructure in United States

2,095 active opportunities · Updated October 2026

Explore current software engineer infrastructure jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

Hiring demand

51/100

steady · 562 related jobs

Hiring trend

-79.1%

Job postings compared with the previous 30 days

Remote options

15.8%

Share of matching jobs listed as remote

Typical salary

$177.2K – $177.2K/yr

Based on 32 salary observations

R
📍 Foster City, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -94.8%

$177.2K – $211K/yr · Jobiba est.

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role Join our Enterprise Platform team and build the infrastructure foundations that enable the world's largest organizations to run Replit within their security and compliance boundaries. As a Software Engineer on this team, you'll design and implement the deployment flexibility, networking capabilities, authorization systems, and data controls that enterprises require, from single-tenant architectures and private connectivity to custom policy enforcement and customer-managed encryption. You'll work at the intersection of cloud infrastructure and enterprise requirements, partnering with Platform Engineering, Security, and Sales to ship capabilities that unlock adoption at demanding organizations. What You'll Do Build enterprise deployment infrastructure: Design and implement single-tenant and dedicated deployment options, enabling customers to run Replit with the isolation guarantees their security posture requires. Implement private networking capabilities: Build VPC peering, private connectivity, and static IP configurations that allow enterprises to integrate Replit into their existing network architectures. Design authorization services: Build the authorization infrastructure that enforces custom enterprise policies; enabling fine-grained access controls, custom permission models, and policy enforcement that integrates with customers' existing identity and governance systems. Ship data protection features: Implement bring-your-own-key (BYOK) encryption, customer-managed keys, and data residency controls that give enterprises ownership over their most sensitive data. Develop infrastructure automation: Write Terraform modules and automation that enable reliable, repeatable enterprise deployments across reg

TypeScriptPythonGCPKubernetes
S
📍 United States· Full-time· Remote
✓ Quality checkedDemand 51/100Company trend -95.9%

$177.2K – $211K/yr · Jobiba est.

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team In this role, you would join Stripe's Vulnerability Management team, whose mission is to "Surface vulnerabilities at scale across Stripe." Our vision is to create a culture of continuous excellence in managing vulnerabilities. We want to enhance customer trust by giving users context about threats and vulnerabilities affecting Stripe's systems. We aim to be a key partner in Stripe's risk management by providing visibility into vulnerabilities across Stripe's products and services. What you’ll do As a Software Engineer focused on Vulnerability Management at Stripe, you will use your software engineering expertise to find and prioritize vulnerabilities in our systems. Working closely with engineers across the company, you will drive the timely remediation of discovered vulnerabilities, playing a key role in Stripe's overall security and risk strategy. In addition, you will continuously improve Stripe's security defenses by enhancing our vulnerability management processes and selecting effective scanning tools to uncover weaknesses. Your core responsibilities as a Vulnerability Management Software Engineer will involve building data- and agent-backed systems to surface vulnerabilities across Stripe. In addition, you will coordinate fixes to prevent exploits that could impact Stripe or our users and serve as an advisor on security risks. All of this requires collaborating cross-functionally to advocate for practices that strengthen the safety of

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -86.4%

$177.2K – $211K/yr · Jobiba est.

About the Team The Cloud Agents team builds product infrastructure for long-running agents in the cloud: orchestration, sandboxing and isolation, secure environment connectivity, secrets and identity, observability, reliability, and cost controls. These agents securely connect to diverse developer and customer environments and use tools to accomplish goals. We partner closely with product, research, and infrastructure teams to turn agentic capabilities into dependable platforms for OpenAI products and developers building on OpenAI. About the Role We are looking for an experienced software engineer to help build and scale our cloud agent platform. You will design and operate systems for orchestrating agents at scale. You will work closely with product engineers on ChatGPT, API, and Codex to define the right abstractions and enable them to ship products quickly. Strong backend or infrastructure experience is important; experience with Python, Rust, distributed systems, cloud infrastructure, or product platforms is especially helpful. In this role, you will: Design and scale the orchestration, sandboxing and storage systems that run agentic workloads for Codex, ChatGPT, and the OpenAI API. Partner with product engineers to build a platform that enables them to ship quickly and turn feedback into robust abstractions. Improve reliability, security, performance, and cost efficiency for long-running agents. Deploy services that can operate across different environments and clouds. Your background might look something like: 9+ years of professional engineering experience, excluding internships, in relevant roles at technology and product-driven companies. Experience leading large-scale backend, platform, or infrastructure projects from ambiguous problem statements to production systems. Proficiency in one or more backend languages such as Python, Go, Rust, TypeScript, or similar, and the ability to move across service, platform, and product boundaries. Strong understanding

TypeScriptPythonAWSRest
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -86.4%

$177.2K – $211K/yr · Jobiba est.

About the Team OpenAI’s Platform and Infrastructure Engineering organization advances the mission of deploying artificial general intelligence (AGI) for the benefit of all by delivering secure, scalable, and resilient technology solutions. Our team builds and maintains robust infrastructure that safeguards OpenAI’s data and systems while ensuring employees are well-equipped and seamlessly connected. By prioritizing security, reliability, and user-centric solutions, we empower OpenAI employees to drive impactful AI research, corporate operations, and product innovation. About the Role As a Software Engineer: Internal Applications, Enterprise, you will build internal products that make technology support and administration safer, faster, and less dependent on manual intervention. You will help reduce reliance on broadly privileged human actions, turn recurring technology problems into paved paths, and build agentic systems that can help resolve tickets end to end. A core part of the role is building the interfaces that bring employees, AI agents, and human responders together in a shared ITSM experience, with the right context, controls, and handoffs at each step. We are seeking engineers who enjoy working across frontend and backend layers on ambiguous, high-leverage enterprise problems. You should bring strong product judgment, solid backend engineering fundamentals, and an interest in building software that changes how technology support, system administration, and agent-assisted operations are delivered. The best fit will care as much about the quality of the operator and employee experience as the correctness of the backend systems behind it. In this role, you will: Build frontend experiences that let employees request help, let agents gather context and take safe actions, and let human responders review, approve, or take over without losing the thread. Reduce reliance on broadly privileged manual actions by replacing them with narrow, auditable, policy-aware aut

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -86.4%

$177.2K – $211K/yr · Jobiba est.

About the Team API Enterprise Controls is part of the API Infrastructure organization and owns the platform capabilities that help developers, startups, and enterprises adopt the OpenAI API securely and confidently. We build the systems underneath our APIs and developer platform across authentication and identity, service accounts and key management, secure networking, compliance, auditability, observability, and operational controls. Our users are developers and teams running critical applications on OpenAI, and we partner closely with Product, go-to-market, security, and infrastructure teams to turn their most important needs into reliable, intuitive platform capabilities. About the Role We are looking for an exceptional backend software engineer to help define and ship the enterprise capabilities our API Platform needs to scale.; this is a product-engineering role grounded in deep backend systems. You will work across databases, streaming systems, request routing, authentication, and developer-facing APIs while bringing strong product judgment, developer empathy, and attention to the small details that make a platform easier to understand, trust, and operate. You will lead large cross-functional initiatives, work closely with Product and go-to-market teams, engage directly with sophisticated users, and carry ambiguous needs from discovery through design, launch, and iteration. In this role, you will: Own backend product capabilities end to end across authentication and identity, service accounts and key controls, secure networking, compliance, observability, and operational workflows. Partner with Product, go-to-market, security, infrastructure teams, and sophisticated customers to identify needs, shape the roadmap, and lead large cross-functional projects from design through launch. Design developer-facing APIs, system behavior, configuration, error handling, safe defaults, auditing, and notifications with exceptional care for the details that define a great dev

TypeScriptPythonAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -86.4%

$177.2K – $211K/yr · Jobiba est.

About the Team API Multimodal builds the developer-facing products and infrastructure that bring OpenAI’s image, audio, and real-time model capabilities into the world. We are responsible for high-scale APIs for image generation, speech transcription, speech generation, and low-latency voice interactions. We partner closely with Research and Inference to bring frontier model capabilities to developers and use customer feedback to improve our models. About the Role As a software engineer on API Multimodal, you will build and operate the products and distributed systems behind OpenAI’s image, audio, and real-time APIs. You will work across model integration, API design, and production infrastructure to turn new research capabilities into reliable developer experiences. This hands-on role combines backend and systems depth with product judgment: you will own projects end to end, partner with Research, Inference, and Safety, and help make multimodal AI useful at scale. Model training experience is not required. In this role, you will: Design, build, and ship developer-facing APIs and backend services that serve frontier models. Architect low-latency streaming, request, session, and model integration systems that make complex multimodal interactions reliable and intuitive at scale. Work directly with Research to bring new model capabilities into production, shape the systems around them, and incorporate feedback from real-world developers and customers. Own the availability, latency, scalability, and cost efficiency of the services you build. Own projects from technical design and implementation through launch and ongoing iteration, while raising the team’s engineering standards. Your background might look something like: 7+ years of professional experience, excluding internships, in backend, infrastructure, platform, or product engineering roles. A track record of designing, building, and operating production backend services, developer-facing APIs, or distributed syste

TypeScriptPythonAWSRest
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -86.4%

$177.2K – $211K/yr · Jobiba est.

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software is reliable, testable, and ready to ship. We design and maintain automated test frameworks, hardware-in-the-loop labs, and release pipelines that keep quality signals trustworthy and enable rapid, safe product launches. Our work spans developer tools, automation, systems integration, and cross-team collaboration to ensure every release meets the highest standards. About the Role As a Software Engineer, Quality and Developer Tools , you will build and own the systems that validate our device software—from test frameworks and regression infrastructure to hardware-in-the-loop labs and release gates. You’ll design the tooling and automation that keep quality signals trustworthy, integrate them into CI/CD, and make it easy for engineers and QA vendor technicians to execute reliable, repeatable workflows. We’re looking for engineers with deep experience in software quality, automation, developer tooling, and hardware-software integration who thrive on building scalable, reliable systems for validation and release readiness. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Test infrastructure & frameworks: Design, implement, and maintain a unified test framework for device software across unit, integration, system, and end-to-end testing, with reproducible runs and integrations with GitHub, Linear, and Slack. CI/CD integration & releases: Integrate test suites with Buildkite, enforce promotion criteria for staging and production, auto-file regressions, and publish traceable artifacts and release notes. Hardware-in-the-loop lab design & orchestration: Plan and bring up racks, power and networking systems, and orchestration for device testing; support automated flashing, provisioning

PythonAWSCI/CDGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -86.4%

$177.2K – $211K/yr · Jobiba est.

About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,

AWSKubernetesLinuxRest
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -94.8%

$177.2K – $211K/yr · Jobiba est.

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role Join our Enterprise Platform team and build the infrastructure foundations that enable the world's largest organizations to run Replit within their security and compliance boundaries. As a Software Engineer on this team, you'll design and implement the deployment flexibility, networking capabilities, authorization systems, and data controls that enterprises require, from single-tenant architectures and private connectivity to custom policy enforcement and customer-managed encryption. You'll work at the intersection of cloud infrastructure and enterprise requirements, partnering with Platform Engineering, Security, and Sales to ship capabilities that unlock adoption at demanding organizations. What You'll Do Build enterprise deployment infrastructure: Design and implement single-tenant and dedicated deployment options, enabling customers to run Replit with the isolation guarantees their security posture requires. Implement private networking capabilities: Build VPC peering, private connectivity, and static IP configurations that allow enterprises to integrate Replit into their existing network architectures. Design authorization services: Build the authorization infrastructure that enforces custom enterprise policies; enabling fine-grained access controls, custom permission models, and policy enforcement that integrates with customers' existing identity and governance systems. Ship data protection features: Implement bring-your-own-key (BYOK) encryption, customer-managed keys, and data residency controls that give enterprises ownership over their most sensitive data. Develop infrastructure automation: Write Terraform modules and automation that enable reliable, repeatable enterprise deployments across reg

TypeScriptPythonGCPKubernetes
M
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -100%

What we're building Mutiny is the self-improving AI infrastructure for GTM teams to execute faster and close more revenue. Our ambition is to do for revenue velocity what Cursor and Claude Code did for engineering velocity. With Mutiny, everyone in sales and marketing gets a bench of GTM athletes that handle any work across their revenue motion and learn from what's actually moved their deals. In April we re-launched the product as an agent-first platform. Anthropic showcased us as a leader in AI GTM. MRR is growing more than 70% month-over-month, with customers like Uber, Rippling, and Snowflake. We're backed by Sequoia, YC, and Insight, and we're building a generational company. The opportunity Most engineers spend their career making predictable systems faster. You'll spend yours making non-deterministic ones trustworthy. As a senior engineer on our AI product team, you'll architect the Campaign Builder and Agent experiences marketers and sellers open every day to go from idea to personalized assets in minutes. You'll partner directly with product, design, and the founders to define what an agent-first GTM platform should feel like, and your calls on architecture, evals, and guardrails compound across thousands of customer accounts. This role is in person in New York City, five days a week, and we ship weekly. What you'll own The core agent surfaces. Architect and ship the Campaign Builder and Agent experiences end-to-end. Frontend, backend, prompts, evals, the whole stack. Reliability on top of LLMs. Make non-deterministic models feel deterministic at the surface. Build the retries, fallbacks, and orchestration so the customer never sees the failure mode. Evals and guardrails. Define how we measure quality, catch regressions, and keep brand and tone consistent across thousands of customer accounts. Speed and feel. AI products live or die by latency and the loop between intent and output. You'll obsess over both, and use coding agents and agent networks to ship f

TypeScriptPythonAIKotlin
P
📍 New York, New York, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -85.7%

$177.2K – $211K/yr · Jobiba est.

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: The Experience team is at the center of one of the most exciting transitions in software development history — the shift from human-driven to agent-driven product experiences. We own Pinecone's API, clients, authentication, revenue, and observability systems, and right now that means redesigning all of it for a world where AI agents are first-class users alongside humans. This is a wide-scope role. You'll own things end-to-end — from backend architecture to API design to SDK and web surfaces. You’ll be working closely with product, design, and other engineering teams to identify user needs and build the right thing, at the right abstraction level, at the right time. Along the way, you will be building high-leverage platform capabilities that accelerate Pinecone’s product development and user growth systems. We're looking for an engineer who sees this moment for what it is: a rare opportunity to shape how developers and agents interact with a category-defining product. You're not waiting to see how the industry figures out MCP, agentic workflows, and AI-native interfaces — you're already experimenting, already forming opinions, already building. You know that speed and leverage matter more than labor, and you've internalized AI-assisted development not as a productivity trick but as a fundamentally different way of working. Responsibilities: Pioneer our agent experience. Shape how AI agents interact with Pinecone — designing interfaces, protocols (MCP), and tooling that make Pinecone the easiest and most capable platform f

JavaReactAWSAzure
P
📍 United States· Full-time
✓ Quality checkedDemand 51/100Company trend -85.7%

$177.2K – $211K/yr · Jobiba est.

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: The Experience team is at the center of one of the most exciting transitions in software development history — the shift from human-driven to agent-driven product experiences. We own Pinecone's API, clients, authentication, revenue, and observability systems, and right now that means redesigning all of it for a world where AI agents are first-class users alongside humans. This is a wide-scope role. You'll own things end-to-end — from backend architecture to API design to SDK and web surfaces. You’ll be working closely with product, design, and other engineering teams to identify user needs and build the right thing, at the right abstraction level, at the right time. Along the way, you will be building high-leverage platform capabilities that accelerate Pinecone’s product development and user growth systems. We're looking for an engineer who sees this moment for what it is: a rare opportunity to shape how developers and agents interact with a category-defining product. You're not waiting to see how the industry figures out MCP, agentic workflows, and AI-native interfaces — you're already experimenting, already forming opinions, already building. You know that speed and leverage matter more than labor, and you've internalized AI-assisted development not as a productivity trick but as a fundamentally different way of working. Responsibilities: Pioneer our agent experience. Shape how AI agents interact with Pinecone — designing interfaces, protocols (MCP), and tooling that make Pinecone the easiest and most capable platform f

JavaReactAWSAzure
P
📍 New York, New York, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -85.7%

$177.2K – $211K/yr · Jobiba est.

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: Join a team that builds robust, real-time distributed systems for a cutting-edge database. We care about performance, reliability, scalability, and most of all learning and having fun together. Whether you’re a seasoned coder or just getting started, if you’re passionate about technology and eager to learn, you’ll fit right in. Who we are: We show up to work, ready to collaborate and build technologies that make a difference, with people who genuinely care. We chase improvements such as tail latencies, bytes throughput, cache hit rate, and operational cost efficiency. We believe learning is ongoing and that even the most complex problems can have simple solutions. What You’ll Do: Collaborate with teammates to design and build database features that power AI applications. Learn how to tune performance and support reliability in distributed systems (don’t worry, we’ll guide you). Help Pinecone run smoothly on popular cloud providers. Take ownership of your work and grow your skills every day. Have fun. Who You Are: 5+ years of work experience - programming in Rust, Go, C++, or a comparable language. You’re genuinely curious about distributed systems and eager to dive deep into technical challenges. You approach problems with creativity and persistence, and you’re comfortable asking thoughtful questions or seeking feedback. You’re excited to learn, value constructive feedback, and appreciate mentorship. Bonus Points: You have hands-on experience with cloud platforms (AWS, GCP, Azure) or have demonstrated an ability to pick u

AWSAzureGCPAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -86.4%

$177.2K – $211K/yr · Jobiba est.

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

PythonAWSLinuxRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -86.4%

$177.2K – $211K/yr · Jobiba est.

About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing

PythonAWSKubernetesRest

Related career options

Similar roles with stronger pay

Client Service Associate

Demand 46/100 · 8 jobs

$840K – $840K/yr

Salary →
Director of Product

Demand 43/100 · 6 jobs

$382.5K – $382.5K/yr

Salary →
Physical Design Engineer

Demand 43/100 · 8 jobs

$300K – $300K/yr

Salary →
Sr. Engineer

Demand 42/100 · 7 jobs

$300K – $300K/yr

Salary →
Senior Director

Demand 38/100 · 30 jobs

$278.9K – $278.9K/yr

Salary →
Senior Product Designer

Demand 30/100 · 11 jobs

$255.7K – $255.7K/yr

Salary →
🔔

Get new software engineer infrastructure jobs in United States by email

Daily job updates · Unsubscribe anytime