Jobs in United States

Edge Infrastructure Engineer in United States

739 active opportunities · Updated October 2026

Explore current edge infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team The Platform Systems team at OpenAI operates at the intersection of cutting-edge AI and large-scale distributed systems. We build the engineering and research infrastructure required to train OpenAI’s flagship models on some of the world’s largest, custom-built supercomputers. Our team develops core model training software and works deep in the stack - spanning collective communication, compute efficiency, parallelism strategies, fault tolerance, failure detection, and observability. The systems we build are foundational to OpenAI’s research velocity, enabling reliable, efficient training at frontier scale. We collaborate closely with researchers across the organization, continuously incorporating learnings from across OpenAI into the evolution of our training platform. About the Role As a Software Engineer, Platform Systems, you will design and build distributed systems that provide visibility into large-scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical to operating OpenAI’s training stack and is actively evolving to support new use cases and increasingly complex workloads. This role sits at the core of our training infrastructure, blending systems engineering, performance analysis, and large-scale debugging. In This Role, You Will Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high-throughput distributed systems Collaborate with systems, infrastructure, and research teams to evolve platform

AWSRestAIRust
R
📍 Goodyear, Arizona, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $126.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Data Center Engineer , you'll help us scale our Core/Edge Data Centers and hardware infrastructure at a time of incredible growth for our business. At Roblox, you'll have boundless opportunities to shape the future of the Imagination Platform™ and demonstrate your passion for delivering thoughtful solutions in front of a global audience. If you know what it takes to build and operate hardware infrastructure that can sustain millions of concurrent players year-round and you take play as seriously as we do, you'll fit right into our highly experienced and ever-expanding engineering team. You will report to the Technical Lead Data Center Engineer. You will: Develop and maintain the Core/Edge Data Center and hardware infrastructure to meet the large scale and real-time requirements of our Imagination Platform™ to ensure our community has an awesome experience anywhere in the world. This includes all aspects of the server, network infrastructure, power, and environmental life cycles. Own efforts to track and mitigate systemic issues preventing hosts from returning to service. Identify and solve critical problems and prevent them from re-occurring via root cause analysis and giving rec

AWSGitAIGo
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -97.2%

From $127K/yr

Quick readStrong listing-quality and freshness signals

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim

MongoDBAWSAzureGCP
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

AWSKubernetesCI/CDLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

PythonAWSLinuxRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing

PythonAWSKubernetesRest
I
📍 Arizona, Phoenix, United States
✓ Quality checkedCompany trend +116%

Job Details: Job Description: Intel is shaping the future of technology to help create a better future for the entire world. Our work in pushing forward fields like AI, analytics, and cloud-to-edge technology is at the heart of countless innovations. With a career at Intel, you'll have the opportunity to use technology to power major breakthroughs and create enhancements that improve our everyday quality of life. Join us and help make the future more wonderful for everyone. Want to learn more? Visit our YouTube Channel or the link below. Life at Intel The Role and Impact As a System Software Architect, you will play a pivotal role in designing and developing innovative software solutions that drive efficiency and scalability across Intel's manufacturing processes. You will be responsible for architecting robust systems and frameworks that enhance operational performance while ensuring seamless integration with existing technologies. Your contributions will directly impact Intel's ability to deliver world-class semiconductor products efficiently and reliably. Business Group Intel Foundry Services is at the forefront of manufacturing excellence, delivering cutting-edge solutions to meet the complex demands of the semiconductor industry. The group focuses on pioneering advancements in manufacturing technologies, infrastructure optimization, and secure operations, enabling Intel to stay ahead in a fast-evolving landscape. By joining this team, you will contribute to Intel's broader mission of solving technological challenges and empowering innovation worldwide. Key Responsibilities: - Architect scalable and efficient system sof

PythonJavaAIRecruitment
T
📍 Boston, Massachusetts, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a SoC Design Verification Engineer to validate the System Management Controller (SMC) and enable seamless multi-chip integration. In this role, you will design and execute tests, build infrastructure, and debug issues across chiplet-based SoCs. You’ll have the opportunity to work with remote mentorship while contributing to the foundation of scalable multi-die systems. This role is hybrid, based out of Toronto, Ontario, Boston, MA or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Proficient in SystemVerilog, SV-UVM, Python, and C/C++ with strong verification skills. Experienced in writing test plans, building infrastructure, and debugging hardware/software flows. Comfortable working with remote mentorship and distributed teams. Familiar with AI-assisted tools like Copilot, Cursor, and Claude to accelerate verification. What We Need Develop and maintain SMC tests and supporting DV infrastructure. Write, execute, and track test plans for chiplet and multi-chip SoC designs. Use C/C++ to develop tests compiled, loaded, and executed directly on the DUT. Triage, analyze, and debug issues in clos

PythonAWSAIC++
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -90.6%

From $187K/yr

Quick readStrong listing-quality and freshness signals

As a Cloud Security Engineer you will partner with different stakeholders across the organization to secure our cloud infrastructure. As part of the Platform Security organization we secure the building blocks of Datadog’s applications and infrastructure. We do this by building solutions to solve systemic risks and combine an approach of making the secure path easier and the insecure path harder to secure and accelerate the business. We regularly partner with the most bleeding edge internal products and are working to solve and build solutions to enable our safe usage of AI. We also develop AI based solutions to enable security at scale. We are looking for a Service Mesh and Kubernetes focused security specialist to help round out an incredibly strong infrastructure security focused group. You will rotate through a variety of internal projects and gain deep exposure to Datadog’s infrastructure. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Solve our most challenging cloud infrastructure security problems starting with our core building blocks and golden paths. Enable our engineers to build and ship secure solutions quickly. Build and extend Datadog’s Platform Security solutions. Leverage and influence the direction of Datadog’s products to secure our infrastructure, and provide internal feedback that enables our teams to improve the products for ourselves and our customers. Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or related scientific field or equivalent professional experience. Passionate about advocating for and implementing solutions to complex problems, at-scale, in a large multi-cloud environment. You don’t want to just provide security recommendations, you want to help imple

PythonAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team The Applied Foundations team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale. The Applied Foundations team is at the front lines of defending against financial abuse, scaled attacks, and other forms of misuse that could undermine the user experience or harm our operational stability. Integrity Foundations provides the core building blocks and infrastructure for this work. About the Role At OpenAI, our mission is to advance AI in a way that is safe, reliable, and aligned with broad societal values. The applied foundations role is crucial for maintaining the trustworthiness of our platforms. You will be pivotal in developing robust defenses against a spectrum of adversarial behaviors that threaten our ecosystem. In this role, you'll work with our entire engineering team to design and implement systems that detect and prevent abuse, promote user safety, and reduce risk across our platform. You'll be at the forefront of our efforts to ensure that the immense potential of AI is harnessed in a responsible and sustainable manner. In this role, you will: Develop and enhance systems to detect and prevent various forms of abuse including financial fraud, botting, and scripting. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. Assist with response to active incidents on the platform and build new tooling and infrastructure that address the fundamental problems. You might thrive in this role if you: Have at least 3 years of professional software engineering experience. Have experience setting up and maintaining production backend services and data pipelines. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Are self-directed

PythonAWSAzureKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team OpenAI’s Applications Engineering organization builds and operates the products (such as ChatGPT & Codex) that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role We’re hiring Backend Software Engineers to design and implement safe services and infrastructure that power our core products. What You’ll Do Architect, build, and improve scalable backend systems and APIs. Drive performance, reliability, and safety across distributed services. Implement data storage, retrieval, compute, and integration solutions. Participate in long-term architectural planning and technical design reviews. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. You Might Thrive Here If You: Have strong experience with distributed systems, APIs, and backend languages (e.g., Go, Python, Rust, C++). Have experience setting up and maintaining production backend services and data pipelines. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Enjoy building resilient services that handle large scale and complexity. Are self-directed and enjoy figuring out the best way to solve a particular problem Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team OpenAI’s Applications Engineering organization builds and operates the products that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role You will be a Senior Android engineer on OpenAI’s Applied Foundations team, building the core mobile experiences that power how users sign up, manage their account, family features, pay for services, stay safe, and interact with OpenAI’s products with confidence. This role is about creating high-quality products as well as reusable Android foundations that product teams across different OpenAI apps depend on to ship quickly while meeting the highest standards for security, reliability, and user trust. You’ll own complex client-side systems spanning UI, networking, local state, payment integrations and Apple platform integrations, and work closely with backend, product, and safety partners to shape the architecture that supports OpenAI’s mobile ecosystem at global scale. You might thrive in this role if you: Have 4+ years of professional software engineering experience. Have a proven track record of building high-quality Android applications in production. Are fluent in Kotlin (and/or Java) and familiar with Android development tools and architecture components. Prioritize performance, security, and user experience in mobile development. Enjoy working cross-functionally to bring ambitious product ideas to life. Care deeply about performance, security, and user experience. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We

JavaAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team OpenAI’s Applications Engineering organization builds and operates the products that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role You will be a Senior iOS engineer on OpenAI’s Applied Foundations team, building the core mobile experiences that power how users sign up, manage their account, family features, pay for services, stay safe, and interact with OpenAI’s products with confidence. This role is about creating high-quality products as well as reusable iOS foundations that product teams across different OpenAI apps depend on to ship quickly while meeting the highest standards for security, reliability, and user trust. You’ll own complex client-side systems spanning UI, networking, local state, payment integrations and Apple platform integrations, and work closely with backend, product, and safety partners to shape the architecture that supports OpenAI’s mobile ecosystem at global scale. In this role, you will: Build and ship new experiences on iOS that showcase the power of AI. Optimize app performance, reliability, and responsiveness at global scale. Design and maintain shared iOS frameworks and primitives for account, trust, and commerce flows that are used across OpenAI’s mobile apps. Establish robust testing frameworks and refine app architecture for long-term maintainability. Collaborate with product, design, research, and backend teams to deliver high-impact features. Provide technical leadership to shape the future of OpenAI’s iOS platform. You might thrive in this role if you: Have 4+ years of professional software engineering experience. Hav

AWSRestAISwift
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage distributed data pipelines, collaborate closely with researchers to translate requirements into robust systems, and harden pipelines that serve as the backbone for OpenAI's rapid iteration cycles. We’re looking for engineers who are detail-oriented, have strong experience with distributed systems, and excel at building reliable infrastructure in high-stakes environments. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain data infrastructure systems such as distributed compute, data orchestration, distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability, and security. Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient. Partner with researchers to deeply understand requirements and translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation. You might thrive in this role if you: Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data. Are detail-oriented and bring rigor to building and maintaining reliable systems. Demonstrate excellent software enginee

AWSRestMachine LearningAI
I
📍 Arizona, Phoenix, United States
✓ Quality checkedCompany trend +116%

Job Details: Job Description: As a Facilities Mechanical Engineer, you will play a pivotal role in ensuring the reliable operation and maintenance of Intel's advanced mechanical systems, supporting our cutting-edge manufacturing, cleanroom, and research and development activities. Your work will directly impact Intel's ability to innovate and deliver world-class technologies by maintaining the critical infrastructure that enables peak operational performance. This role offers a unique opportunity to drive engineering excellence, collaborate with cross-functional stakeholders, and shape the future of Intel's facilities globally. Key Responsibilities: Ensure the availability, reliability, and maintenance of mechanical systems such as oil-free air systems, HVAC systems, chilled water plants, boilers, environmental abatement systems, and compliance exhaust systems. Support daily tactical efforts to meet safety, reliability, and environmental targets for mechanical systems. Design and analyze mechanical systems and equipment, troubleshoot systematic issues, and provide evaluations, recommendations, and solutions to resolve problems. Develop engineering scopes of work for mechanical projects and conduct design reviews to ensure system integrity and performance. Conduct feasibility studies and testing on new and modified designs, while overseeing prototype fabrication and design testing. Partner with internal business units and operations teams to deliver engineering solutions tailored to their specific needs.

Project ManagementRecruitment
🔔

Get new edge infrastructure engineer jobs in United States by email

Daily job updates · Unsubscribe anytime