About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Industrial Compute team is central to this mission, setting the core infra strategy and implementing this vision. From site selection to the buildout process, this team sits at the intersection of commercial, technical, strategy, and operations, interacting with teams and executives inside and outside of OpenAI. About the Role The Clean Energy and New Technology Lead will own infrastructure clean energy and emerging energy technology strategy and execution, identifying and deploying scalable solutions that enable resilient, low-carbon compute and data center growth. The role will work closely with regulatory and policy teams to align infrastructure expansion with OpenAI’s long-term environmental and operational objectives. This is an individual contributor lead role and does not have direct reports initially. The role will evaluate where emerging energy technologies can materially improve reliability, cost, carbon, speed, or resilience; translate those options into practical deployment pathways; and help ensure OpenAI’s infrastructure growth remains aligned with sustainability considerations. Key Responsibilities Evaluate emerging energy solutions such as clean firm power, advanced storage, grid flexibility, low-carbon backup power, heat reuse, water-related energy efficiency, and other scalable technologies where relevant. Identify pilot opportunities and deployment pathways that can move promising energy technologies from concept to commercially and operationally credible execution. Translate technical options into clear reliability, cost, schedule, carbon, regulatory, and operational implications for infrastructure decision-making. Partner with energy regulatory, policy, procurement, engineering, deployment, finance, legal, and site-readiness teams to align technology and sustainability choices with
Jobs in United States
Storage And Backup Engineer in United States
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current storage and backup engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems underpin data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences. As frontier models and agents evolve how they use memory, history and snapshots, the underlying architecture increasingly shapes the capabilities products can deliver—and their latency, reliability, cost and efficiency. About the Role We are looking for a technically deep TPM to independently define and lead multiple programs across data platforms, online databases and storage infrastructure. You will connect model, product and data-consumer requirements to architecture, and work with the relevant engineering teams to take new capabilities through production adoption and repeatable expansion. The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The challenge is not simply forecasting more resources: it is making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment and validation. A data pipeline, database query, file operation or execution snapshot can affect whether a product or agent succeeds; you will connect those outcomes to the systems underneath. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate model, product and data-platform needs into precise access patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency. Partner with engineering to transform data and storage architecture into repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together.
From $345K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on Creator Services Data, you’ll be leading the company’s efforts to build the next generation Data Storage systems to power the millions of experiences on the Roblox Platform. We run the mission critical cloud services, Data Stores , Memory Stores , and Badges , which are crucial for storing game state such as inventory and scores, implementing leaderboards, server lists, and trading, and tracking player progress and achievements. Our team is also responsible for building dashboards to provide insights to Creators using cloud services including Client/Server Performance , Data Stores , and Memory Stores . Finally, our team owns the Roblox Extended Services platform, which provides the capability for large experiences to purchase additional resources for existing services like Data Stores and new services built around compute and generative AI. At its core, this team is focused on solving complex back end distributed systems and storage problems at scale. However, our scope extends to full stack projects spanning all the way from the infrastructure layer, through data storage and data pipelines, microservices, telemetry, game servers,
From $192K/yr
Coordination Systems provides foundational distributed systems building blocks for internal Datadog platforms. Our services cover sharding, consensus, resource protection, configuration distribution, and much more. We are looking for a manager to lead the Coordination Systems - Storage team. This team provides essential configuration storage and distribution systems that are depended upon by almost every service and pod at Datadog. We power critical runtime configuration (e.g. feature flags), complex control planes (e.g. dynamic sharding configuration), and much more. Storage is one of four subteams within Coordination Systems. If successful, the candidate will have opportunities to lead other growing and impactful areas such as Resource Protection. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: (Describe role responsibilities here/max 6 bullets) Lead a core team of 5 engineers (distributed, with majority in NYC) Lead ceremonies, prioritize and delegate project Stay hands-on with the code, e.g. isolated features, small remediations, investigation follow ups Stay actively involved in operations, incidents, root cause analysis, etc. Constantly promote a culture of operational excellence, organizing gamedays, conducting operational reviews, staying proactive with reliability Who You Are: (Describe role qualifications here/max 6 bullets) Strong distributed systems skills, able to understand and account for a variety of failure modes, well-versed in end-to-end o11y, validation testing, simulation setup, etc. Worked on platform teams before, providing critical infrastructure to internal stakeholders Experienced in handling significant incidents, both as a responder and follow-up ow
From $177.2K/yr
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . We’re looking for a Staff Software Engineer to help build the next generation of Pinterest’s big data storage platform. You’ll work on some of the most exciting big data open source technologies — especially Apache Iceberg — at exabyte scale to power the data infrastructure that helps Pinners discover and do what they love. As a Staff Software Engineer, you’ll serve as a technical leader and hands-on contributor, designing and building highly scalable storage systems for Pinterest’s data lake. You’ll partner closely with teams across data, ML/AI, analytics, and infrastructure to evolve our storage and metadata management capabilities, enabling efficient, reliable, and governed access to data at massive scale. What you’ll do: Design, implement, and optimize Pinterest’s exabyte-scale data lake storage platform. Lead complex technical projects and
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Events Analytics Platform (EAP) team is responsible for the infrastructure that powers all of Sentry's time-series data and searching capabilities across billions of events with sub-second latency. We started this initiative by building Snuba, the primary storage and query service for Sentry's event data powered by ClickHouse, and we are now focused on unlocking deeper visibility and reporting across the terabytes of event data our users generate. As a Senior Software Engineer, you will lead efforts to push the boundaries of data visibility at Sentry. You will do this by expanding the capabilities of our search infrastructure, building new capabilities on top of our state-of-the-art storage layer and increasing the performance and integrity of Sentry’s core data services. You will also help shape Infrastructure's technical direction at Sentry and collaborate with Product and other Engineering teams to turn that vision into a reality. If you want to solve the hard problems that come with scaling event data into the petabyte range, this could be the job for you. In this role you will: Expand EAP's ability to deliver data at world-class speed and reliability. Architect and automate services and systems to scale reliably under growing demand. Make architectural trade-offs that balance product requirements with engineering constraints. Maintain and grow the team's code quality initiatives by regularly reviewing code and contributing to design decisions. Lead design and discussions around deliverables the team is working towards. Improve the maintainability and developer experience of the codebases EAP owns. Exa
About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua
$220K – $450K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides developer-first observability to over 4 million developers worldwide. The Events Analytics Platform (EAP) team is at the heart of that mission: it powers how all of Sentry's event data, such as errors, transactions, spans, profiles, replays, and metrics, is stored, queried, and analyzed. It also powers Sentry's latest AI push, Seer. The EAP team makes it possible for developers to efficiently search and debug across massive volumes of data, providing the context needed to understand and fix issues quickly. This team is also a cornerstone of Sentry's long-term strategy to become a context assembly and telemetry platform that unifies different signals so developers can see the complete picture. As an engineering manager on the EAP team, you will lead a group of engineers building and scaling one of Sentry's most critical data platforms. You will be responsible for driving architectural evolution, ensuring system stability, and mentoring a talented team. This is a highly visible leadership role with direct ties to Sentry's long-term product and platform strategy. What you'll do Grow and develop a team of engineers with high expectations for ownership and impact Set the technical and strategic direction for the team, balancing short-term stability with long-term architectural evolution Drive development of core EAP features, including support for complex analytical queries, dynamic routing logic across fidelity levels, storage and compute separation, and modern patterns for analytical storage Ensure EAP can support the workload demands of AI agents and MCP servers that unlock new AI capabilities fo
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As an AI Accelerator Systems Software Technical Program manager at OpenAI, you will help bring our chips/system hardware roadmap to life, navigating an array of technical and partnership challenges. We’re looking for people excited to push the frontiers of computing by navigating technical explorations and are passionate about building. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Manage the end-to-end software development from design to implementation for our AI acceleration systems, working across technical, cross-functional and external stakeholders Lead planning and scheduling of AI system software designs with our strategic partners and vendors Coordinate and lead internal resources and communication for efficient interaction with partners and vendors. You might thrive in this role if you: Have experience as a software technical program manager for data center system products (server, GPU, TPU, networking, storage and so on) taking products from concept to volume in a data center environment ensuring the systems scale with high quality Know end-to-end software development program management techniques from concept, design, production, deployment into the data center Want to help design some of the world’s largest supercomputing systems, working at the edge of complex hardware challenges Enjoy working with and enabling world-clas
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a Hardware Chips Programs Manager at OpenAI, you will help bring our chips hardware roadmap to life, navigating an array of technical and partnership challenges. We’re looking for people excited to push the frontiers of computing by navigating technical explorations and are passionate about building. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Manage the design and implementation planning of our ML acceleration hardware, working across technical, cross-functional and external stakeholders Lead planning and scheduling of chip hardware designs with our strategic partners and vendors Coordinate and marshal internal resources and communication for efficient interaction with partners and vendors. You might thrive in this role if you: Have experience as a technical program manager for data center hardware products (server, GPU, TPU, networking, storage and so on) Know the whole end-to-end system program management from concept, design, production, deployment into the data center Have some experience with System SW programs through NPI Want to help design some of the world’s largest supercomputing systems, working at the edge of complex hardware challenges Enjoy working with and enabling world-class AI Researchers and Engineers Are passionate about the technical program function, and enjoy independently owning and delivering on your tea
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron's Global Supplier Quality organization is seeking a Controller Quality Principal Engineer to lead the quality strategy, qualification, and continuous improvement of storage and memory controller (ASIC/SoC) manufacturers supporting Micron's SSD, embedded, and storage solutions portfolio. This is a senior technical leadership role responsible for driving controller supplier quality performance from design qualification through mass production and field support The successful candidate will collaborate across functions with ASIC Development Team, Compose Engineering, Product Engineering, Dependability, Manufacturing, and Commodity Management, as well as directly with controller IC vendors, foundries, third party reliability labs and OSAT (outsourced assembly and test) partners, to ensure controller quality, reliability, and supply continuity meet Micron's standards. This role can be based in Taiwan, Hyderabad, or San Jose and will work extensively across time zones with global partners and suppliers! Develops, evaluates, revises, and applies technical quality assurance protocols/methods to inspect and test in-process raw materials, production equipment, and finished products. Ensures activities and items are in compliance with both company quality assurance standards and applicable government regulations. Performs analysis and identifies trends in the inspection of finished products, in-process materials and bulk raw materials, and recommends corrective actions when vital. Ensures that established manufacturing inspection, sampling and statisti
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. In this role, you will: Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Collaborate with engineering and security teams to drive deployment of security enhancements and control changes across broad-scale infrastructure. Tackle high-impact projects such as checkpoint encryption, network isolation, secret management, and machine identity, while continuously raising the security bar for emerging AI workloads. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets to adapt to evolving challenges. You will thrive in this role if you have: Deep understanding of security principles, best practices, and common vulnerabilities. A proactive mindset, with the ability to identify and address secu
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Principal Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. As a principal engineer, you will set technical direction and drive execution on high-impact infrastructure security programs, partnering across various orgs at OpenAI to deliver durable controls that raise the security bar at OpenAI scale. In this role, you will: Own end-to-end security outcomes for one or more critical infrastructure areas, including multi-quarter strategy, roadmap, and delivery. Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Lead cross-functional programs to deploy security enhancements and control changes across broad-scale infrastructure, balancing security guarantees with reliability and velocity. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? The Data Infrastructure team at Cohere is responsible for the storage and data movement layer underlying every model training run. We're building the unified storage layer that feeds our training workloads. It needs to serve petabytes of training data and model checkpoints fast enough to keep thousands of GPUs busy across several training clusters. In this role, you’d have an opportunity to build this system from the ground up. You’d be a key contributor, working on a problem few teams have had to solve at this scale. In this role, you will: Design, build, and operate the distributed storage system that feeds model training and evaluation. Run this system multiple on Kubernetes clusters at petabyte scale. Work with researchers and training-infra teams on how jobs actually read and write data, and turn that into throughput, latency, and durability requirements Work through the networking, I/O, and consistency problems of moving large datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as the measures of success You may be a good fit if you have: Strong storage fundamentals,
From $285.5K/yr
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . As a principal engineer on the Online Systems team, you’ll join a team that powers Pinterest’s most business-critical online systems at massive scale, driving the reliability, efficiency, and evolution behind every core Pinner and Advertiser experience. You'll lead major efforts like multi-region deployment and Kubernetes migration, set the standard for operational excellence, and define the long-term vision for our online serving infrastructure, supporting machine learning and product innovation across the company. This is an opportunity for high-impact technical leadership, broad visibility, and cross-functional influence at the heart of Pinterest’s platform. What you’ll do: Improve reliability, scalability and infra efficiency for Pinterest’s critical online systems across storage and caching, online service and realtime analytics syste
Other cities to consider
More places hiring for this role
Get new storage and backup engineer jobs in United States by email
Daily job updates · Unsubscribe anytime