Datadog's Software Engineers with Systems depth leverage their experience with systems and tooling to build software that ensures Datadog remains reliable, performant, and secure. For this track, their Software Engineering experience may resemble the Distributed Systems track, but is typically applied in combination with their systems experience to build and run internal platforms and tools that our products are built on. These people typically have deep experience building and managing large cloud infrastructure deployments, or leading reliability efforts for orgs similar to ours, or building release machinery to allow hundreds or thousands of devs to do their jobs without stepping on each others' toes. The systems and tooling where they may have experience depth may include (but not limited to): bazel, build tooling, cassandra, CDN, chef, configuration management, container orchestration, consul, docker, elasticsearch envoy, haproxy, kafka, kubernetes, load balancing, network architecture, postgres, redis, release management, RPC frameworks, service discovery, spinnaker, terraform, zookeeper. Bonus: You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. #LI-KM5 Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan. Th
Jobs in United States
Platform Engineering Director in United States
3,618 active opportunities · Updated October 2026
Showing
15 jobs
Explore current platform engineering director jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. About our Team: Micron’s Industrial and Physical AI team is driving the transformation of semiconductor manufacturing through Autonomous Operations, AI, robotics, and digital twin technologies! We develop and deploy innovative solutions across Micron’s global fabrication and assembly/test facilities, enabling smarter, safer, and more efficient operations at scale. Position Overview: We are seeking a hands-on Full-Stack AI Engineer to design, build, and deploy production-grade AI applications that support Micron's Autonomous Operations initiatives. This role owns the end-to-end development lifecycle, from data pipelines and AI models to APIs, web applications, digital twin integrations, and cloud/edge deployments, delivering impactful solutions for engineers, operators, and business leaders worldwide. Responsibilities: Design, architect, and deliver end-to-end AI products, including data ingestion pipelines, feature engineering, model training/inference, APIs, user interfaces, and application monitoring. Build and maintain modern front-end applications using React, Angular, or Streamlit, supported by backend services in Python and FastAPI. Develop scalable integrations between manufacturing systems, robotics platforms, AMRs, sensor networks, and enterprise applications to enable intelligent factory operations. Design and implement digital twin environments using platforms such as NVIDIA Omniverse, Gazebo, or Unity Robotics Hub to support simulation, validation, and o
$175K – $275K/yr
A Career with Point72’s Technology Team As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. What you’ll do Develop and maintain the information security policy and standards library, aligning it with business priorities, regulatory expectations, and control objectives Lead independent assessments of the information security program, including regulatory examinations and third-party evaluations Identify technology risks across a complex business environment and drive mitigation aligned with our firm’s standards and control expectations Investigate data privacy inquiries and privacy-related events, assess business impact, and coordinate timely response activities Partner with technology, legal, compliance, and business stakeholders to translate risk findings into practical remediation plans Advise control owners on policy interpretation, risk treatment, and evidence expectations for assessments and examinations Prepare clear reports for management on team activity, emerging risks, remediation progress, and key decisions Maintain accurate risk, policy, priv
Manufacturing Engineer (Experienced or Senior) Company: The Boeing Company The Boeing Test and Evaluation (BT&E) Test Operations & Engineering (TO&E) Instrumentation Data System Installation Design team located in Berkeley, MO is seeking Manufacturing Engineers to provision Flight Test Instrumentation hardware on multiple aircraft platforms. On this team, a Manufacturing Engineer (Assembly & Installation) has the primary responsibility for the fabrication and installation of the mechanical and electrical hardware required to install a flight test data system. Design-Build Flight Test Manufacturing Engineers work the entire production cycle from review of engineering, manufacturing plan authorship, fabrication sourcing (inside and outside of Boeing) and detail planning, and continuing on through installation planning and shipside support. Manufacturing Engineers are owners of the build plan and ensure compliance to design configuration, build plan, and process requirements. Manufacturing Engineering performs Design for Manufacturing and Assembly (DFMA) analysis and develops an integrated build plan consisting of work instructions for assembly/fabrication, kitting, tooling, inspection requirements, and other outputs to ensure fabrication meets design, test, and program requirements. BT&E is currently hiring for a broad range of experience levels including Experienced and Senior level Manufacturing Engineers. Position Responsibilities: Synthesis of test hardware & instrumentation requirements in support of flight test Review of design models and drawings, creating/reviewing redlines and coordinating with design to incorporate updates Authorship of manufacturing plans to facilitate build
Principal Data Privacy Architect Description - Job Summary - Role Purpose • Lead and oversee complex, cross-functional privacy and data protection programs from strategy through implementation, ensuring alignment across business, technical, legal, and compliance stakeholders. • This role will design and implement scalable, AI-ready data privacy architecture across enterprise data environments, applications, and AI-enabled workflows. • The Principal Data Privacy Architect will serve as a hands-on subject matter expert responsible for embedding privacy-by-design, consent enforcement, data sovereignty, data loss prevention, and compliance controls into large, complex global data environments. • The architect will partner closely with Data Engineering, Cybersecurity, Legal, Privacy, AI Governance, Product, and Enterprise Architecture teams to ensure customer, employee, partner, and sensitive enterprise data is accessed, processed, shared, retained, and protected in a compliant, secure, and trustworthy manner. - Why This Role Matters • Architect for Trust & Scale: Build reusable privacy architecture patterns that enable secure, compliant, and scalable data usage across platforms, products, and regions. • Enable Responsible AI: Design privacy guardrails for AI agents, generative AI, RAG pipelines, model inputs and outputs, embeddings, vector stores, and automated data workflows. • Reduce Risk While Enabling Innovation: Translate privacy, consent, regulatory, and data sovereignty obligations into practical engineering controls that accelerate business outcomes. Responsibilities - Think Customer First • Embed customer trust, transparency, and privacy-by-design principles into enterprise data platforms and customer-facing applications. • Design consent-aware data access and usage p
Become a part of our caring community Job Description Summary The Lead Solutions Architect provides architecture leadership for CenterWell Home Health programs and platforms, shaping conceptual and reference architectures, governing solution designs, and aligning delivery teams to cloud and data strategies. The scope includes high-priority initiatives as well as interoperability and provider-data integrations that span CenterWell and Humana Insurance. As Lead Solution Architect, you'll be the senior individual contributor on a team with broad accountability across CenterWell's dispensing pharmacy portfolio — mail order, specialty, retail, and associated platforms. You'll own the architectural vision for complex, multi-system initiatives, shape how technology decisions get made, and act as a connective force between business strategy, engineering execution, and enterprise standards. You will operate within CenterWell IT – Cross-CenterWell Architecture. You will collaborate with product, engineering, EA Activation, security, data, and operations. You will engage governance forums to enable Integrated Health across CenterWell and Humana Insurance. Key Activities Quickly conduct structured knowledge transfer with existing architects and relevant stakeholders to capture critical in-flight designs and decisions. Review current initiatives and establish an architectural roadmap aligned with organizational priorities. Develop or refine reference architectures and design patterns for core platforms and solutions. Collaborate with governance and compliance teams to validate designs against enterprise standards and regulatory requirements. Define integration strategies and solution blueprints for key systems and data flows. Establish architecture review processes and decision forums to support de
$170K – $250K/yr
A Career with Point72’s Technology Team As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. What you’ll do Optimize cloud financial operations to maximize value from cloud investments, including rapidly growing artificial intelligence (AI) and machine learning workloads Provide actionable insights on cloud spend, SaaS license optimization, and emerging AI cost drivers, including model inference and usage-based consumption Implement tooling, tagging standards, and processes that improve cost visibility and optimization across cloud, SaaS, and AI workloads Monitor large language model API consumption and GPU-intensive infrastructure to identify cost trends, anomalies, and optimization opportunities Build financial models to forecast cloud, SaaS, and AI expenditures for budgeting cycles, commitment decisions, and vendor negotiations Design cost allocation, tagging, showback, and chargeback models that attribute spend to the teams, applications, and use cases driving it Educate engineering and business owners on cloud financial management practices th
Principal Embedded SW/FW Engineer (Bringup) - Austin, Tx, USA Job Summary We have an exciting opportunity to be part of a collaborative, cross-functional development team validating cutting-edge, high-performance AI chips and platforms. You will play a key role in supporting new product introductions and post-silicon validation. Working within the Post-Silicon Validation team, you will be involved with bringing first silicon to life, functionally validating it and working closely with many other teams to help it become a fully characterised and working product, reporting project status/progress to program management on a regular basis. You will have the opportunity to provide technical guidance to other engineering team members. In this role, you can leverage our experience and industry knowledge to architect and drive implementation of continuous improvements to test infrastructure and processes. The Team The Post-Silicon Bringup team sits within the Architecture and Validation team, we are responsible for bringup and validation of new silicon when it returns from manufacture, enabling and supporting the production SW and FW teams to bring up their software and supporting the Silicon Characterisation team. Responsibilities and Duties Plan, design, develop and debug silicon validation tests in bare metal C/C++ on FPGA/Emulator prior to first silicon Deploy silicon validation tests on first silicon and debugging them Develop automated test framework and regression test suites in Python to optimize validation efficiency Collaborate closely with engineers from many other disciplines on a variety of topics Work with Validation and Production Test engineering peers to implement best practices and continuous improvements to test methodologies Analyse test results, identify and debug failures/defects Contribute to shared test and validation infrastructure Provide feedback to architects Candidate Profile Essential: Understanding of ML
About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. About the Role We’re looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production — directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal
Data Architect, AI and Supply Chain Insights Description - Job Summary We are seeking a Supply Chain Data Architect who combines strong data architecture expertise with a data product mindset. The role is responsible for owning and evolving a portfolio of Data Products, designing scalable and reusable data models, and ensuring data assets are trusted, governed, and optimized for reporting, analytics, Data Science, Machine Learning, and AI use cases. The successful candidate will work closely with business stakeholders to understand processes, priorities, and opportunities, translating them into data strategies and architecture solutions that maximize business value and adoption. A strong understanding of data engineering concepts is foundational and required, including how data is acquired, transformed, and delivered across modern data platforms. The primary focus of the role is on data architecture, data modeling, data product maturity, business engagement, and AI ready design. Success is measured by the ability to create scalable, business ready data products that enable self service analytics, advanced insights, continuously leveraging the evolving AI technologies for our Supply Chain Stakeholder. The ideal candidate combines deep expertise in data architecture and modeling with business acumen, critical thinking, problem solving and a passion for innovation. They partner with business leaders, Data Engineers, Data Scientists, Product Owners, and AI practitioners to transform complex business challenges into data driven solutions Responsibilities Own and evolve a portfolio of Supply Chain Data Products, ensuring they are certified, compliant, governed, and aligned to the latest Data Governance standards while maximizing business value and adoption. Partner with busi
From $192K/yr
Distributed Systems engineers at Datadog design, implement and run in production the foundational platforms powering our applications. Your data pipelines will ingest, store, analyze and query in real-time billions of events per second from companies all over the globe. The platforms are optimized for durability, high availability, low latency, internet-scale footprint and operability. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Build fault-tolerant, horizontally scalable solutions running in multi-tenant environments Write in Go, Java Rust or C++, amongst other languages Use Kafka, Redis, Cassandra, Elasticsearch and other open-source components Own meaningful parts of our service, have an impact, grow with the company Who You Are: 6+ years of experience You have a BS/MS/PhD in a scientific field or equivalent experience You have significant backend programming experience in one or more languages (Go, Java, Rust, C++) You have been exposed to working on problems (high durability / low latency /…) You can get down to the low-level when needed You care about simple designs and performance You want to work in a fast, high-growth startup environment that respects its engineers and customers You have demonstrated ability to use AI coding tools in day-to-day workflows and validate, critique, and refine AI-generated output. Bonus: you’re motivated to push the boundaries of how AI can improve software engineering best practices and contribute to building AI-enabled products. This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. Datadog values peo
From $265K/yr
We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. OVERVIEW The Developer Experience team at Instacart is the force multiplier for our 1000+ engineers. We define and build the platforms, tools, and practices that enable Instacart engineers to ship high-quality software faster — from AI-powered code review to build infrastructure, developer onboarding, and local development environments. Our success is measured in developer velocity: how quickly an engineer can take an idea from concept to production. As a Staff Software Engineer on Developer Experience, you will lead two of the pillars that Instacart engineering is built on: Bazel (our build system) and Go (one of our core backend languages). You will own the technical direction for how thousands of daily builds run and how our Go services are structured, from remote build execution and caching to the frameworks, libraries, and patterns Go engineers use every day.
About the Team The People Technology team builds and operates the systems that support how OpenAI hires, develops, and supports its people. The team brings together People Systems and People Innovation Labs, a product engineering group focused on rethinking how we find and retain exceptional talent and help employees do their best work. People Systems owns the company’s core people-technology ecosystem, including platforms such as Workday and Ashby. People Innovation Labs builds new employee and People Team experiences on top of that foundation, including OpenHouse, our internal employee hub, and AI-powered products and automations. Together, we are working toward a model in which our enterprise systems provide reliable data, controls, and core business logic, while employees and managers can complete more of their work through simple, integrated, and AI-native experiences. About the Role We are looking for a People Systems Lead to manage the People Systems team and shape how our core systems evolve. You will be responsible for the reliability and effectiveness of our current environment while helping us move beyond the constraints of traditional enterprise software. This includes stabilizing and improving platforms such as Workday and Ashby, designing the integrations that connect them to the broader technology ecosystem, and partnering with People Innovation Labs to surface workflows through OpenHouse, Slack, and AI-powered experiences. This role requires someone who is comfortable moving between strategy, technical design, and team leadership. You should understand People systems deeply, be able to work through integration and architecture decisions with engineers, and translate complex organizational needs into scalable solutions. You will also manage vendor relationships, develop the People Systems team, and drive alignment across People, Engineering, Finance, Security, Legal, and other partners. This role could be a fit for someone who has grown up in People S
About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Through strategic partnerships and self-built campuses, we are scaling one of the world's fastest-growing AI infrastructure platforms. The Supply Chain organization ensures critical infrastructure components—from compute systems and networking equipment to integrated rack solutions—are sourced, manufactured, qualified, and delivered with the speed and reliability required to support frontier AI development. We partner closely with Hardware Engineering, Manufacturing Quality Engineering, Infrastructure Delivery, Hardware Operations, Finance, and suppliers worldwide to build a resilient, scalable supply chain capable of supporting rapid infrastructure expansion. As Industrial Compute continues to grow, Supply Chain serves as the operational bridge between engineering innovation and large-scale infrastructure deployment. About the Role We are seeking a Supply Chain Manager to lead strategic execution across sourcing, supplier operations, manufacturing quality, and infrastructure delivery for OpenAI's AI infrastructure portfolio. This role will oversee a multidisciplinary team responsible for strategic sourcing, manufacturing quality engineering, and technical program management while partnering closely with engineering, finance, hardware operations, and deployment teams. You will drive supplier strategy, manufacturing readiness, production planning, quality performance, and operational execution across the full hardware lifecycle. Success requires balancing long-term supplier strategy with day-to-day execution. You'll establish scalable operating mechanisms, strengthen supplier partnerships, manage complex cross-functional programs, and ensure OpenAI can rapidly deploy AI infrastructure without compromising quality, cost, or reliability. This is a people leadership role responsible for developing a high-performing organization while driving operati
About the Team The Stargate organization is responsible for building and scaling the physical infrastructure systems that power OpenAI’s next generation of AI training and inference platforms. This includes the manufacturing, deployment, and operational execution required to bring large-scale compute infrastructure online globally. The team operates at the intersection of data center infrastructure, hardware manufacturing, supply chain, deployment operations, and systems planning. We partner closely across Infrastructure Strategy, Manufacturing Operations, Capacity Planning, Supply Chain, Deployment, and Engineering to execute one of the largest infrastructure scale-outs in the industry. About the Role We are seeking a Technical Program Manager, Rack Delivery to drive operational execution across rack manufacturing, site readiness, and deployment coordination for Stargate infrastructure programs. This role will serve as a key connective layer between manufacturing partners, deployment teams, and infrastructure readiness programs to ensure rack production and delivery timelines remain aligned with site availability and deployment sequencing. You will help manage operational execution across contract manufacturers (CMs), support build planning and RCCA processes, and coordinate deployment readiness across multiple concurrent infrastructure programs. You will also partner closely with Demand Planning teams to translate strategic planning inputs into actionable SKU-level manufacturing and delivery schedules. This role is ideal for someone who thrives operating across ambiguity, manufacturing operations, infrastructure deployment, and large-scale cross-functional execution. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation support. Key Responsibilities Drive cross-functional coordination between rack manufacturing, deployment operations, and site readiness programs. Manage operational execution acros
Other cities to consider
More places hiring for this role
Get new platform engineering director jobs in United States by email
Daily job updates · Unsubscribe anytime