About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software is reliable, testable, and ready to ship. We design and maintain automated test frameworks, hardware-in-the-loop labs, and release pipelines that keep quality signals trustworthy and enable rapid, safe product launches. Our work spans developer tools, automation, systems integration, and cross-team collaboration to ensure every release meets the highest standards. About the Role As a Software Engineer, Quality and Developer Tools , you will build and own the systems that validate our device software—from test frameworks and regression infrastructure to hardware-in-the-loop labs and release gates. You’ll design the tooling and automation that keep quality signals trustworthy, integrate them into CI/CD, and make it easy for engineers and QA vendor technicians to execute reliable, repeatable workflows. We’re looking for engineers with deep experience in software quality, automation, developer tooling, and hardware-software integration who thrive on building scalable, reliable systems for validation and release readiness. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Test infrastructure & frameworks: Design, implement, and maintain a unified test framework for device software across unit, integration, system, and end-to-end testing, with reproducible runs and integrations with GitHub, Linear, and Slack. CI/CD integration & releases: Integrate test suites with Buildkite, enforce promotion criteria for staging and production, auto-file regressions, and publish traceable artifacts and release notes. Hardware-in-the-loop lab design & orchestration: Plan and bring up racks, power and networking systems, and orchestration for device testing; support automated flashing, provisioning
Jobs in United States
Aws And Tooling Platform Lead in United States
2,026 active opportunities · Updated October 2026
Showing
15 jobs
Explore current aws and tooling platform lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. About The role As a Vulnerability Management Engineer, you will support the identification, assessment, prioritization, and remediation of vulnerabilities across applications and infrastructure. Working under the guidance of senior team members, you will assist in understanding how vulnerable dependencies enter an application, identifying remediation options, and engaging with engineering teams to track fixes. You will contribute to the maintenance of internal vulnerability-management tools, such as scripts, documentation, and reporting. The ideal candidate will have a desire to grow their AppSec expertise, will be eager to learn about modern security tooling and automation, and will be comfortable using AI tools like Claude to assist with documentation, investigation, and scripting tasks while following company security and data-handling requirements. What you’ll do Perform regular vulnerability assessments using different tools. Regularly drive remediation and reporting of cataloged vulnerabilities. Assess discovered vulnerabilities and properly prioritize their scope, impact and necessary response actions. Conduct security reviews of our products and production infrastructure. Contribute to vulnerability management, application security and/or offensive/red-team operations. Engage in security audit
About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. Build the operational and technical skill set that runs global trade At Flexport, we believe global trade can move the human race forward. We're shaping the future of a $10T industry with technology and people who understand it end to end. The Rotational Development Program (RDP) exists because that combination — real supply chain fluency plus the ability to build and automate — doesn't exist at scale in the market today. So we're building it ourselves. This is an 18-month program, not a rotation for rotation's sake. Every assignment is live, revenue-generating work. Every participant leaves with a certification in AI and low-code tooling running alongside their operational and commercial training. And every graduate places into a role with real scope: primarily Account Management, with paths into Automation Engineering, Forward Deployed Engineering/Consulting, Sales Engineering, or a traditional Ops/Sales/Customs track. What you'll do You'll rotate through three functions, six months each, with AI/low-code curriculum running continuously across all three: Gateway Ops (Air or Ocean). Own the shipment lifecycle. Execute bookings, manage exceptions, and hold data int
$295K – $380K/yr
About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Senior Software Engineer, ML Systems & Training Infrastructure, you will be a deeply hands-on engineering force multiplier for the robotics team. You will help keep the training framework and surrounding infrastructure healthy, review and improve code quickly, debug failures across ML systems and infrastructure, and unblock researchers and engineers when the path from idea to working training job gets rough. We’re looking for people who love writing, reading, reviewing, and fixing code; who can get productive quickly in unfamiliar systems; and who bring strong practical judgment without a lot of ego or process overhead. This role will be based in San Francisco, CA and be expected in office 5 days per week and offer relocation assistance to new employees. In this role, you will: Review, improve, and clean up code across training frameworks and adjacent infrastructure. Identify risky or low-quality changes before they land, and raise the code quality bar without slowing the team down. Debug issues across ML training systems, GPUs, clusters, networking, and related infrastructure. Help researchers and engineers unblock broken training jobs, flaky workflows, and brittle internal tooling. Improve the reliability, maintainability, and usability of the robotics team’s training framework. Move quickly on practical engineering problems that directly affect team velocity. You might thrive in this role if you: Have strong software engineering fundamentals and excellent code review judgment. Have experience with ML systems, training fr
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Threat Intelligence team protects OpenAI’s technology, people, research, and infrastructure by proactively identifying and disrupting adversaries who seek to compromise our systems or misuse our models. We investigate sophisticated threats, build tooling to scale and augment analysis, and deliver intelligence that shapes security strategy and equips leadership with timely, risk-aware insights. We combine technical depth, investigative rigor, and strong cross-functional partnerships to uncover threats and drive impact across OpenAI’s security and research organizations. About the Role As a Technical Threat Investigator at OpenAI, you will help protect the company from sophisticated adversaries targeting OpenAI and the broader ecosystem, as well as those attempting to misuse our models in support of cyber operations. This is a deeply investigative role. You will independently conduct complex, end-to-end investigations into capable threat actors to understand their behavior, infrastructure, emerging techniques, and how AI is integrated into their workflows. You’ll use these insights to proactively identify malicious activity and drive detection, disruption, enforcement, and safety improvements across the company. You’ll translate your investigative findings into durable solutions that scale impact. You’ll build and own lightweight tooling, automate where it matters, and create AI-assisted workflows to make investigations faster, more repeatable, and more effective over time. In this role, you will: Conduct deep, end-to-end investigations into sophisticated threat actors interacting with OpenAI’s models, products, and broader ecosystem. Think like an adversary — model attacker behavior, anticipate misuse patterns, and proactively hunt for, identify, and disrupt malicious activity. Leverage internal telemetry, OSINT, vendor data, a
About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Team Overview The TechOps team is the technical foundation that is a major stakeholder in keeping the company running. We own the systems, tools, and infrastructure that every Plaid employee depends on, from identity and endpoint management to help desk support, office infrastructure, and the internal tooling that powers day-to-day productivity. What sets us apart from a traditional IT team is how we approach our higher-level goals. We treat corporate infrastructure like an engineering problem: configuration lives in code when possible, endpoint provisioning is automated, access assignment is self-service, and we're always looking for ways to make our systems more reliable and our support burden smaller. We're a small team with broad ownership and high standards. We work closely with Security, Engineering, and People teams to make sure Plaid's internal environment is secure, scalable, and ready for where the company is going, whether that's a new office, a new compliance requirement, or a new way of working enabled by AI tooling. Role Overview In this role, you'll take part in our on-call rotation for a few hours each week, but this isn't just a break/fix IT position. It's a chance to build, improve
We anticipate the application window for this opening will close on - 28 Sep 2026 Careers that change lives start here. Medtronic is a global leader in healthcare technology with a Mission to alleviate pain, restore health, and extend life. Our 95,000 employees work across more than 150 countries to put patients first — developing innovative medical technologies that improve the lives of 72+ million patients each year. Your unique talents will help shape the future of healthcare while building a career grounded in purpose, growth, and impact. A Day in the Life Principal Software Engineer Careers That Change Lives At Medtronic, we push the limits of what technology can do to make tomorrow better than yesterday, and that makes it an exciting and rewarding place to work. The Device Embedded Systems group in Medtronic’s Electrophysiology Therapies (EPT) operating unit designs, develops, and verifies embedded applications used in implantable medical devices. As a Principal Software Engineer, you will use your strong DevOps and CI/CD expertise to advance the tools and workflows used by the Device Embedded Systems group. You will play a leading role in the group’s transition to GitHub Enterprise, the adoption of GitHub Copilot and AI-assisted development agents, and the delivery of select tools in partnership with a dedicated AWS cloud team. We are looking for a highly motivated, self-starting software engineer who wants to innovate in a fast-paced environment and improve people’s quality of life. The ideal candidate should have experience designing, implementing, and improving modern CI/CD processes and developer productivity tooling, including building workflows and automation frameworks from the ground up on GitHub Enterprise. Key technical skills desired include: • CI/CD pipeline design and imple
About the Role As a member of the Data team within the Go-to-Market organization, you will help build a data-driven culture, improve decision-making, and advance strategic initiatives through analytics. This is a full-stack data role spanning data modeling, metric definition, visualization, analysis, and self-service tooling. You will build trusted, scalable data sources and products that give the business reliable, actionable insights. The work calls for judgment: you will choose the tool, approach, and level of investment that best fit each problem, from a focused analysis to a durable production data product. As a core partner to the GTM organization, you will address both foundational and ad hoc analytics needs. You will turn complex data into clear narratives that help technical and non-technical audiences understand what is happening, why it matters, and what they should do next. In This Role, You Will Partner closely with GTM teams to proactively identify high-impact questions and translate business needs into data models, metrics, analyses, and scalable technical solutions. Define, source, validate, and operationalize the metrics that guide the business, helping teams incorporate them into planning and day-to-day decisions. Lead cross-functional data projects across established and emerging business areas, including setting the data strategy for greenfield domains. Build scalable data models and pipelines that integrate and transform data from multiple sources into trusted, accessible datasets. Create dashboards, reports, analytical tools, and other data products that enable stakeholders to answer questions independently. Own the lifecycle of metrics, analytical models, and data products from initial exploration and prototyping through production and ongoing maintenance. Choose the most effective approach for each problem—whether an analysis, metric, data model, visualization, or self-service product—based on the audience, urgency, complexity, and expected v
From $125K/yr
About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Lead Engineer on the Product Catalog Manager Team, you will help set the technical direction for the systems that power Stitch Fix’s product data ecosystem. You will work on the tools, workflows, and data models that support the full product lifecycle, from new style creation and catalog enrichment to product readiness, validation, and downstream product experiences. You will own complex problem spaces from discovery through delivery, translate business and merchandising needs into scalable technical solutions, and lead execution across ambiguous, cross-functional initiatives. This role requires strong technical judgment, deep ownership, clear communication, and the ability to influence partners across Engineering, Product, Merchandising, Data Science, and Operations. Your work will directly impact product data quality, catalog accuracy, merchandising efficiency, product readiness, and the client experience. Responsibilities: Own and evolve critical catalog systems, including product onboarding, attribute management, data enrichment, validation workflows, and product readiness tooling. Design and operate scalable services and data models that ensure product information is accurate, complete, consistent, and available to downstream systems. Drive discovery with Product, Merchandising, Data Science, and Operations partners to identify high-impact problems, evaluate tradeoffs, and define clear technical roadmaps. Independently lead initiatives from concept through production rollout, including tec
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We are hiring for a leader for the newly forming Detection and Response team at Plaid. Our mission is to protect Plaid's financial infrastructure by detecting and responding to suspicious activity across the company. We are responsible for the entire lifecycle of detection and response, including detection infrastructure, AI triage and response, investigation tooling, Red Teaming, and Fraud Operations. Security is foundational to the trust thousands of businesses and millions of consumers place in Plaid, and we work directly to reduce risks and enable our business to move faster and safer. As the Head of Detection and Response, you will be the founding leader responsible for standing up and evaluating Plaid's detection and response team. You will lead a specialized technical team of analysts and engineers, gain deep experience partnering with the CISO and cross-functional engineering leaders, and build critical security infrastructure at scale. This is a unique leadership opportunity to manage a team that encompasses traditional security operations alongside Red Teaming and Fraud Operations, directly impacting Plaid's security posture and long-term stability. Responsibilities: Form and set up Plaid’
About the Team OpenAI's Industrial Compute organization is building and operating the infrastructure foundation for the next generation of AI. Infrastructure Operations works across facilities, hardware, network operations, incident management, data center engineering, delivery teams, and external partners to bring capacity online safely, understand its operational state, and improve it over time. As OpenAI's data center portfolio grows across first-party and partner-delivered capacity, the organization needs clear goals, trusted data, repeatable processes, and systems that make ownership, risk, readiness, and performance visible. This role will help build the operating mechanisms that allow Infrastructure Operations to scale with rigor. About the Role We are seeking a Technical Program Manager to own the systems, data, reporting, governance, and program-management backbone for Infrastructure Operations. Reporting to the Delivery & Operations Lead, you will translate strategy into executable goals and operating cadences, turn operational needs into software and data solutions, and create the mechanisms that keep a rapidly evolving organization aligned and accountable. This role will also own the current 1P+3P delivery-tracking layer within Operations: milestones, delivery timelines, quantity forecasts, risks, decisions, and executive reporting. You will partner closely with 1P Delivery Program Management, Compute TPMs, Data Center Engineering, construction, commissioning, and operations leaders to ensure that delivery information becomes complete, usable input for readiness, handover, and ongoing operations. You will own program health and the operating system around it: the goals, data definitions, workflows, reporting, decision paths, and follow-through that help functional DRIs execute. The ideal candidate is comfortable in ambiguity, technically fluent enough to implement real systems, and relentless about converting scattered information into durable mechan
From $114K/yr
The Assurance, Risk and Compliance (ARC) Initiatives team at MongoDB owns the governance and delivery of key cross-functional security risk and compliance initiatives. The team designs and executes programs that support compliance audits, risk assessments, common control frameworks, operating cadences, and executive reporting that strengthen the organization’s assurance, risk management and compliance objectives. The policy and controls governance pillar is responsible for the structure, standards and operating mechanisms that keep MongoDB’s security policies, standards, procedures, and controls governance processes current, aligned, auditable and scalable across the organization. This includes ownership of the policy lifecycle, common controls framework governance, issue management, and the review cadences and cross-functional coordination needed to maintain strong governance maturity and audit readiness. This role sits under the Assurance, Risk and Compliance function within the Global Security Office and reports to the Director of ARC Initiatives. This role will be based remotely in the United States Responsibilities: Scope of Ownership Policy governance program ownership, including policy lifecycle management, documentation standards, review and approval cadences, change tracking, and exception governance Controls governance ownership, including common controls framework lifecycle management, control harmonization, framework mapping, and processes that support audit readiness and scalable control oversight Governance over supporting systems and workflows, including Jira, GRC tooling, documentation repositories, and reporting structures, that enable consistent execution and visibility Issue management and remediation governance, including intake, triage, tracking, and reporting for timely closure of findings Executive-ready reporting and metrics for policy health, controls maturity, policy exceptions and broader program effectiveness Program Leadership Own
Other cities to consider
More places hiring for this role
Get new aws and tooling platform lead jobs in United States by email
Daily job updates · Unsubscribe anytime