Job Details: Job Description: Ocotillo Technology Fabrication Yield is looking for Integration Technicians with a strong computing background and capabilities to support all Ocotillo Technology Fabrication platforms and beyond technologies. The candidate will work closely with a team of Integration Technicians to execute experiments/pilots, develop process flows and lot plans for new product introductions, and support taskforces. As a Process Integration and Yield Technician, you will play a critical role in Intel's high-tech manufacturing process, your contributions will directly impact Intel's innovation, enabling the delivery of advanced technologies that power the future of computing. This is an exciting opportunity to grow your technical expertise, work alongside engineering teams, and be part of a collaborative environment where your skills make a tangible difference. The Process Integration and Yield Technician will be responsible for but not limited to: Sustaining factory floors with lot plan/operation/route/process flow checks/rework/critical queue time edits, shipping errors, automation errors, process flow adjustments, wafer shipments, and scraps. Delivering lot plan owner training to the Ocotillo Technology Fabrication organization Collaborate with integration engineers to gather data needed for lot disposition Respond to and contain discrepant material in the fab Conduct engineering tests and detailed experimental testing to collect data or assist in research work related to integration, process, and yield. Provide pilot lot set up and cross site shipment support. Process change support and implementation via global flow changes, global route changes, and lot conversions The ideal candidate should exhibit the following behavioral traits: Self-starter an
Jobs in United States
Platform Engineering Lead in United States
3,618 active opportunities · Updated October 2026
Showing
15 jobs
Explore current platform engineering lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Manufacturing Engineer (Experienced or Senior) Company: The Boeing Company The Boeing Test and Evaluation (BT&E) Test Operations & Engineering (TO&E) Instrumentation Data System Installation Design team located in Berkeley, MO is seeking Manufacturing Engineers to provision Flight Test Instrumentation hardware on multiple aircraft platforms. On this team, a Manufacturing Engineer (Assembly & Installation) has the primary responsibility for the fabrication and installation of the mechanical and electrical hardware required to install a flight test data system. Design-Build Flight Test Manufacturing Engineers work the entire production cycle from review of engineering, manufacturing plan authorship, fabrication sourcing (inside and outside of Boeing) and detail planning, and continuing on through installation planning and shipside support. Manufacturing Engineers are owners of the build plan and ensure compliance to design configuration, build plan, and process requirements. Manufacturing Engineering performs Design for Manufacturing and Assembly (DFMA) analysis and develops an integrated build plan consisting of work instructions for assembly/fabrication, kitting, tooling, inspection requirements, and other outputs to ensure fabrication meets design, test, and program requirements. BT&E is currently hiring for a broad range of experience levels including Experienced and Senior level Manufacturing Engineers. Position Responsibilities: Synthesis of test hardware & instrumentation requirements in support of flight test Review of design models and drawings, creating/reviewing redlines and coordinating with design to incorporate updates Authorship of manufacturing plans to facilitate build
$170K – $250K/yr
A Career with Point72’s Technology Team As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. What you’ll do Optimize cloud financial operations to maximize value from cloud investments, including rapidly growing artificial intelligence (AI) and machine learning workloads Provide actionable insights on cloud spend, SaaS license optimization, and emerging AI cost drivers, including model inference and usage-based consumption Implement tooling, tagging standards, and processes that improve cost visibility and optimization across cloud, SaaS, and AI workloads Monitor large language model API consumption and GPU-intensive infrastructure to identify cost trends, anomalies, and optimization opportunities Build financial models to forecast cloud, SaaS, and AI expenditures for budgeting cycles, commitment decisions, and vendor negotiations Design cost allocation, tagging, showback, and chargeback models that attribute spend to the teams, applications, and use cases driving it Educate engineering and business owners on cloud financial management practices th
Principal Embedded SW/FW Engineer (Bringup) - Austin, Tx, USA Job Summary We have an exciting opportunity to be part of a collaborative, cross-functional development team validating cutting-edge, high-performance AI chips and platforms. You will play a key role in supporting new product introductions and post-silicon validation. Working within the Post-Silicon Validation team, you will be involved with bringing first silicon to life, functionally validating it and working closely with many other teams to help it become a fully characterised and working product, reporting project status/progress to program management on a regular basis. You will have the opportunity to provide technical guidance to other engineering team members. In this role, you can leverage our experience and industry knowledge to architect and drive implementation of continuous improvements to test infrastructure and processes. The Team The Post-Silicon Bringup team sits within the Architecture and Validation team, we are responsible for bringup and validation of new silicon when it returns from manufacture, enabling and supporting the production SW and FW teams to bring up their software and supporting the Silicon Characterisation team. Responsibilities and Duties Plan, design, develop and debug silicon validation tests in bare metal C/C++ on FPGA/Emulator prior to first silicon Deploy silicon validation tests on first silicon and debugging them Develop automated test framework and regression test suites in Python to optimize validation efficiency Collaborate closely with engineers from many other disciplines on a variety of topics Work with Validation and Production Test engineering peers to implement best practices and continuous improvements to test methodologies Analyse test results, identify and debug failures/defects Contribute to shared test and validation infrastructure Provide feedback to architects Candidate Profile Essential: Understanding of ML
About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. About the Role We’re looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production — directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal
From $192K/yr
Distributed Systems engineers at Datadog design, implement and run in production the foundational platforms powering our applications. Your data pipelines will ingest, store, analyze and query in real-time billions of events per second from companies all over the globe. The platforms are optimized for durability, high availability, low latency, internet-scale footprint and operability. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Build fault-tolerant, horizontally scalable solutions running in multi-tenant environments Write in Go, Java Rust or C++, amongst other languages Use Kafka, Redis, Cassandra, Elasticsearch and other open-source components Own meaningful parts of our service, have an impact, grow with the company Who You Are: 6+ years of experience You have a BS/MS/PhD in a scientific field or equivalent experience You have significant backend programming experience in one or more languages (Go, Java, Rust, C++) You have been exposed to working on problems (high durability / low latency /…) You can get down to the low-level when needed You care about simple designs and performance You want to work in a fast, high-growth startup environment that respects its engineers and customers You have demonstrated ability to use AI coding tools in day-to-day workflows and validate, critique, and refine AI-generated output. Bonus: you’re motivated to push the boundaries of how AI can improve software engineering best practices and contribute to building AI-enabled products. This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. Datadog values peo
About the Team The Stargate organization is responsible for building and scaling the physical infrastructure systems that power OpenAI’s next generation of AI training and inference platforms. This includes the manufacturing, deployment, and operational execution required to bring large-scale compute infrastructure online globally. The team operates at the intersection of data center infrastructure, hardware manufacturing, supply chain, deployment operations, and systems planning. We partner closely across Infrastructure Strategy, Manufacturing Operations, Capacity Planning, Supply Chain, Deployment, and Engineering to execute one of the largest infrastructure scale-outs in the industry. About the Role We are seeking a Technical Program Manager, Rack Delivery to drive operational execution across rack manufacturing, site readiness, and deployment coordination for Stargate infrastructure programs. This role will serve as a key connective layer between manufacturing partners, deployment teams, and infrastructure readiness programs to ensure rack production and delivery timelines remain aligned with site availability and deployment sequencing. You will help manage operational execution across contract manufacturers (CMs), support build planning and RCCA processes, and coordinate deployment readiness across multiple concurrent infrastructure programs. You will also partner closely with Demand Planning teams to translate strategic planning inputs into actionable SKU-level manufacturing and delivery schedules. This role is ideal for someone who thrives operating across ambiguity, manufacturing operations, infrastructure deployment, and large-scale cross-functional execution. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation support. Key Responsibilities Drive cross-functional coordination between rack manufacturing, deployment operations, and site readiness programs. Manage operational execution acros
About the Team The Core Services team is responsible for building and managing foundational services. It acts as the bridge between core infrastructure (e.g. compute, storage, networking) and product engineering teams, and enables product teams to move fast, build reliably, and scale efficiently. About the Role As a software engineer in the core services team, you will design and operate critical backend platforms such as caching systems, workflow orchestration, metadata stores, and file services. You’ll focus on building highly reliable, scalable, and performant systems that serve as the backbone of our products. We’re looking for people who are passionate about building infrastructure that empowers product teams, love working on distributed systems challenges, and enjoy creating well-designed APIs and abstractions that accelerate development. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain shared infrastructure services such as caching layers, workflow orchestration (Temporal), metadata stores, and file storage services. Collaborate with product teams to provide scalable, reliable primitives that abstract the complexities of distributed systems. Improve performance, resilience, and scalability of core services that power customer-facing applications. You might thrive in this role if you: Have experience with distributed systems, caching infrastructure (e.g., Redis, Memcached), metadata storage (e.g., FoundationDB), or workflow orchestration (e.g., Temporal, Cadence). Have experience running containerized services in cloud environments and integrating them into automated build/test/release (CI/CD) workflows. Understand trade-offs in consistency models, replication strategies, and performance optimization in multi-region systems. Excel at communication and collaboration with cross-functional teams, and are obsesse
Ground Systems Integration Engineer (Experienced or Senior Level) Company: The Boeing Company Boeing Defense, Space & Security (BDS) is seeking a Ground Systems Integration Engineer (Experienced or Senior Level) to support an Air Dominance Fixed Wing Proprietary Program in Berkeley, MO . Step into a fast-paced, cutting-edge program where your expertise will drive the design, integration, and testing of advanced, cloud-based systems that empower mission planning, debrief, and tactical Command & Control (C2) solutions for fixed-wing platforms. As a key contributor, you will support the development, analysis, integration, and testing of innovative engineering solutions for critical Ground System capabilities, including but not limited to: Mission Planning and Debrief Command and Control Systems Network and Communication Architectures Situational Awareness Enhancements You will take ownership of developing prioritized mission systems digital threads, ensuring seamless support for ground systems from initial design through to final delivery. You will work with a high-performing, cross-functional team in an agile environment, driving next-generation capabilities from design through delivery. This role offers the chance to innovate with open-architecture, model-based designs while collaborating across disciplines to support critical defense missions. If you’re passionate about advancing mission-critical systems and thrive in a collaborative, fast-moving environment, this is your opportunity to make a significant impact. Why Join Us? Impactful Work: Be part of a team that plays a crucial role in
About the Team Consumer Monetization builds the experiences and systems that power how customers purchase and pay for OpenAI products. Our scope spans purchasing flows, payments, subscriptions, and billing, along with the shared capabilities that support new products, offers, and distribution channels. We own both customer-facing experiences and the underlying platforms that power them. We partner closely with Product, Design, Growth, Data Science, and engineering teams across OpenAI to make purchasing effective and reliable, and new offerings easier to launch and monetize. About the Role We’re looking for experienced Staff+ engineers to evolve the payments, billing, and subscription capabilities that support OpenAI’s growing product portfolio. You’ll tackle problems where correctness, reliability, and flexibility are essential: supporting new billing requirements, managing the billing lifecycle, synchronizing state across internal systems and external providers, and enabling new products and commercial models. Your work may span several areas based on your expertise and team priorities: Billing and monetization capabilities: Extend billing capabilities and improve integrations and state consistency across systems to support new products and business models. Subscriptions: Orchestrate purchases, renewals, plan changes, cancellations, and recovery, ensuring customers are charged correctly and receive the right benefits. Payments: Expand payment capabilities through processor integrations, routing, and broader payment-method coverage. Risk and integrity: Partner with risk and integrity teams to integrate controls into purchasing flows, reducing abuse while protecting legitimate customer experiences. You’ll help set technical direction while remaining hands-on in implementation and delivery. This is an opportunity to solve complex engineering problems at scale, connect architecture decisions to customer and business outcomes, and help other engineers take on broader ow
About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. Our team works across systems software and product engineering to build reliable consumer devices and the platforms behind them. We develop the connectivity and networking foundations that support communication across OpenAI products and systems. About the Role As an Operating Systems Engineer focused on connectivity and networking, you will design, develop, and maintain the OS capabilities that enable reliable, secure, and efficient communication across OpenAI products and systems. Your work will span Wi-Fi and Bluetooth frameworks, IP networking, and advanced network services and policy. You’ll develop OS services, libraries, and interfaces for a broad range of connectivity needs, make design decisions across software boundaries, and carry solutions through development, integration, and production. In this role, you will: Build connectivity foundations: Design, implement, and maintain OS services, frameworks, and APIs for Wi-Fi, Bluetooth, and IP networking. Develop reusable network capabilities: Build connection management, network configuration and selection, service discovery, and routing capabilities. Enable secure communication: Develop network services and policies for secure communication, traffic management, and isolation across varied network environments. Resolve issues across the stack: Investigate correctness, concurrency, interoper
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Job Summary As a Business Systems Engineer on the GTM Engineering team, you will design, build, and operate the automation and AI systems that power ClickUp's Go-To-Market business — spanning core business products and the integration platforms that connect them. This role is AI-native at its core: you won't just maintain existing workflows, you'll actively advance our GTM systems with intelligent agents, LLM-powered automations, and next-generation integration patterns built on MCP, ClickUp Super Agents, and modern iPaaS tooling. You'll sit at the intersection of business systems architecture and applied AI — partnering with Sales, Finance, Revenue Operations, and fellow GTM Systems engineers to eliminate toil, accelerate revenue workflows, and build the automated, AI-augmented infrastructure the company runs on. This is a hands-on engineering role for someone deeply fluent in enterprise business systems, excited about deploying production AI, and committed to genuine ownership of the platforms they build — directly supporting GTMSOE's broader mission of operational excellence across the GTM org. Key Responsibilities AI-Native Automation & Agent Development Design and build AI-powered automations and agentic workflows across the GTM tech stack — including Salesforce, NetSuite, Workato, and MuleSoft — to eliminate manual effort and accelerate business operations. Develop and deploy ClickUp Super Agents and LLM-based automations to automate tasks such as deal data enrichment, quote generation assistance, order validation, revenue recognition triggers, and exception handling in quote-to-cash workflow
About the Team At OpenAI, our Trust, Safety & Risk Operations teams safeguard our products, users, and the company from abuse, fraud, scams, regulatory non-compliance, and other emerging risks. We operate at the intersection of operations, compliance, user trust, and safety working closely with Legal, Policy, Engineering, Product, Go-To-Market, and external partners to ensure our platforms are safe, compliant, and trusted by a diverse, global user base. The Global Safety Response Operations team within the org provides 24/7 coverage for user safety, risk, and regulatory escalations across OpenAI’s products, handling the highest-priority cases that require human judgment and rapid response. The team operates as the core escalations management and delivery arm of OpenAI’s safety operations, ensuring that our products remain safe and aligned with our policies while enabling timely, empathetic, and consistent user support. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Please note: This role may involve exposure to sensitive content, including material that is sexual, violent, or otherwise disturbing. About the Role We’re looking for experienced Trust, Safety, and Risk Operations analysts who have subject matter expertise in one or more of the following areas: policy enforcement and content moderation, fraud and scam prevention, developer risk, or privacy and regulatory escalations. You’ll be on the front lines of safety escalation management, helping to triage and resolve urgent and sensitive cases. You’ll work across subject matter areas, systems, and processes to ensure operational excellence, develop process improvements and automations, and surface insights and trends. This is a 24/7 global operation that requires flexibility to work rotating shifts, including nights, weekends, and holidays, as part of an on-call coverage model. We use a hybrid work model of 3 days in the office per week and offer r
About the Team The Core Models team helps shape how OpenAI’s frontier models are built, measured, and launched. We work across Research, Engineering, Model Design, Data Science, and Product to turn advances in model capabilities into reliable, useful experiences for people. Our scope includes model planning and launches as well as building data flywheels, evaluations and measurement systems to ensure our models have strong capabilities and behavior. About the Role As a Product Manager for the Core Models team, you'll be at the forefront of defining and guiding the future of how our AI models work in real-world applications. You will connect user needs to model and systems decisions: how prompts are understood; how information is aggregated and made useful for training and evaluation data; and how capabilities move from research prototypes into the mainline model and launch stack. You will operate comfortably across research, infrastructure, and consumer product surfaces, creating clarity where ownership and technical boundaries are still emerging. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate user and product goals into clear model requirements, system architecture choices, and research priorities across query understanding, indexing, retrieval, ranking, tool boundaries, data, training, inference, and evaluation. Build closed learning loops that turn product usage, explicit feedback, and other user signals into datasets, evaluations, experiments, training priorities, and launch decisions. Define success across offline evaluations and online product metrics, balancing model quality, usefulness, latency, safety, reliability, and cost. Partner closely with post-training research, applied product engineering, Model Design, and Data Science to integrate capabilities into the mainline model stack. Create reusable platforms and operatin
About the Team This team builds and operates the systems that enable OpenAI researchers to run reliable, scalable, and efficient research workflows. The team sits close to research and works across infrastructure, systems, and automation to make sure researchers have the tools and environments they need to move quickly. The work spans software engineering, infrastructure, systems administration, cluster operations, and reliability engineering. As OpenAI’s infrastructure evolves from bespoke bare-metal systems toward more standard, scalable platforms, the team needs engineers who can understand how systems work end-to-end and build the right abstractions without reinventing the wheel. About the Role As a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability - while also writing software to make these systems better over time. This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work. We expect you to: Build and operate reliable infrastructure for research workloads and research-facing services. Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services. Improve cluster bootstrapping, provisioning, automation, and deployment workflows. Debug issues across networking, compute, storage, orchestration, and service reliability layers.
Other cities to consider
More places hiring for this role
Get new platform engineering lead jobs in United States by email
Daily job updates · Unsubscribe anytime