Jobs in United States

Platform Architecture Manager in United States

3,692 active opportunities · Updated October 2026

Explore current platform architecture manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Mechanical and Thermal Laboratory Technician Position Summary Graphcore is a globally recognized leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data center hardware that provide the specialized processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Austin, which will play a central role in Graphcore's work building the future of AI computing. Responsibilities The Mechanical and Thermal Laboratory Technician is a hands-on technical role supporting the development and validation of advanced AI hardware systems for data center environments. Working as part of a cross-functional engineering team, this individual will be responsible for executing mechanical and thermal laboratory testing, supporting product validation activities, prototype fabrication and assisting with troubleshooting and root-cause analysis of complex hardware systems. Requirements Associate degree in Mechanical Engineering Technology or a related technical field preferred. Equivalent combinations of education, training, and relevant experience will be considered, including experienced non-degreed candidates or candidates with degrees in unrelated disciplines. Minimum of 5 years of experience working in mechanical laboratories, machine shops, test labs, or similar technical environments. Experience with server hardware platforms and data center equipment. Knowledge of Direct Liquid Cooling (DLC) systems and their implementation in server environments. Experience operating forklifts, pallet jacks, and other material-handling equipment. Key Responsibilities Execute mechanical and thermal test plans to validate hardware designs a

AISEMTraining
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. The Team The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Responsibilities and Duties Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. Own execution against program achievements and report validation progress and risks. Triage test failures, collect debug data, and collaborate on root cause analysis. Track

PythonCI/CDLinuxAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Manufacturing Test Engineer – Server Hardware Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Austin, which will play a central role in Graphcore's work building the future of AI computing. Role Overview We are seeking an experienced Manufacturing Test Engineer to support high-volume server manufacturing from board-level test through system-level production test. This role will work closely with an ODM manufacturing partner to define, implement, validate, and optimize the manufacturing test strategy for L6 board-level products , including ICT, MDA, and Board Functional Test , as well as support L10 system-level manufacturing test . The ideal candidate has strong experience in server hardware manufacturing, Linux-based test environments, diagnostic test coverage, fixture requirements, yield improvement, and root cause corrective action processes. This role requires both technical depth and hands-on manufacturing execution experience, with the ability to drive best practices across test development, factory readiness, quality planning, and ongoing production support. Key Responsibilities Manufacturing Test Strategy and Planning Work with ODM partners to define and execute the manufacturing test strategy for L6 board-level production . Develop and review test plans covering: In-Circuit Test, or ICT Manufacturing Defect Analyzer, or MDA Board Functional Test Diagnostic coverage requirements Manufacturing line test flow Fai

PythonLinuxRestAI
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the foundational infrastructure that will define how the world thinks about and deploys AI, and we want the sharpest, most curious people to help us do it. As a Lead Data Scientist on our Analytics and Data Insights team, you'll tackle problems that don't have textbook answers yet; shaping go-to-market strategy for technology that's still being invented, designing the experiments that prove or kill our biggest bets, and helping enterprises understand what foundational AI actually means for their bottom line. You'll own the full analytical lifecycle, from framing the right questions and building the models, to leading a team that delivers answers leadership can act on. As a Lead Data Scientist, you will: Drive the mission forward. Own the science: design and lead experimentation programs including A/B tests, multi-armed bandits, causal inference studies, that directly map to product and go-to-market decisions. Build predictive models that matter: develop and deploy models for forecasting, segmentation, propensity scoring, and opportunity sizing across Cohere's core business lines. Lead and grow a tea

PythonSQLGitAI
C
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the data infrastructure behind some of the most demanding AI training workloads in the world, and we want sharp, curious people to help us do it. In this role, you'll build and maintain the high-performance data layer our Modeling teams rely on for training and evaluation jobs. As a Software Engineer, Data Infrastructure, you will: Work directly on petabyte-scale storage infrastructure, and the networking and performance challenges that come with it. Collaborate daily with researchers and engineers who are some of the best in the world at what they do. You may be a good fit if you have: 4+ years of experience working on data storage infrastructure Strong command of Python Kubernetes experience, especially on the storage side (Persistent Volumes, CSI drivers, etc.) The ability to transform unstructured data into performant datasets across diverse storage backends including S3, GCS, and POSIX Experience with distributed data processing frameworks such as Apache Beam, Spark, or Flink [Nice-to-have] Familiarity with modern analytics tooling such as BigQuery, Airflow, or dbt Genuine excitement about AI.

PythonKubernetesGitAI
C
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the foundational infrastructure that will define how the world thinks about and deploys AI, and we want the sharpest, most curious people to help us do it. As a member of our Analytics & Data Insights team, you'll tackle the kind of problems that don't have textbook answers yet, launch products that didn't exist a year ago, and help enterprises understand what foundational AI actually means for their bottom line. As a Data Engineer, you will: Work directly on new customer experiences built on one of the most advanced AI systems in the world Collaborate daily with researchers and engineers who are some of the best in the world at what they do Run implementations end-to-end and see initiatives through to real outcomes Partner across research, marketing, sales, and finance to help define how Cohere grows, with your recommendations feeding directly into products and strategy You may be a good fit if you have: 5+ years of experience working on production-grade data processing systems Strong command of Python and SQL Experience with distributed data processing frameworks such as Apache Beam, Spark, or

PythonJavaSQLKubernetes
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the foundational infrastructure that will define how the world thinks about and deploys AI, and we want the sharpest, most curious people to help us do it. As a Data Scientist on our Analytics and Data Insights team, you'll work on problems that don't have textbook answers yet; building the analytics that shape product and company strategy, designing the experiments that prove or kill our biggest bets, and helping enterprise customers understand what foundational AI actually means for their bottom line. You'll own analytical work end-to-end: from framing the right questions and building the models to shipping insights, tools, and results that product leaders, sales teams, and enterprise customers rely on. As a Data Scientist, you will: Drive the mission forward. Build bleeding-edge agentic analytics: Agentic analytics is far from a solved problem, and we want to be the company that solves it. We need the sharpest minds with the curiosity, drive, and focus required to build the tools necessary to bring order and clarity to real world data. Define AI impact measurement: own the end-to-end analytics st

PythonSQLGitAI
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the hardware platforms that support the next generation of AI systems. The Opportunity We are looking for a recent graduate or early-career engineer to join the Mech/Thermal Platform team as a Graduate Mechanical Engineer. You will contribute to the mechanical design, thermal validation, and verification of advanced AI hardware and data center systems. You will work with experienced mechanical, thermal, hardware, and systems engineers throughout the development lifecycle. The role combines 3D computer-aided design, prototype evaluation, lab testing, data analysis, troubleshooting, and clear engineering documentation. What You Will Do Create and update mechanical parts, assemblies, and drawings using Creo or comparable 3D CAD software. Take ownership of defined mechanical design tasks from requirements and concepts through detailed design, review, release, and verification. Apply mechanical design, materials, manufacturing, heat transfer, and thermodynamics principles to engineering decisions. Develop mechanical and thermal test plans and procedures for prototypes and development systems. Set up and operate laboratory equipment, collect accurate data, analyze results, and document conclusions. Evaluate prototypes and identify mechanical, thermal, assembly, or manufacturability issues. Support design verification, troubleshooting, root-cause analysis, and implementation of verified design improvements. Maintain accurate engineering documentation, including design notes, drawings, test procedures, results, and change r

Artificial IntelligenceAI
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world’s most transformative technologies. Our new AI Engineering Campus in Austin will play a central role in building the future of AI computing. The Opportunity As a Mechanical Engineering Intern, you will contribute to the mechanical design and thermal validation of advanced AI hardware. Working alongside experienced mechanical and thermal engineers, you will gain hands-on experience with 3D computer-aided design (CAD), laboratory setup, test development, prototype evaluation, and engineering documentation. Type: 12-week summer internship Timing: May - August (exact dates to be confirmed) Commitment: Full-time What You’ll Do Contribute to mechanical design activities using Creo or comparable 3D CAD software, including part modeling, assemblies, drawings, and design updates under the guidance of experienced engineers. Work with the thermal engineering team to develop test procedures and help set up laboratory capabilities for evaluating AI hardware. Support thermal and mechanical testing using appropriate laboratory equipment; collect, organize, and analyze test data. Assist the mechanical and thermal design teams with prototype evaluation, troubleshooting, design verification, and documentation of findings. Collaborate with cross-functional engineering partners, communicate progress and issues clearly, and follow applicable laboratory and safety procedures. What You’ll Bring Current enrollment in a bachelor's, master's, or doctoral program in mechanical engineering or a related discipline during the internship. Students in other disciplines with relevant mechanical engineering exper

Artificial IntelligenceAI
NR
📍 Atlanta, Georgia, United States· Full-time
✓ High-confidence listingCompany trend -75%

From $98K/yr

Quick readStrong listing-quality and freshness signals

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity At New Relic, we provide our customers with real-time insights, so they can innovate faster. Our software provides deep observability across the stack, enabling software teams to solve their customer’s problems, accelerate digital transformation, and make DevOps work. You will be at the heart of the teams supporting New Relic’s infrastructure and will work on a team that provides global service mesh and load balancing solutions. We provide these services on-premises, as well as using our multi-cloud infrastructure. We support each other to do our best work through positive communication and continuous improvement. What You’ll Do As a key member of our Infrastructure team, you will design and operate a scalable, resilient ingress data plane that directly impacts the value we provide to our customers. By ensuring the stability and performance of our global service mesh and load balancing solutions, you drive the foundational reliability that the entire New Relic organization depends on to deliver real-time insights. You will leverage advanced automation and infrastructure-as-code to accelerate development speed, allowing our engineering teams to ship safe, incremental changes across a massive fleet with confidence. Your work in evolving our DNS and CDN infrastructure is not just about maintenance; it is about creating a seamless, high-performance environment that enables innovation at scale. Through deep collaboration with Product, Design, and partner platform t

PythonAWSAzureKubernetes
O
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -84.1%
Quick readStrong listing-quality and freshness signals

About the Team API Frontiers turns OpenAI’s frontier models into production APIs that developers can use to build reliable products and agents. We own the core path connecting models to developers through the Responses API, with a focus on safety, reliability, and speed. Working closely with Research, Safety, Codex, and other API teams, we bring new model capabilities into production and improve them through developer feedback. About the Role We are looking for a backend software engineer to build and operate the services behind the Responses API. You will shape API behavior, bring new capabilities from research into production, and make long-running agent workflows dependable and fast. The work combines distributed systems engineering with product judgment: designing useful developer interfaces, managing staged rollouts, and following production issues through to durable fixes. In this role, you will: Design, build, and operate APIs and backend services that bring frontier model capabilities to developers. Partner with Research, Safety, Codex, and API teams to define API behavior and deliver safe, staged launches. Build API capabilities for agent workflows, including task delegation, context sharing, and parallel execution. Strengthen long-running request reliability across timeouts, cancellation, streaming, and background execution. Improve request-processing performance and tail latency through profiling, efficient systems code, and persistent connections. Turn developer feedback and production failures into better observability, diagnostics, and lasting product improvements. Your background might look something like: 5+ years of experience building and operating backend services or developer-facing APIs in production. Strong software engineering fundamentals, with practical knowledge of distributed systems, concurrency, and asynchronous execution. Ability to diagnose production failures and performance bottlenecks using observability data and profiling. Product

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI's mission is to ensure that artificial general intelligence (AGI) benefits all of humanity. The API Platform turns frontier research into reliable capabilities that developers use to build transformative products and services for people around the world. API Safety's goal is to ensure safe deployment of frontier models in the API. We design APIs and systems that help developers share usage context, understand safety events, and apply safeguards tailored to the risk profile of the applications they are building. This work is critical to our frontier model launches and partners closely with teams across API, Integrity, and Safety Research. About the Role We're looking for product-minded software engineers to join a team that is addressing emerging risks at the frontier of model development while building novel solutions for real-world AI deployment. The day-to-day work ranges from solving production challenges to designing new product experiences and safeguards. The right candidate is comfortable balancing tradeoffs across developer experience, latency, reliability, and risk. In this role, you will: Design and build dashboards and APIs for safety controls and customer-facing observability. Develop scalable systems that extend trusted safety capabilities to new use cases, customers, and deployment environments. Partner with Safety Research and Integrity to build safeguards that mitigate emerging risks. Be responsible for the availability, latency, and scalability of safeguards across high-volume API traffic. Own projects from technical design and implementation through launch and ongoing iteration, while raising the team’s engineering standards Your background might look something like: 7+ years of professional experience, excluding internships, in backend, infrastructure, platform, or product engineering roles. A track record of designing, building, and operating production backend services, developer-facing APIs, or distributed systems. Strong s

TypeScriptPythonArtificial IntelligenceAI
NR
📍 Atlanta, Georgia, United States· Full-time
✓ High-confidence listingCompany trend -75%

From $186K/yr

Quick readStrong listing-quality and freshness signals

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity Are you ready to step into a pivotal leadership role where your engineering depth directly shapes the future of our core platform? As our new Engineering Manager, you will lead a talented, distributed team across US and EU time zones, acting as the critical manager bridging regional collaboration. Our Cloud Foundation team is the backbone of the New Relic platform. In this role, you won't just manage tasks; you will mentor and empower engineers, transitioning our operational framework from a reactive state to a culture of proactive ownership and engineering excellence. You will oversee critical global initiatives, including major regional expansions into FedRAMP High / IL4, India, and Australia. If you thrive on solving complex multi-cloud challenges at an exabyte scale while helping engineers grow in their careers, this is your opportunity to make a lasting impact. What you'll do Empower & Mentor: Lead and nurture a high-performing engineering team across the US and EU, facilitating career development, performance growth, and a collaborative team culture. Drive Strategic Ownership: Champion a shift from reactive delivery to proactive technical ownership, establishing best practices for platform reliability and cross-regional alignment. Lead Regional Expansions: Architect and execute key global infrastructure expansions across complex environments (including FedRAMP High / IL4, India, and Australia). Architect for Extreme Scale: Guide decisions around micr

ReactAWSAzureGCP
NR
📍 Portland, Oregon, United States· Full-time
✓ High-confidence listingCompany trend -75%

From $98K/yr

Quick readStrong listing-quality and freshness signals

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity The Telemetry Data Platform group at New Relic builds the foundation for all of our products: data ingest, storage, and query. As an engineer working on NRDB, you’ll be contributing directly to the proprietary telemetry database technology at the core of our business. We own our software from top to bottom and are directly responsible for its quality and reliability. Each member of the team shares our pager rotation and will occasionally be on-call to respond to system failures; so we prioritize work that keeps the lights on and the pager quiet, in addition to the work that powers all of our new products and streams of data. If the idea of working on systems that process millions of messages per second and handle exabytes of data excites you, then you may be an excellent fit! What you'll do Develop new features with a focus on optimizing performance and efficiency Collaborate with the team to implement scalable solutions and enhance application performance Identifying and acting on opportunities to improve the reliability of our services This role requires 2+ years of professional experience in distributed SaaS software development. Proficiency in Java programming, expertise with algorithms and data structures, and building high-throughput software following best-practices. Deeper understanding of distributed systems and their core challenges. Experience using the command line to manage, investigate, and fix things when they’re broken. Expe

JavaSQLMySQLMongoDB
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team API Enterprise Controls is part of the API Infrastructure organization and owns the platform capabilities that help developers, startups, and enterprises adopt the OpenAI API securely and confidently. We build the systems underneath our APIs and developer platform across authentication and identity, service accounts and key management, secure networking, compliance, auditability, observability, and operational controls. Our users are developers and teams running critical applications on OpenAI, and we partner closely with Product, go-to-market, security, and infrastructure teams to turn their most important needs into reliable, intuitive platform capabilities. About the Role We are looking for an exceptional backend software engineer to help define and ship the enterprise capabilities our API Platform needs to scale.; this is a product-engineering role grounded in deep backend systems. You will work across databases, streaming systems, request routing, authentication, and developer-facing APIs while bringing strong product judgment, developer empathy, and attention to the small details that make a platform easier to understand, trust, and operate. You will lead large cross-functional initiatives, work closely with Product and go-to-market teams, engage directly with sophisticated users, and carry ambiguous needs from discovery through design, launch, and iteration. In this role, you will: Own backend product capabilities end to end across authentication and identity, service accounts and key controls, secure networking, compliance, observability, and operational workflows. Partner with Product, go-to-market, security, infrastructure teams, and sophisticated customers to identify needs, shape the roadmap, and lead large cross-functional projects from design through launch. Design developer-facing APIs, system behavior, configuration, error handling, safe defaults, auditing, and notifications with exceptional care for the details that define a great dev

TypeScriptPythonAWSRest
🔔

Get new platform architecture manager jobs in United States by email

Daily job updates · Unsubscribe anytime