ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. At Baseten, we are building the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale, the network is the computer. We are looking for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building block in our infrastructure and unlocking the next generation of distributed inference optimizations. THE OPPORTUNITY Networking and compute are no longer separate disciplines; they are converging. The massive throughput of H100, B200, and NVL72 architectures enables and demands a new approach where communication is co-optimized alongside computation. We are entering an era where the network is an active accelerator, leveraging smart hardware offloads and direct interconnects to ensure that data movement operates at wire-speed. In this role, you will go beyond network configuration to architect the software fabric that unifies thousands of GPUs into a cohesive operating system. While you will leverage the best of the open-source ecosystem, you won't be limited by it. Where off-the-shelf solutions stop, you will build from scratch, engineering the primitives required to co-optimize communication and compute for Disaggregated Serving, Wide Expert Parallelism (WideEP), and lightening cold starts. WHAT YOU'LL DO Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack,
Jobs in United States
Software Architect in United States
2,125 active opportunities · Updated October 2026
Showing
15 jobs
Explore current software architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview Join a team responsible for building the authentication and security solutions that help protect digital interactions across Mastercard's Identity Solutions platform. As a Senior Software Engineer, you will design, develop, and support secure, scalable, and high-performing applications that enable trusted identity verification, authentication, and fraud prevention capabilities. In this role, you will take ownership of complex technical challenges, contribute to software design and architecture decisions, and partner closely with product, security, and platform teams to deliver reliable, production-ready solutions. You'll play a key role in advancing engineering excellence through secure development practices, system reliability, automation, and continuous improvement while mentoring other engineers and influencing technical direction across the team. Role •Design, build, test, deploy, and maintain scalable, cloud-native applications and microservices •Develop REST APIs using Java and Spring Boot, focusing on performance, scalability, and reliability •Translate requirements into well-structured designs and architecture, ensuring maintainability and security •Lead and contribute to system design discussions, aligning with architectural standards and best practices<
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. As a Senior Security Engineer at Vanta, you’ll own projects with impact across the business to help us run an efficient and highly effective security team. The security team at Vanta ensures that we are a trustworthy steward of sensitive data. We also contribute subject matter expertise to the product, sales, marketing, support, and engineering functions, given the nature of our business. You’ll join Vanta’s Security organization, which provides essential security operational services, is directly involved in the software development process and building tools to make it easy for developers to ship products securely, sets policies and standards regarding enterprise-wide security requirements, and offers advisory services to enable our business to thrive while effectively managing risk. If you’re someone who has high initiative and enjoys problem solving while having impact at a high-growth company, we would love to hear from you! What you’ll do as a Senior Security Engineer at Vanta: Participate in team exercises to identify potential security risks, including threat modeling and tabletop scenarios Contribute to complex prioritization discussions around which risks are the most important to solve next Plan projects to address the risks we prioritize, and coordinate with cross-functional stakeholders across the company to execute those projects Build maintainable programs to implement operational excellence where ongoing work is needed to achieve our goals (e.g. vulnerability management) Partner with engineering teams to architect secure software, address security concerns, and build a strong security culture Build, customize, a
We’re looking for a technical architecture expert with hands-on development experience to define the architecture of the Retail and Consumer Packaged Goods Platform and collaborate with engineering to build the platform. This is a technical engineering role and requires familiarity with NVIDIA’s GPU acceleration libraries, experience in Retail, CPG, logistics, or marketplaces, and knowledge of enterprise AI architectures will make you an ideal candidate. Join us and take the lead in building the retail platform and help us develop the future of AI in retail! Your mission is to leverage NVIDIA’s existing retail and supply chain blueprints, along with proven workflows and best practices codified through customers and ISV engagements, to define the architecture and build the Retail & CPG platform. You will be responsible for architecting and collaborating with engineering teams across NVIDIA, to build the platform. Built on NVIDIA’s full stack—including Agentic AI, Physical AI, and accelerated computing—this platform will power our three strategic pillars: digital commerce, supply chain, and intelligent stores. What you'll be doing: You will work with Retail & CPG product management, business development, and developer relations teams to harness their deep expertise in supply chain and retail to architect and build a platform that makes it easier for our ecosystem and customers to build AI solutions at scale. The platform will use our NVIDIA Nemo microservices and Nemotron open-source models, with initial priority on Agentic AI for supply chain, digital commerce and employee productivity. This role will initially be an individual contributor with developer experience building AI agents and will collaborate with a strong technical retail and supply chain team, as well as horizontal engineering team(s). Define product, technical vision and architecture Develop platform s
$295K – $380K/yr
About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Senior Software Engineer, ML Systems & Training Infrastructure, you will be a deeply hands-on engineering force multiplier for the robotics team. You will help keep the training framework and surrounding infrastructure healthy, review and improve code quickly, debug failures across ML systems and infrastructure, and unblock researchers and engineers when the path from idea to working training job gets rough. We’re looking for people who love writing, reading, reviewing, and fixing code; who can get productive quickly in unfamiliar systems; and who bring strong practical judgment without a lot of ego or process overhead. This role will be based in San Francisco, CA and be expected in office 5 days per week and offer relocation assistance to new employees. In this role, you will: Review, improve, and clean up code across training frameworks and adjacent infrastructure. Identify risky or low-quality changes before they land, and raise the code quality bar without slowing the team down. Debug issues across ML training systems, GPUs, clusters, networking, and related infrastructure. Help researchers and engineers unblock broken training jobs, flaky workflows, and brittle internal tooling. Improve the reliability, maintainability, and usability of the robotics team’s training framework. Move quickly on practical engineering problems that directly affect team velocity. You might thrive in this role if you: Have strong software engineering fundamentals and excellent code review judgment. Have experience with ML systems, training fr
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore fosters continuous learning and innovation. Job Summary Reporting into the Systems Engineering organisation, the Distinguished Engineer, End-to-End Security Architect will define and lead the security architecture for Graphcore’s inference service platform. This role is responsible for establishing a comprehensive security strategy spanning platform, infrastructure, networking, service operations, customer assurance, and compliance readiness. Working across multiple engineering and operational functions, the successful candidate will provide technical leadership, drive security requirements, and ensure the platform delivers robust protection, resilience, and trust for customers. The Team You will work closely with teams across security architecture, infrastructure engineering, networking, site reliability engineering, platform software, firmware, data centre operations, compliance, legal, customer engineering, and customer security. The team collaborates across the business to deliver secure, reliable, and scalable AI infrastructure and services while supporting customer assurance, regulatory requirements, and operational excellence. Responsibilities and Duties Own the end-to-end security a
Become a part of our caring community The Lead Software Engineer codes software applications based on business requirements. The Lead Software Engineer works on problems of diverse scope and complexity ranging from moderate to substantial. The Lead Software Engineer standardizes the quality assurance procedure for software. Oversees testing and debugging and develops fixes. Researches complaints and makes necessary adjustments and/or recommendations to resolve complex software related issues. Advises executives to develop functional strategies (often segment specific) on matters of significance. Exercises independent judgment and decision making on complex issues regarding job duties and related tasks, and works under minimal supervision, Uses independent judgment requiring analysis of variable factors and determining the best course of action. Key Responsibilities Technical Architecture and Ownership:** Design and own the end-to-end architecture of Centerwell's AI systems, including LLM-powered clinical tools, RAG pipelines, harnesses, agent-based workflows, and intelligent automation. Make and communicate foundational technical decisions in close collaboration with the broader engineering team. Model Development and Fine-Tuning:** Evaluate, select, and where appropriate guide the fine-tuning of foundation models. Establish model evaluation frameworks that prioritize safety, accuracy, and clinical relevance. Clinical and Product Partnership:** Collaborate closely with product managers, designers, clinicians, and data stakeholders to understand care delivery workflows and translate them into well-scoped, high-impact AI features. HIPAA Compliance and Responsible AI:** Ensure all AI systems are designed, deployed, and monitored in compliance with HIPAA and Humana's Responsible AI standards, including participation i
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are looking for an AI Agent Security Architect to function as the primary technical authority for Replit’s autonomous and AI agent security blueprint. In this critical role, you will design, implement, and maintain the runtime defense systems, guardrail frameworks, and sandboxing architectures that govern AI agents executing code, invoking tools, and reasoning across our platform. You will be a key technical contributor—leading high-impact AI security initiatives and bridging the gap between non-deterministic AI behavior and rigorous cybersecurity controls for both engineering and executive leadership. What You'll Do AI Agent Security Strategy & Technical Execution AI Agent Security Blueprint: Define the long-term vision and architectural patterns for securing autonomous agent workflows, Model Context Protocol (MCP) integrations, multi-turn reasoning loops, and multi-agent coordination. Runtime Guardrails & Policy Enforcement: Architect and deploy dynamic input/output guardrail systems, semantic firewalls, and real-time intent verification filters to prevent goal hijacking, system prompt leaks, and indirect prompt injections. Agent Execution & Tool Sandboxing: Partner with Infrastructure and AppSec teams to design secure, short-lived, micro-isolated environments (e.g., microVMs, WebAssembly, container sandboxes) where agents can dynamically execute code, run shell commands, and interact with host operating systems safely. Agentic Threat Modeling & Red Teaming: Conduct specialized threat modeling against non-deterministic systems. Lead automated and manual AI red-teaming initiatives to uncover vulnerabilities in RAG context pipelines, vector stores, and tool-calling interfaces. Identity
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
From $265K/yr
We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacarts Data Infrastructure organization builds and operates the systems that power our company’s data ecosystem, including a modern open data lakehouse on Apache Iceberg, a multi-engine compute platform for stream and analytical workloads, and self-serve tooling that helps Product, Data Science, ML, Ads, Finance, and engineering teams move fast with data. We’re looking for a Staff Software Engineer, Data Infrastructure to join our Data Governance and Foundations Team. In this role, you’ll serve as a senior technical leader owning the architecture and delivery of our open lakehouse foundation, governance and access patterns, and multi-engine compute strategy—balancing today’s reliability with the next three to five years of scale, maturity, and cost efficiency. You’ll collaborate closely with engineering leadership and stakeholders across Data Science,
SUMMARY STATEMENT We are looking for a Solution Architect to design the technical solutions behind our client engagements and give delivery teams a clear, workable path from concept to production. You will work across enterprise data, software applications and GenAI - translating complex business problems into practical architectures that delivery teams can build and scale. This could include architecting an agentic workflow for clinical operations, a conversational analytics product grounded in enterprise data, or an AI-enabled decision platform for commercial teams. You will work directly with clients, define the architecture, test the most important technical decisions yourself and establish the foundations for successful delivery. This is an architecture-first role with meaningful hands-on engineering: you will stay close enough to implementation to prove the architecture works and support it through production delivery, without becoming the primary engineer for every component. You will also help shape the reusable patterns, technical standards and accelerators behind Lynx’s growing AI-native life sciences practice. KEY RESPONSIBILITIES Solution Architecture Own the end-to-end solution architecture for client engagements, including data models, system design, integration patterns and technology choices. Translate business requirements into clear technical designs and implementation paths that delivery teams can build from. Design solutions spanning enterprise data, APIs, applications, cloud platforms and GenAI capabilities. Lead technical discovery with clients: understand requirements, assess existing systems and identify dependencies, constraints and delivery risks. Present architectural options and trade-offs clearly to technical teams, business stakeholders and senior leaders. Make pragmatic decisions across build speed, cost, scalability, security and maintainability. Review key implementation decisions and remain
Job Requisition ID # 26WD100611 Position Overview Autodesk is a global leader in design and make technology, with expertise across architecture, engineering, construction, design, manufacturing, and entertainment. Our software and services empower innovators everywhere to solve challenges big and small—from greener buildings to smarter products to more compelling media and entertainment. At Autodesk, we believe that when you have the right tools to work and think flexibly, you have the power to transform what actually needs making. We provide our customers with technology to help them achieve better outcomes for their products, businesses, and the world. We are looking for a highly motivated software engineer to join the PSET-Access group. The team builds and operates systems that enable secure, reliable, and scalable access to product capabilities and resources, and now in the process to develop the next generation to modernize the access management capabilities. In this role, you will work with engineers and cross-functional partners to design, build, test, deploy, and operate production software. You will primarily develop backend services using Java or Go, while owning well-defined components and projects, contributing to technical decisions, and helping improve the reliability, maintainability, and developer experience of the systems the team owns. This is a strong opportunity for an engineer who has solid software engineering fundamentals and is ready to grow their technical depth and ownership. Responsibilities Design, implement, test, deploy, and maintain production-quality backend software, prim
From $195K/yr
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Community Support Engineering (CS Eng) is responsible for the world-class technology, architecture, and solutions that power Airbnb’s global community support operations. The products and solutions we build support our customers, front line agents, and business process operations teams. They are a key driver of enabling Airbnb’s core business. As a member of CS Eng, you will have an opportunity to impact every person who travels or hosts on Airbnb. Under CS Eng, the Agent Core Products team has an opportunity for a Senior Engineer to help drive our initiatives. The team builds the core software that empowers our global support agents. The work has a direct impact on agent experience and impacts the quality of service we provide to our guests and hosts. The Difference You Will Make: We are seeking a highly skilled and motivated engineer who is passionate about making a difference through their work. As a Senior Software Engineer, you will work in a team of talented and diverse software engineers to build solutions that improve the agent experience. You will play a significant role in shaping the technical vision and then delivering a solution that is flexible, efficient and scales with the needs of the business. Each individual brings their own unique skill set, experiences, thought leadership and technical expertise to solve these technical challenges for Airbnb. We are a high-impact team focused on driving double-digit percentage improvements across our key metrics through various strategic projects. We continuously advance the quality and efficiency of our product, emp
Overview: The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Software Engineer to join our Data Acquisition team. Responsibilities: Own and lead engineering projects in the area of data acquisition including web crawling, data ingestion, and search. Collaborate with other sub-teams, such as Data Processing, Architecture, and Scaling, to ensure smooth data flow and system operability. Work closely with the legal team to handle any compliance or data privacy-related matters. Develop and deploy highly scalable distributed systems capable of handling petabytes of data. Architect and implement algorithms for data indexing and search capabilities. Build and maintain backend services for data storage, including work with key-value databases and synchronization. Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks. Conduct and analyze experiments on data to provide insights into system performance. Qualifications: BS/MS/PhD in Computer Science or a related field. 4+ years of industry experience in software development. Experience with large web crawlers a plus Strong expertise in large stateful distributed systems and data processing. Proficiency in Kubernetes, and Infrastructure-as-Code concepts. Willingness and enthusiasm for trying new approaches and technologies. Ability to handle multiple tasks and adapt to changing priorities. Strong communication skills, both written and verbal. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an
About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun
Other cities to consider
More places hiring for this role
Get new software architect jobs in United States by email
Daily job updates · Unsubscribe anytime