The NVIDIA PerfTech team is looking for an outstanding Senior C++ Software Engineer to build the next generation of AI-powered developer tools. You will apply deep expertise in modern C++, systems architecture, performance, and debugging to extend our developer-tools ecosystem into the emerging world of agentic software. In this role, you will help develop Genie, NVIDIA’s company-wide AI knowledge and developer-productivity service. You will build production-quality C++ components and integrations while contributing to agentic workflows and retrieval systems that help engineers find information, understand complex systems, and work more effectively. What You’ll Be Doing: Architect and develop production-quality C++ components, APIs, and integrations for NVIDIA’s AI-powered developer-tools ecosystem. Build reliable, performance-sensitive systems connecting native C++ tools with Genie’s retrieval and agentic capabilities. Design maintainable interfaces between C++ applications and AI services exposed through Python, MCP, REST APIs, and structured tool calling. Contribute to agentic workflows, retrieval systems, ingestion pipelines, evaluation infrastructure, and enterprise integrations. Drive technical initiatives from early exploration through production deployment, working closely with multidisciplinary engineering teams. Mentor engineers and help establish best practices for production-quality AI-powered systems. What We Need to See: Bachelor’s or Master’s degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience. 5+ years of professional experience developing complex production software in modern C++. Strong knowledge of software architecture, concurrency, debugging, performance optimization, and maintainable API design. Exper
Jobiba hiring network
Systems Architect Jobs
10,000 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current systems architect jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking an experienced Principal Hardware Diagnostics Engineer to design and develop diagnostics software used to monitor hardware health and diagnose system-level issues across Graphcore’s AI infrastructure platforms. This role focuses on building diagnostics agents, tools, and analytics frameworks that enable engineers and automation systems to identify, isolate, and resolve hardware issues across blade-level servers and rack-scale clusters. The Team Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. The Systems Engineering and Platform Validation team ensures Graphcore’s AI compute platforms are reliable, diagnosable, and operationally robust at scale. The team co
About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus
About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that will run DoorDash's data, analytics, and ML compute for the next five years and beyond. You will own the deep, cross-cutting problems that span the runtime, the shuffle service, the scheduler, and the overall service reliability — making the architectural calls that compound across the platform's lifetime. You will partner with the Engineering Manager on technical roadmap, hiring, and team shape, and act as the senior technical voice in cross-team partnerships with Data Engineering, ML Platform, and product engineering teams that depend on the platform. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Set the multi-year technical direction for an in-house Spark-on-Kubernetes platform — runtime, shuffle, scheduler, reliability — and make the architectural calls that compound for years. Own the deepest distributed-systems problems on the team: shuffle architecture, multi-tenant scheduling, runtime performance, and the failure modes that only show up at scale. Partner with the Engineering Manager on technical roadmap, hiring, inte
MongoDB is seeking a Vice President, Global Specialization Leader to build and scale a world-class specialist organization focused on our highest-growth, highest-innovation product surfaces: Agentic applications, Voyage AI, Vector Search, and Analytics. This is a foundational, builder-oriented leadership role for someone who thrives at the intersection of emerging technology, early-stage product-market fit, and enterprise go-to-market execution. This leader will own the global specialist motion end-to-end — driving revenue, shaping product direction through frontline signal, and building a team of Specialist Sales and Specialist Solutions/Systems Architects that can sell, demo, and pilot next-generation AI-native workloads with speed and credibility. We are looking to speak to candidates who are based in Austin, NYC, Palo Alto, or San Francisco for our hybrid working model. What You'll Do Own a global number; Take full accountability for a consumption and/or bookings target across Agentic, Voyage AI, Vector Search, and Analytics, and build the plans, forecasting discipline, and pipeline hygiene needed to hit it predictably Build and lead a global specialization team; Recruit, manage, and develop a team of Specialist Sales leaders and Specialist Systems/Solutions Architects across regions, ensuring consistent coverage, quality, and career growth Operate like a startup inside MongoDB; These are early product-market fit motions — build the playbooks, qualification criteria, and champion strategies from scratch rather than inheriting mature process. Move fast, test, and iterate Drive fast, structured product feedback loops. Serve as the connective tissue between the field and Product/Engineering, translating customer and prospect signal into prioritized feedback that shapes the roadmap for Agentic, Voyage AI, Vector Search, and Analytics Own demo and pilot excellence; Ensure the team has best-in-class demo assets and a repeatable, fast pilot/PoC meth
About the Team The Monetization team is a new cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products—including next-generation ads experiences—that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to help build the core infrastructure behind OpenAI’s monetization and ads systems. In this foundational role, you’ll architect and implement distributed systems that power OpenAI’s monetization stack—focusing on reliability, performance, privacy, and large-scale operation. You’ll work across backend, systems, and platform layers to define and implement 0→1 infrastructure, partnering closely with Product, Design, and Research to shape the future of monetized AI experiences. Your work will enable both internal and external teams to build on safe, scalable, and robust monetization primitives. This role is exclusively based across our San Francisco & Seattles sites. We offer relocation assistance to new employees. In this role, you will: Design and build the foundational backend and infrastructure powering OpenAI’s monetization and ads systems Architect large-scale distributed systems that
About the Team The Cooperative AI team is scaling OpenAI with OpenAI. We are building an AI powered knowledge system that evolves and learns as our products, systems and customers evolve. We leverage our state of the art models, technologies, and products (some external, some still in the lab) to assist or completely automate robust operations supporting both internal and external customers. We support OpenAI customers and internal partners globally, powering systems from customer support to integrity to product insights. We are a self-contained multi-disciplinary team, who enjoy a lightning fast feedback loop with customers at scale, some of whom sit just a few pods away. We iterate fast, and engineer for reliable long-term impact. We're constantly looking for the similarities and patterns in different types of work, and focus on building simple primitives, to apply world class knowledge to many domains. The work of this team exemplifies use of OpenAI technologies. We build systems so everyone can see the leverage that is possible with well designed AI-based implementations. We do this by working through internal use cases focused on Customers (specifically knowledge systems, automation systems, and automated agent systems) to prove impact, then we scale. About the Role We’re looking for Software Engineers who're passionate about blending production-ready platform architecture with new tech and new paradigms. You’ll push the boundaries of OpenAI’s newest technologies to enable interactions and automations that are not only functional, but delightful. We value proactive, customer-centric engineers who can get the foundational details right (data models, architecture, security) in service of enabling great products. In this role, you will: Own the end-to-end development lifecycle for new platform capabilities and integrations with other systems Collaborate closely with engineers, data scientists, information systems architects, and internal customers to understand th
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity We're looking for a Senior Director/Director, GTM Tools & Technology to own the systems and tooling that power Postman's entire go-to-market organization. This is a strategic leadership role sitting at the intersection of business operations, systems architecture, and AI innovation. You'll be the internal product owner for our GTM tech stack — from Salesforce at its core, to every platform that plugs into it, across both pre-sale and post-sale. You'll define the vision and drive the roadmap for how we use technology and AI to make our GTM teams faster, smarter, and more effective. This role demands both strategic range and operational depth: you'll partner directly with GTM executives across Sales, Marketing, Customer Success, Professional Services, Revenue Operations, and Finance, while leading a team to deliver on a high bar. The standard we're hiring to is high. We want someone who can architect for where Postman is going, not just maintain what exists — and who will lead meaningfully on AI adoption as a first-class mandate. What You’ll Do GTM Systems Strategy & Roadmap Define and execute the roadmap for
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Forward Deployed Infrastructure Engineers (FDIEs) build, operate, and maintain the infrastructure that powers Palantir’s platforms and production deployments. As an FDIE intern, you’ll work alongside full-time FDIEs to deploy and operate Palantir software across real production environments, automate manual processes, and develop novel solutions to infrastructure challenges using tools like Foundry and Apollo. Every day looks different — you might be debugging a distributed systems issue, building automation to replace a manual runbook, or designing infrastructure improvements that scale across multiple deployments. You’ll be treated as a full member of the team, with real ownership over the work you take on. Core Responsibilities As an FDIE intern, your responsibilities look similar to those at a small startup, with the resources, stability, and mentorship of an established tech company. You’ll work in small teams with minimal supervision and own end-to-end execution of real infrastructure projects. Your day might span discussing systems architecture with fellow engineers, debugging a production issue, building automation to eliminate a manual process, or deploying new Palantir products across production environments. FDIE interns are treated just like full-time engineers, with significant freedom and ownership over their work. Specifically, you can expect to: Deploy and operate Palantir software across production environments, including monitoring, alerting, configuration management, and upgrades Debug, improve, and optimize Palantir’s services and infra
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Forward Deployed Infrastructure Engineers (FDIEs) build, operate, and maintain the infrastructure that powers Palantir’s platforms and production deployments. As an FDIE intern, you’ll work alongside full-time FDIEs to deploy and operate Palantir software across real production environments, automate manual processes, and develop novel solutions to infrastructure challenges using tools like Foundry and Apollo. Every day looks different — you might be debugging a distributed systems issue, building automation to replace a manual runbook, or designing infrastructure improvements that scale across multiple deployments. You’ll be treated as a full member of the team, with real ownership over the work you take on. Core Responsibilities As an FDIE intern, your responsibilities look similar to those at a small startup, with the resources, stability, and mentorship of an established tech company. You’ll work in small teams with minimal supervision and own end-to-end execution of real infrastructure projects. Your day might span discussing systems architecture with fellow engineers, debugging a production issue, building automation to eliminate a manual process, or deploying new Palantir products across production environments. FDIE interns are treated just like full-time engineers, with significant freedom and ownership over their work. Specifically, you can expect to: Deploy and operate Palantir software across production environments, including monitoring, alerting, configuration management, and upgrades Debug, improve, and optimize Palantir’s services and infra
About the role: We are seeking a Senior Backend Engineer with deep backend engineering expertise and proficiency in one or more major programming languages (e.g., Python, Java, Go, Rust, or Kotlin), along with a strong understanding of AI models and agents. As a core member of our AI Engineering team, you will collaborate with data scientists, ML engineers, and product managers to build scalable, production-ready infrastructure and APIs that power intelligent systems. What you'll be doing: As a Senior Backend Engineer in the AI Engineering team, you will: Build and maintain reliable, scalable backend services to support AI agent execution and orchestration. Develop AI agent systems for complex operational workflows using LangChain, LangGraph, LiteLLM, and Langfuse. Orchestrate a hybrid model stack that includes OpenAI and Google Gemini alongside self-hosted and fine-tuned LLMs like Gemma and Llama. Build and maintain integrations with clinical systems (FHIR, EMR). Drive observability and reliability using OpenTelemetry, Datadog, and Langfuse. Design APIs (GraphQL, REST), background workers, and event-driven systems that interface with AI inference engines and agent runtimes. Collaborate with Data Science, ML, and engineering teams to deploy AI features and improve the performance, scalability, and reliability of backend systems. Participate in code reviews, knowledge sharing, and mentoring to elevate the team’s technical capabilities. What we're looking for: 6+ years of backend engineering experience, with strong proficiency in more than one major programming language (such as Python, Java, Go, Rust, or Kotlin). Solid understanding of AI systems architecture and experience working in environments involving AI agents, LLMs, or inference pipelines. Proven experience in building and scaling backend APIs, microservices, and background jobs. Strong experience with relational and NoSQL databases (e.g., PostgreSQL, MySQL, MongoDB, Redis), including schema des
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE Baseten's Compute org is in hyper growth. As it scales, the systems and workflows that keep supply and demand balanced across our GPU fleet need to get more sophisticated, and this role exists to make sure they do. Compute sits at the center of how Baseten allocates, forecasts, and manages the capacity that powers every customer inference request. The team that supports this work, C3, runs on a mix of internal tooling, manual processes, and systems that haven't fully kept pace with the scale of the problem. This role exists to close that gap. You'll design, build, and ship AI-powered workflows that give the Compute and C3 teams real leverage, automating the manual, repetitive, and error-prone parts of the capacity lifecycle so the team can focus on judgment calls that actually need a human. We want someone who can walk in, audit what exists today, identify what's missing or broken, and start shipping fast. You know when to reach for an existing internal tool and when to build something custom in Claude Code. You think two to three steps ahead about how the thing you build today fits into the broader capacity systems architecture tomorrow. And you bring a point of view on our stack, on what we should be building, and on where AI can do something existing tooling simply can't. RESPONSIBILITIES Ship AI-powered workflows for Compute and C3 : build the agents and automations that give capacity analysts, ops leads, an
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. The Role We are looking for an Engineering Leader to manage and scale multiple product lines in the Voice, BPO, and Workforce Management space. This is a high-impact leadership role that sits at the intersection of real-time voice systems, operations research, and data-intensive platform engineering. You will report directly to the Head of Engineering and own the engineering organization that builds the infrastructure powering Ema’s Voice AI Employees, Agent QA, auto-learning pipelines, rich analytics, and workforce optimization capabilities — all operating as scalable, multi-tenant systems deployed across global geographies. You will collaborate with Product, ML/AI, and Go-to-Market teams to translate customer needs into production systems that handle high volumes of voice data, deliver real-time insights, and continuously improve through automated learning loops. As the owner of multiple product lines, you will balance roadmap priorities across Voice, BPO operations, and WFM (work force management) — ensuring each product evolves cohesively while meeting distinct customer needs. What You Will Do Scalable Multi-Tenant Systems Architect and build multi-tenant systems that serve ent
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the FlashArray team to build high-performance, resilient storage software that powers mission-critical applications worldwide. As a core member of an agile engineering group, you will design and deliver zero-downtime algorithms that directly shape our enterprise storage and public cloud offerings. In this role, you will partner closely with global systems engineers and product teams to translate complex technical challenges into scalable, real-world solutions. Your work will directly impact how thousands of global enterprises manage, protect, and scale their data seamlessly. WHAT YOU’LL DO Design & Deliver Resilient Systems: Architect and implement high-performance algorithms for enterprise storage products, ensuring platform reliability, end-to-end delivery from concept to release, and six-nines availability. Expand Cloud Architecture: Extend core platform capabilities into public cloud environments (such as AWS), driving performance and agility for both traditional IT and cloud-native applications. Drive Problem Solving & Quality: Analyze and resolve complex systems software challenges, optimizing storage internals and contributing to continuous continuous platform upgrades. Collaborate & Mentor: Partner with multidisciplinary engineering peers to review code, refine architecture, and maintain high engineering standards across distributed software projects. WHAT YOU BRING Systems Programming Exp
We're looking for a Senior Operations Director for Revenue & Services who will serve as a trusted advisor to the Chief Commercial Officer, Professional Services leadership and other senior leaders. This role will shape how Professional Services operates, scales, forecasts, and performs, while driving alignment across Operations, Revenue, Sales and Finance. This role will connect services to broader business strategy, with a focus on cross-functional alignment, profitable growth and corporate results. WHAT YOU’LL DO: Shape and evolve the operating model for Revenue and Services, includingProfessional Services to enable scalable, predictable and profitable growth: forecasting methodology, capacity planning, utilization and margin management, and delivery-to-revenue alignment Set the strategic roadmap for services operations and operating infrastructure, including systems architecture (resource management, CRM integration, financial systems, data and reporting) and where to invest next Act as the strategic partner to Professional Services leaders translating business strategy and operating plans into performance targets. Represent services performance directly to executive leadership and the board Partner with President, Chief Commercial Officer, Finance leaders to evaluate profitable and growth opportunities, operating leverage, organizational capabilities, technologies and external partnerships. Lead deal desk strategy for services and deal governance, including pricing frameworks, SOW structuring, and margin guardrails, in partnership with Sales and Legal Partner in quarterly and annual planning for services capacity and revenue targets, and defend the plan directly to Finance and executive leadership Establish and drive a consistent operating cadence and process framework, including KPIs, business reviews, planning, forecasting, and reporting across the full customer lifecycle, identifying where PS operations should own versus influence
Get new systems architect jobs by email
Daily job updates · Unsubscribe anytime