Jobiba hiring network

Mission Systems Engineers Jobs

6,816 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current mission systems engineers jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently. You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. EXAMPLE INITIATIVES You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services. Building AI-assisted tooling for incident triage and response. RESPONSIBILITIES Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking. Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code. Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution. Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations. Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, con

kubernetesgitmachine learning
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for a Customer Marketing Manager who can own the full customer evidence motion at Baseten: building the systems that capture customer stories, running the co-marketing programs that amplify them, and developing the channels and assets that get those stories in front of the right people. Our customers are ML engineers and AI teams deploying serious workloads — and the stories they tell about what they've built matter. We've earned trust with some of the most demanding technical teams in the industry, and this role exists to turn that trust into evidence. RESPONSIBILITIES Co-Marketing Execution Serve as the DRI for every customer co-marketing launch end to end — managing timelines, coordinating internal and external stakeholders, and driving the process from first outreach to final publication Own the single source of truth for what's in flight across all customer co-marketing activity Coordinate with design, social, and sales to ensure every asset is built, approved, and distributed correctly Customer Evidence & Asset Library Own the customer evidence library: written case studies, video stories, customer quote repository, logo library, and sales snippets ensuring all assets stay current and are tagged and accessible for sales and marketing use Run the monthly operating rhythm: new logo additions from closed-won opportunities, asset updates, and customer health checks Programs & Channels I

machine learningaigo
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Platform team within the Infrastructure organization. As a senior member of Baseten's Platform Team, you will own the systems that let every engineer at Baseten prove their code works before it reaches production. Our product runs mission-critical AI inference for customers who measure downtime in dollars per second, which means our internal bar for correctness, performance, and failure tolerance has to be exceptional. Your focus is the full testing stack: fast and reliable unit test tooling, integration harnesses that spin up realistic environments on demand, load and performance testing for GPU-backed inference workloads, and resilience testing that deliberately breaks things so our customers never have to find out what happens when a node dies mid-request. This is a builder role with org-wide leverage. You won't be writing tests for other teams — you'll be building the frameworks, harnesses, and feedback loops that make writing good tests the path of least resistance, and you'll set the standards for what "well-tested" means at Baseten. RESPONSIBILITIES Own Baseten's testing strategy end to end — define the standards, the tiers, and the tooling that engineering teams build against. Build and maintain unit, integration, load and performance testing frameworks Design end to end test infrastructure that provisions realistic dependencies

pythondockerkubernetes
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab

kubernetesmachine learningai
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use. As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently. This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users. EXAMPLE INITIATIVES Blog Posts https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/ https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/ https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid RESPONSIBILITIES Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference Work across the stack, from customer-facing features to low-le

kubernetesci/cdrest
View job →
B
Baseten
📍 New York• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for a Marketing Operations Manager who can own and harden the systems layer of Baseten's Marketing engine. Marketing at Baseten is scaling fast — more spend, more campaigns, more model launches, more inbound. The systems underneath (Our ESP, CRMs, forms, tracking, routing, alerting) need an owner who treats them like production infrastructure that cannot go down. When something breaks, it costs us time, pipeline, and trust in the data. This role exists so it doesn't break. You’ll simultaneously build for the future and re-think assumptions about our tech stack in the age of agents. This is an offensive play that gives the rest of the team leverage and superpowers to hit our ambitious goals. This is NOT an IT or service role. This is a core member of the marketing team who implements technology to achieve outcomes. RESPONSIBILTIES Own the marketing tech stack end-to-end: ad platforms, email systems, tracking, pixels, forms, connectors. Build defense-in-depth on inbound: spam/bot protection, rate limiting, email/domain validation, sync gating — and the alerting to catch anomalies before they hit sales or leadership dashboards. Enforce data integrity: UTM governance, campaign membership, lifecycle stages, lead scoring and routing logic, field-level hygiene, canonical metric definitions. Operationalize the web request pipeline with our dev agency: structured briefs, tickets, SLAs, and launch-day runb

pythonrestmachine learning
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Engineering Manager for Baseten's Cloud Platform team, you will directly manage a team of cloud platform engineers responsible for building the systems and processes that keep our infrastructure scalable, reliable, and efficient — from automated deployments and monitoring to performance optimization and incident response. You are a people-first leader with a strong cloud infrastructure background. You set a high bar for reliability and operational excellence, engage credibly in technical discussions and code reviews, and know how to build a culture of ownership and accountability. You'll spend most of your time close to the work: unblocking your team, shaping technical direction on day-to-day decisions, and developing your engineers. At Baseten, we work closely with our users to understand their struggles operationalizing ML — you'll keep your team connected to that mission and translate user learnings into better infrastructure. RESPONSIBILITIES Recruit, hire, and grow a high-performing team of cloud platform engineers; provide ongoing coaching, feedback, and career development through regular 1:1s. Set clear performance expectations, hold a high bar, and create an environment where engineers do their best work. Foster a culture of ownership, accountability, and continuous improvement. Drive day-to-day technical decisions through design reviews, code reviews, and architectural discussions; translate th

kubernetesci/cdgit
View job →

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Engineering Manager (Player & Coach), you will lead and mentor a team of Forward Deployed Engineers focused on building, scaling, and optimizing LLM inference workloads for Baseten customers. Applying both hands-on technical ownership and managerial leadership, you will guide your team through the processes of designing, deploying, and managing high performance, low latency AI applications on Baseten’s platform. FDE at Baseten is not a sales function – we are a mix of engineering, product, and customer architects who contribute to the core Baseten codebase, drive large portions of our feature roadmap, and execute on complicated customer engagements. You will also partner with product, infrastructure, and other customer engineering teams to ensure that large language models (LLMs) and other generative AI systems deliver best-in-class performance, reliability, and cost efficiency in production environments. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Leadership & Team Management Lead, mentor, and grow a team of Forward Deployed Engineers, providing guidance on technical direction, project execution, and professional deve

pythondockermachine learning
View job →

At Datadog, we're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. APM at Datadog is on its way to redefine how users interact with their telemetry. We are integrating intelligence directly into troubleshooting workflows to help engineers find root causes faster, navigate complex distributed systems seamlessly, and optimize application performance with minimal cognitive load. APM provides deep visibility from end-user interactions to backend services and we are now expanding this foundation with new AI-driven insights, guidance, and automation. As a Product Manager II for APM, you will work with world-class engineers, designers, and partner product teams to shape the future of Distributed Tracing, Performance Analysis, and Intelligent Troubleshooting. You will help build advanced capabilities that scale to thousands of customers and make sophisticated observability workflows accessible to every engineer, from experts to beginners. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What you will do: Develop a deep understanding of APM customers, their performance challenges, telemetry workflows, and competitors Lead conversations with design partners and strategic customers to uncover real-world performance issues, validate product assumptions, and guide solutions from early prototypes through General Availability Define and deliver the next generation of APM features with engineering and design, especially agentic on

microservicesaigo
View job →

At Datadog, we're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. APM at Datadog is on its way to redefine how users interact with their telemetry. We are integrating intelligence directly into troubleshooting workflows to help engineers find root causes faster, navigate complex distributed systems seamlessly, and optimize application performance with minimal cognitive load. APM provides deep visibility from end-user interactions to backend services and we are now expanding this foundation with new AI-driven insights, guidance, and automation. As a Product Manager II for APM, you will work with world-class engineers, designers, and partner product teams to shape the future of Distributed Tracing, Performance Analysis, and Intelligent Troubleshooting. You will help build advanced capabilities that scale to thousands of customers and make sophisticated observability workflows accessible to every engineer, from experts to beginners. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What you will Do: Develop a deep understanding of APM customers, their performance challenges, telemetry workflows, and competitors Lead conversations with design partners and strategic customers to uncover real-world performance issues, validate product assumptions, and guide solutions from early prototypes through General Availability Define and deliver the next generation of APM features with engineering and design, especially age

microservicesaigo
View job →

At Datadog, we’re on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Observability Data Platform (ODP) is the backbone of everything Datadog delivers – powering how data is ingested, stored, routed, and surfaced across every product at planet scale. As a Senior Product Manager for ODP, you will work with world-class engineers and cross-functional partners to shape how the platform is deployed, controlled, and operated. You will define product direction across the control plane and data layer, translate complex infrastructure trade-offs into clear roadmap decisions, and help customers get the most from their observability investment – regardless of architecture, topology, or scale. At Datadog, we place value in our office culture – the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You Will Do: Develop a deep understanding of the Observability Data Platform customers – platform engineers, SREs, and product managers that own the product verticals – their infrastructure challenges, deployment topologies, and cost-to-serve trade-offs. Define product direction across multiple ODP surfaces, including the control plane and data layer, by articulating clear problem statements and desired outcomes, and partnering with engineering on technical approach and sequencing Lead conversations with design partners and strategic customers to understand real-world platform pain points, validate product assumptions, and guide solutions from early prototypes through General Availability Develop a co

aigorust
View job →
M
Mongodb
📍 Dublin• Full-time
1mo ago

MongoDB Technical Services Engineers use their exceptional problem solving and customer service skills, along with their deep technical experience, to advise customers and to solve their complex MongoDB problems. Technical Service Engineers are experts in the entire MongoDB ecosystem - database server, drivers, cloud and infrastructure. This also includes services such as Atlas (database as a service), or Cloud Manager (which helps customers with automation, backup and monitoring of their MongoDB systems). Our engineers combine their MongoDB expertise with passion, initiative, teamwork and a great sense of humor to achieve exceptional results for our customers. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Why MongoDB is a fantastic place to work and build your career Be a part of the company that’s reinventing the database, passionate about innovation and speed Enjoy a fun, inspiring culture that is engineering focused Work with exceptionally talented people around the globe Learn, contribute, and make an impact on the product and community Cool things you’ll do MongoDB is on a mission to change the way people think about databases. Along the way, our customers encounter questions and issues about how our approach to databases works for their use case. In Technical Services, it's our job to help these people. You'll be working alongside our largest customers, solving their complex issues - resolving questions on architecture, performance, recovery, security, and everything in between. You'll be an expert resource on standard methodologies in running MongoDB at scale, whatever that scale may be. You'll be an advocate for customers' needs - working with our product management and development teams on their behalf. And you'll contribute to internal projects, including software development of support tools for performance, benchmarking, and diagnostics. In addition, you will also be responsible for mentoring and ramping new

pythonjavasql
View job →
M
Mongodb
📍 Palo Alto• Full-time• From $90K/yr
1mo ago

MongoDB Technical Services Engineers for Named Accounts, as part of the Premium Services team within Technical Services, will use their exceptional problem solving and customer service skills, along with their deep technical experience, to advise customers and solve their complex MongoDB problems. They are experts in one or more components of the MongoDB ecosystem - database server, drivers, our management suite, and services such as Cloud Manager (the online product we developed for customers for automation, backup, monitoring, and analysis of their MongoDB systems), and MongoDB Atlas. Our engineers combine their MongoDB expertise with passion, initiative, teamwork, and a great sense of humor to achieve exceptional results for our customers. We're looking to speak with candidates based in San Francisco or Palo Alto for our hybrid working model. Cool things you’ll do MongoDB is on a mission to change the way people think about databases. Along the way, our customers encounter questions and issues about how our approach to databases works for their use cases. In Technical Services, it's our job to help these people. You'll be primarily working alongside a handful of our largest enterprise customers - building a relationship, intimately guiding their use of MongoDB products and services, coordinating the resolution of their complex issues, - answering questions on architecture, performance, recovery, security, and everything in between. You'll be an expert resource on best practices for running MongoDB at scale, whatever that scale may be. You'll be an advocate for customers' needs, - interfacing with our product management and development teams on their behalf. What you need You should have 7+ years of database industry experience deploying and managing operational production databases at scale, both on premise and in the cloud. We encourage you to apply even if you’ve never used MongoDB before. We consider all candidates with an eye for those who are sel

javascriptpythonjava
View job →
I
Instacart
📍 Canada - Remote (ON, BC• Full-time• Remote• From C$168K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview About the Role We are currently seeking a Senior Software Engineer to join our Agentic Analytics Platform team — the team responsible for the AI-for-Data charter inside Instacart's Data Infrastructure org. You'll design and build LLM-powered systems that transform how data practitioners (data scientists, data engineers, analysts, PMs) interact with data at Instacart — from natural-language data access and AI-assisted SQL, to automated metadata generation, to embedding intelligent capabilities across our broader data infra ecosystem. This is a hands-on role at the frontier of applied AI inside a large, modern data stack. About the Team The mission of the Instacart Self-Serve organization is to improve the productivity of data practitioners through easy-to-use, self-serve tools. Agentic Analytics is the team chartered with bringing AI and LLMs into that miss

REMOTEpythonsqlai
View job →
O
1mo ago

About the Team OpenAI’s mission is to ensure that general-purpose artificial intelligence benefits all of humanity. The Engineering Acceleration team builds products that multiply the effectiveness of OpenAI’s technical teams, helping engineers, researchers, and product teams understand complex systems, learn from what they ship, and operate reliably at scale. As AI changes how software is built, we have an opportunity to rethink engineering workflows from first principles. We’re creating tools and shared systems that turn complex data, experimentation, and technical workflows into clear decisions and useful action. About the Role In this role, you’ll lead design across two connected product areas: a real-time data exploration and observability experience for investigating large-scale system and product behavior, and an experimentation platform for safely launching changes, measuring their impact, and deciding whether to ramp, iterate, or roll back. This is more than a dashboard-design role. You’ll define the interaction models that take someone from a vague question or unexpected signal to a trustworthy answer and clear next step. You’ll work closely with engineers, researchers, data scientists, and product teams to understand the mechanics of their work and make dense technical systems coherent without flattening the details that matter. You’ll also help establish greater consistency across OpenAI’s enterprise and internal tools, developing durable patterns that support AI-native workflows and enable other designers to build more effectively. This role is based in our Seattle, WA or San Francisco, CA offices. We offer relocation assistance to new employees. In this role, you will: Lead end-to-end design for data-intensive products used by engineers, researchers, and product teams. Shape the complete learning loop: instrument, launch, observe, investigate, evaluate, decide, and iterate. Create clear, high-craft workflows for querying, filtering, comparison, drill-d

awsrestai
View job →
🔔

Get new mission systems engineers jobs by email

Daily job updates · Unsubscribe anytime