Jobiba hiring network

System Engineer Jobs

10,000 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current system engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab

kubernetesmachine learningai
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto

pythonci/cdgit
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg

pythonawsgcp
View job →
A
Anyscale
📍 Remote• Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: Optimizing performance of large-scale workloads on Ray Stability and stress testing infrastructure Improving fault tolerance (HA) As part of this role, you will: Leading cross-team projects while mentoring junior team members Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements

restmachine learningai
View job →
A
Anyscale
📍 Remote• Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: - Optimizing performance of large-scale workloads on Ray - Stability and stress testing infrastructure - Improving fault tolerance (HA) As part of this role, you will: Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements to Ray core Improve the testing process for Ray to make re

restmachine learningai
View job →

Senior Software Engineer - Analytics Compute Platform Team About the Role & Team Every chart, insight, and experiment result a customer sees in Amplitude passes through the compute layer that this team owns. Our team sits between Amplitude's front-end analytics products/ rest APIs/ MCPs and Nova, our proprietary analytics database, and owns the compute API that translates a user's question into a fast, correct answer. Our mandate: enable a highly performant, reliable, and flexible way to compute insights. Those three goals are often in tension the more flexible we make the system, the harder it is to keep it fast and simple, and a lot of the interesting engineering work on this team lives in that tradeoff. Day to day, you'll work closely with engineers on data management, marketing analytics, Session Replay, Experiment, CDP, Query, and various product teams, since they're all consumers of what we build. As a Senior Software Engineer, you will Design and build core parts of the query/compute engine that powers analysis across Amplitude's product suite Evolve the compute API and semantic layer that other engineering teams build features on top of, so that it can support a more generic table Help unify and evolve our core data model, including how non-event data (profile properties, lookup tables, etc.) is represented alongside event data Improve the performance, reliability, and scalability of query planning and execution on the Amplitude query engine Partner with teams across Session Replay, Experiment, CDP, Query, and frontend to understand their needs and shape the compute layer around them Take ownership of projects end-to-end, from design through rollout You'll be a great addition to the team if you have Strong backend or distributed-systems engineering experience, ideally touching OLAP databases, query engines, or data infrastructure Experience with SQL, query planning/execution, or building APIs that many other engineering teams depend on Comfort reasoning

sqlgitrest
View job →
A
Amplitude
📍 Remote• Full-time• From $1.6M/yr
1mo ago

Software Engineer II, Growth at Amplitude (View all jobs) San Francisco Bay Area About The Role & Team The Growth organization is focused on helping users realize long-term value from Amplitude. The team plays a critical role in that journey by driving expansion within existing accounts. Many of our customers start with just a few users — often in product or data teams. Our job is to unlock value for everyone else. We build features that help new users set up, get activated, leverage AI for frictionless insights, highlight relevant parts of the product they might otherwise miss, and expand usage across teams and departments. This is especially impactful in large, complex organizations, where getting from the first few users to broad adoption is a multiplier for retention and revenue. We work across Amplitude’s core product suite and integrations, partnering closely with Product, Design, and GTM. We ship fast, test often, and use data to guide what we build and where we invest. Our work touches activation, feature adoption, habit formation, and expansion — the entire lifecycle of turning someone from a new user into a loyal advocate. This is a role for a product-minded engineer who wants to own meaningful projects end-to-end and help shape the strategy and systems behind how customers adopt Amplitude. As a Software Engineer on Growth, you will: Design and build product features that drive activation, engagement, and cross-team expansion Work across the full stack using TypeScript, React, Node.js, and GraphQL Build experiences that improve user onboarding, feature discovery, and product education Create third party connections, allowing other products to integrate with Amplitude Collaborate with Product, Design, GTM, and other engineers to prioritize and deliver impact Use data to identify opportunities and measure outcomes Leverage AI coding tools to speed up development while staying grounded in system understanding Improve the quality, reliability, and maintain

typescriptreactnode.js
View job →

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Senior Software Engineers lead and mentor engineers, delivering high-value products for our customers and infrastructure that enables our business to scale. Vanta’s team and technology surface are growing quickly, and it’s essential that we invest in the right abstractions and systems to enable us to scale with our business. As a Senior Software Engineer, you’ll be responsible for setting technical direction to enable our product and infrastructure to scale with our business, driving complex projects across our technical stack, and mentoring our talented engineering team. Your past experience will be leveraged to enable and accelerate Vanta’s growth. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We’re growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and to contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! The Integrations Platform team mission is to power the world’s largest trust automation ecosystem, enabling any person or agent to build, connect, and automate trust seamlessly. We own Vanta’s integration ecosystem, which currently includes over 400 integrations across Cloud Providers (AWS, Azure, GCP), Identity Providers, Mobile Device Management (MDM), and Human Resources Information System (HRIS). We are focused on developing the Integration Platform. This includes creating shared primitives for authentication, lifecycle, observability, and publishing to ensure all integrations are built on the same fou

typescriptreactnode.js
View job →

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Software Engineers build and own high-value products for our customers and the infrastructure that lets our business scale. Vanta's team and technology surface are growing quickly, and it's essential that we invest in the right abstractions and systems to scale with our business. As a Software Engineer, you'll build and own full-stack features and platform primitives across our stack, work closely with the teams and partners who depend on what you ship, and grow into broader technical ownership. Your work will directly accelerate Vanta's growth. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We're growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and to contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! The Integrations Platform team mission is to power the world's largest trust automation ecosystem, enabling any person or agent to build, connect, and automate trust seamlessly. We own Vanta's integration ecosystem, which currently includes over 400 integrations across Cloud Providers (AWS, Azure, GCP), Identity Providers, Mobile Device Management (MDM), and Human Resources Information System (HRIS). We are focused on developing the Integration Platform. This includes creating shared primitives for authentication, lifecycle, observability, and publishing to ensure all integrations are built on the same foundation. Our North Star is to eliminate the barrier to building integrations entirely. We aim to enable any

typescriptreactnode.js
View job →
S
Supabase
📍 Remote• Full-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the role Auth , written in Go (server) and with client libraries for TypeScript , SSR and for other frameworks and technologies, is one of the most popular products in the Supabase stack. We are seeking someone to help us build new and maintain existing Auth features. What you’ll be responsible for Designing and implementing secure, scalable authentication features in Go and TypeScript. Working across the stack: from server-side protocols to client-side libraries for frameworks like Next.js. Owning the performance, reliability, and scalability of the Auth server across Supabase's infrastructure. Contributing to the evolution of our Auth architecture, including support for OAuth, OIDC, SAML, and other protocols. Planning and executing safe database migrations across a large fleet of Postgres instances. Building and improving observability: metrics, tracing, alerting, and dashboards to keep the system healthy at scale. Writing and reviewing RFCs as part of our product development process. Collaborating with engineers across Supabase to ensure a seamless experience for developers using our tools. Supporting the community and responding to developer feedback on GitHub, Discord, and other channels. You might be a good fit if you (Required) Have 4+ years of professional experience writing and shipping Go in production. (Required) Have 2+ years of professional experience working on an authentication system (implementing protocol support, maintenance at scale). (Required) Have strong relational database experience (Postgres or MySQL); Postgres experience is a bonus. Have strong knowledge of TypeScript in addition to Go (languages used daily). Have strong knowledge of web technology fundamentals (

javascripttypescriptjava
View job →
S
Supabase
📍 Remote• Full-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We're looking for engineers to join the team building Supabase Lite , a lightweight, TypeScript-native implementation of Supabase. It runs interchangeably on SQLite and Postgres and ships a PostgREST and Auth compatible API, so applications written against @supabase/supabase-js work as-is. It exists because AI builders and platform partners need a database they can give to every prototype without the cost or the wait: sub-second provisioning, a footprint small enough to run inside a sandbox, and an upgrade path to full Supabase when an app graduates to production. This is a role with a lot of agency. Small team, working product, patterns still to be set - you'll own large areas end to end and make real decisions from day one. What You'll Be Responsible For In this role, you'll: Grow API compatibility with the Supabase client, keeping behavior faithful to hosted Supabase and backed by conformance tests Bring more of the Supabase stack to SQLite Build and harden the path for upgrading a project from Supabase Lite to full Supabase Own the developer experience end to end: the CLI, local tooling, and docs Build and harden the hosted offering on a scale-to-zero, near-zero-cost footing, kept loosely coupled so the implementation can be swapped without a rewrite Work directly with design partners and turn real-world migration and integration friction into a sharper product Engage with the open-source community and contribute back to the broader Supabase stack. You Might Be a Good Fit If You Have substantial backend or full-stack experience with strong fluency in TypeScript, and have built or contributed to a backend system, framework, or developer platform before Are a strong generalist,

typescriptsqlrest
View job →
S
Supabase
📍 Remote• Full-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We're looking for a SDK Engineer - JavaScript to join the SDK team and own a core part of how developers talk to Supabase from JavaScript and TypeScript. Our JS/TS SDKs are the front door to the platform: database queries, auth, storage, realtime subscriptions and edge functions, used by a very large number of developers and by most of the AI coding tools building on Supabase today. This is a library-authoring role, and the work happens in the open: public repos, public issue trackers, and API decisions whose consequences are permanent. The type system here is a design surface rather than a formality. If you get satisfaction from an inference that Just Works, from a migration guide that saves people an afternoon, and from an issue tracker that isn't a graveyard, you'll like it here. This role is ideal for someone who thrives in async, fast-paced environments and is excited about building developer tools that scale to millions. What You'll Be Responsible For Build and evolve our JavaScript/TypeScript SDKs. Start new things. Alongside the maintenance work, you're expected to find what's missing, make the case for it, and build it. Some of what we ship next doesn't exist yet, and nobody will hand you the list. Work in the open. These are public repos: you'll triage inbound issues, review and shepherd outside contributions, keep CI and release automation healthy, and hold a clear, kind line on scope. Own the type experience. Keep generated database types flowing correctly through the client APIs so autocomplete and inference stay correct in real codebases. Keep us honest across runtimes. Node, Deno, Bun, browsers, Cloudflare Workers, React Native/Expo, including dual ESM/CJS publishi

javascripttypescriptjava
View job →
S
Supabase
📍 Remote• Full-time
1mo ago

Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We’re looking for a Senior Postgres Engineer to join our Postgres Team and help maintain and expand the stability and functionality of our hosted Postgres offering . You’ll work closely with customers, partners, product, engineering, support and success, helping us maintain a secure, stable, performant, and functional Postgres foundation. This role is ideal for someone who thrives in async, fast-paced environments and is excited about building developer tools that scale to millions. What you'll own: Build and maintain PostgreSQL extensions in C and Rust, with a deep understanding of internals — parser, planner, WAL mechanics, and MVCC Diagnose and resolve issues in managed PostgreSQL deployments, including custom extension failures, core dump analysis, and performance bottlenecks Own idempotent deployment pipelines across thousands of running PostgreSQL instances, including testing and rollout strategies Manage complex extension ecosystems — compatibility, upgrade paths, and conflict resolution Work with PostgreSQL's background worker framework, shared memory management, and hook system to build reliable, scalable functionality Collaborate closely with customers, partners, product, engineering, support, and success teams to maintain a secure, stable, and performant Postgres foundation What you bring: Deep expertise in PostgreSQL internals - query planner, executor, and storage engine mechanics Proven experience building PostgreSQL extensions in both C and Rust Strong knowledge of PostgreSQL's permission model : RLS, roles, and grant systems Experience troubleshooting production issues in managed PostgreSQL environments, including custom extension issues, performance bottlenecks, and resource co

sqlpostgresqlai
View job →
S
Sentry
📍 San Francisco• Full-time• $220K – $450K/yr
1mo ago

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role AI and machine learning are reshaping how developers debug, monitor, and ship software, and Sentry is uniquely positioned to lead that shift. We sit on a novel and massive dataset of real production errors, spans, and logs from tens of thousands of engineering organizations — the kind of signal that makes ML genuinely useful, whether it's a clustering model that groups related issues, a ranking system that surfaces the right alert at the right time, or an agent that proposes a fix. We're looking for an Engineering Manager to lead and grow our Machine Learning Engineering team. This team owns the full spectrum of ML at Sentry: classical techniques like clustering, ranking, anomaly detection, and embeddings that quietly power core product surfaces today, alongside the LLM-based and agentic systems shaping where the product is headed. You'll partner closely with product, design, and engineering leaders to decide where ML belongs in our products, what kind of ML actually fits the problem, and how we translate that work into experiences millions of developers rely on every day. In this role you will Set technical direction across the team's full ML surface area — from classical models for clustering, ranking, and anomaly detection to LLM-based and agentic systems — and make sharp calls about which approach fits each problem Define how the team evaluates and monitors ML systems in production, from offline metrics to online experimentation to model and agent observability Stay hands-on enough to review code and model designs, contribute to architecture discussions, and unblock engineers on complex ML problems Define

restmachine learningai
View job →
R
Ramp
📍 New York• Full-time• From $10K/yr
1mo ago

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About Production Engineering Production Engineering is Ramp's infrastructure ownership layer. We exist to make Ramp faster, more reliable, and more scalable — and we do that by being embedded in the problems, not adjacent to them. A few things that define how we operate: One team, one company, one objective. There is no "infra team" and "product team" — there is Ramp. We share the company's goals as our own. When a product team struggles with reliability or scalability, that is our struggle. If reliability or scalability is at risk, we own it. We don't wait to be invited, and we don't ask whose code it is. If a system is slow, if it breaks, if it won't scale — that's ours to lead, regardless of where it lives in the stack. We go first, and we go fast. When the path isn't obvious, we don't wait for someone else to find it. We move with urgency, propose the solution, align the stakeholders, and stay in until it's done — not until our ticket is closed. We lead the way. We find the next problem before it finds us. And when we solve it, we don't just fix

awsci/cdrest
View job →
🔔

Get new system engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More system engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring System Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.