Jobiba hiring network

Ai Systems Engineer Jobs

10,000 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current ai systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. Database Observability is a critical pillar of New Relic's platform strategy. We are looking for an experienced Engineering Manager (M3) to lead a senior, high-performing team building next-generation Database Observability products. Your team will own the entire lifecycle of critical telemetry data flows from lightweight database agents and high-throughput ingestion pipelines to intelligent DB recommendation engines and autonomous DB AI Agents. You will lead a team that includes senior and Lead-level engineers with deep domain expertise in distributed systems and AI. Your primary value will come from setting strategic technical direction, enabling their best work, and fostering a high-accountability culture while partnering closely with Product and Design to deliver features that directly drive New Relic's Database Observability. What you'll do Manage a full-stack engineering team (6–8 engineers) spanning backend systems, database telemetry, agent engineering, and UI workflows. Own end-to-end delivery sprint planning, roadmap execution, system quality, and operational excellence for critical database ingesti

pythonjavasql
View job →
I
Instacart
📍 BC, Canada• Full-time• Remote• From C$196K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacart is the North American leader in online grocery, and we’re building the operating system for the grocery industry so people can access the food they love and more time to enjoy it together. We’re looking for an Engineering Manager to lead our Catalog Enrichment team. This team builds the AI-native platform and pipelines that create, enrich, and maintain product attributes across a catalog of tens of millions of products from more than 100,000 retailer locations. Catalog data is the foundation of our marketplace: when attributes are complete and accurate, search returns the right results, recommendations feel personal, and advertisers can target with confidence. In this role, you’ll lead a team of engineers building systems that blend large language models, classical inference, workflow orchestration, and human review into configurable pipeli

REMOTEaigo
View job →
A
Amplitude
📍 San Francisco• Full-time• $165K – $247K/yr
18 days ago

Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. About the Role Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Senior Platform Engineer, you'll own medium-to-high-complexity platform projects end-to-end and help shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll partner with Staff engineers and product teams to make Kubernetes effortless across the engineering org, building self-service automation and scalable AWS infrastructure that lets product teams ship faster, safer, and with less cognitive load. If you're excited about building the systems that other engineers will rely on every day, this role is for you. Key Resp

pythonawsgcp
View job →
DU
DoorDash USA
📍 San Francisco• Full-time• From $1.6M/yr
19 days ago

About the Team The Storage teams build and operate online stateful systems and abstractions that are reliable, efficient, secure and easy to use for DoorDash Engineering. The teams are responsible for understanding Product Engineering’s evolving needs and developing platform and infrastructure capabilities to serve them. The team currently supports CockroachDB, Cassandra, Kafka and Redis as well as data abstraction services to reduce the complexity of interacting with storage systems for Product Engineers. About the Role The Storage team is building and operating a high-performance, scalable, and reliable data abstraction layer that optimizes both efficiency and reliability. Our goal is to create a platform that manages itself and fades into the background—empowering engineers to focus on delivering product experiences our customers love. This role is available across two teams within Storage, each solving unique and high-impact challenges: One team is building the orchestration layer for DoorDash’s storage platform—unifying lifecycle management, operations, and self-serve APIs for databases and streaming systems, turning complex, stateful infrastructure into reliable, developer-friendly services used across the company. One team builds and operates the distributed data platform powering DoorDash's largest stateful workloads -- including Cassandra, which backs critical product surfaces across DoorDash, Wolt, and Roo. You'll design high-throughput data abstractions, smart clients, and platform services that make distributed data reliable and easy to work with at multi-petabyte, multi-million-QPS scale, with opportunities to go deep on distributed systems internals and contribute to the open-source Cassandra ecosystem. If you're passionate about distributed systems, developer experience, and building foundational infrastructure at scale, we'd love to hear from you. You must be located in San Francisco, Sunnyvale, Seattle, or the New York Metro Area for this hybrid pos

javasqlredis
View job →
A
Amplitude
📍 Remote• Full-time• $165K – $247K/yr
1mo ago

About the Role Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Senior Platform Engineer, you'll own medium-to-high-complexity platform projects end-to-end and help shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll partner with Staff engineers and product teams to make Kubernetes effortless across the engineering org, building self-service automation and scalable AWS infrastructure that lets product teams ship faster, safer, and with less cognitive load. If you're excited about building the systems that other engineers will rely on every day, this role is for you. Key Responsibilities Lead high-impact platform projects — design and ship capabilities that move the needle on developer experience, reliability, or security, and set the bar for quality, testing, and safe deployment practices. Build the AI-augmented platform. Design tooling and workflows that help engineers get more out of AI-assisted development — think infra primitives that are easy to reason about, automated review, and policy-as-code that keeps the guardrails strong as AI shifts how code gets written. Own Infrastructure-as-Code for Kubernetes, AWS, and GCP using Terraform, Helm, Kustomize, and emerging tooling — and make it consumable enough that an LLM can safely PR against it. Evolve our CI/CD backbone (Argo CD / Workflows / Rollouts, GitHub Actions) to make deploys faster, safer, and easier to reason about. Instrument and operate. Drive observability with Datadog and Amplitude, own dashboards and SLOs, and use the data to push reliability forward. Participate in on-call, lead incident response when needed, and turn postmortems into durable platform improvements. Reduce toil and tech debt with pragmatic remediation

pythonawsgcp
View job →
G
19 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe

pythonaigo
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
21 days ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a product manufacturing & quality engineer, who will be responsible for driving technical initiatives related to the manufacturing, quality and reliability of our AI supercomputer hardware systems to ensure product success from concept to launch and through mass production. You’ll have the opportunity to coordinate with functional SMEs and work with a wide range of stakeholders, from design engineering and operations teams, TPMs, external industry vendors and partners to ensure that all products are developed and delivered on time and to the highest quality standards. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees In this role, you will: Own the integrated manufacturing and quality readiness for a product across L6, L10, and L11, with clear gates, milestones, deliverables, owners, and closure criteria. Lead readiness of process flows, tooling, fixtures, assembly operations, test interfaces, and production controls. Review and contribute to work instructions. Translate product requirements into qualification plans, process controls, test requirements and acceptance criteria with design engineering and Area SMEs Coordinate and drive execution of product and process qualification, reliability testing, and validation with the relevant SMEs. Maintain traceable evidence that assigned products and processes meet agreed performance, reliability,

REMOTEawsrestai
View job →
A
27 days ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join The Quality Platform team is at the heart of Airbnb’s mission to deliver a seamless, high-quality experience for millions of hosts and guests. We don’t just find bugs — we build the systems that prevent them. Our team sits at the intersection of Quality Engineering, Infrastructure, and Applied AI. We are evolving how software quality is built by integrating LLMs, intelligent automation, and data-driven systems into the testing lifecycle. You will join a high-impact group of engineers focused on building AI-powered quality systems that scale across one of the world’s most complex codebases. The Difference You Will Make: As a Mobile Software Engineer, you will be a key contributor to the development of our Quality Platform across both native mobile stacks. You will help build the foundation for how quality is engineered across Airbnb's iOS and Android ecosystems, developing tools and frameworks that enable our mobile platform to scale while keeping developers productive and confident, regardless of which platform they build on. In this role, you will: Build AI-Driven Solutions: Contribute to AI-native agents that automate repetitive testing tasks and provide intelligent feedback to developers on both platforms.Deliver Scalable Infrastructure: Develop and maintain the high-scale platforms and testing environments used daily by the iOS and Android engineering organizations.Promote Engineering Craft: Implement best-in-class mobile patterns and modularity to improve testability and fault-tolerance across both native codebases.Contribute to Operational Excellence: Ensure our automated syste

ci/cdaikotlin
View job →

The Development Infrastructure team builds the tooling and systems our Asana engineers use every day to bring their ideas to production quickly and reliably. We build and operate the software that drives Asana’s roadmap. Each day, we combine industry best practices and innovation to support this product-focused company. We’re looking for an experienced Software Engineer with a passion for developer infrastructure. You will work with a world-class team of engineers on deploying and operating existing developer tooling, and building new tools to support our global, growing development team. You will have a unique opportunity to design and develop the systems and applications that drive the Asana development experience, lead complex technical projects, and work on cross-functional initiatives to help define the future of software engineering at Asana. This role is based in our Reykjavík office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Innovate the architecture of our development sandboxes to accelerate development iteration. Modernize our build system, tighten iteration loops, and polish frequent cycles in the developer experience. Lead complex developer infrastructure projects from technical design through implementation, rollout, and operation. Partner with engineering teams to identify opportunities to enable teams to develop faster at Asana and help the company achieve our goals faster. Analyze complex developer infrastructure systems to uncover issues, root causes, and areas for improvement. Keep Asana up to date on open-source trends and identify new opportunities for improved development infrastructure. Champion co

awsrestai
View job →
L
Lyft
📍 Toronto• Full-time• From C$172K/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft is looking for a Design Systems Manager to join Design Foundations, the infrastructure layer of Lyft's design organization. Design Systems at Lyft is an integral part of the overall user experience. Our design system spans product surfaces, marketing, and brand, officially known as the Lyft Product Language (LPL) and shipped directly to code. You'll own the craft and execution of that system, partner closely with design, engineering, and marketing, and lead Design systems work through a period of active global growth and capability expansion. This role reports to the Director of Design Foundations and works alongside a dedicated DPM and Engineering managers, and a tight-knit team of design systems practitioners. Responsibilities: Strategy & Vision Set the strategy and roadmap for Lyft Product Language (LPL) across product, marketing, and brand surfaces, in partnership with the Director of Design Foundations and engineering leadership and the Design Systems team Lead the AI transformation of the system: integrate AI into LPL tooling, adoption workflows, and the day-to-day designer experience at Lyft Represent Design Systems in reviews and other senior leadership forums People & Craft Manage and develop a team of design systems designers and illustrators working across iOS, Android, web, and brand Hold the quality bar for system-wide design: cross-platform coherence, accessibility (WCAG), road safety, localization, and brand consistency. Coach the team on making educated tradeoffs between flexibility, coherency, usability, and efficiency Execution & Partnership Partner closely with the Design Systems DPM and engineering managers to maintain system health, sequence work, and unblock teams Own the LPL component contribution and review process end-to-end, from intake through release Commu

L
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft is looking for a Design Systems Manager to join Design Foundations, the infrastructure layer of Lyft's design organization. Design Systems at Lyft is an integral part of the overall user experience. Our design system spans product surfaces, marketing, and brand, officially known as the Lyft Product Language (LPL) and shipped directly to code. You'll own the craft and execution of that system, partner closely with design, engineering, and marketing, and lead Design systems work through a period of active global growth and capability expansion. This role reports to the Director of Design Foundations and works alongside a dedicated DPM and Engineering managers, and a tight-knit team of design systems practitioners. Responsibilities: Strategy & Vision Set the strategy and roadmap for Lyft Product Language (LPL) across product, marketing, and brand surfaces, in partnership with the Director of Design Foundations and engineering leadership and the Design Systems team Lead the AI transformation of the system: integrate AI into LPL tooling, adoption workflows, and the day-to-day designer experience at Lyft Represent Design Systems in reviews and other senior leadership forums People & Craft Manage and develop a team of design systems designers and illustrators working across iOS, Android, web, and brand Hold the quality bar for system-wide design: cross-platform coherence, accessibility (WCAG), road safety, localization, and brand consistency. Coach the team on making educated tradeoffs between flexibility, coherency, usability, and efficiency Execution & Partnership Partner closely with the Design Systems DPM and engineering managers to maintain system health, sequence work, and unblock teams Own the LPL component contribution and review process end-to-end, from intake through release Commu

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. Role Overview We are seeking a Package Reliability Engineer to lead reliability engineering for advanced packages used in high-performance AI and computing systems. The primary focus of this role is to assess package level mechanical and thermal reliability risks and apply thermal and mechanical modeling to optimize package design, material selection, and assembly processes. The engineer will also develop reliability test plans with external partners, identify failure mechanisms, perform root-cause analysis, and recommend practical corrective actions. In this role, you will assess package reliability risks from early architecture development through product qualification and high-volume manufacturing. You will work closely with package design, silicon design, system engineering, manufacturing, and ASIC partners to predict package behavior, develop qualification strategies, resolve reliability issues, and improve overall package robustness and lifetime. In this role you will: Lead reliability test plan and assessments for advanced HPC packages, including risk identification, potential failure-mechanism analysis, root-cause investigation, mitigation planning, and corrective-action development. Drive reliability-focused package design optimization based on thermo-mechanical modeling to improve package reliability, power integrity, thermal performance, mechanical robustness, and platform scalability. Develop, validate, and apply package reliability models and lifetime-prediction

redisawsrest
View job →

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Infrastructure Engineer Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning

pythonsqlaws
View job →

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. Today, we are tapping into the unlimited potential of AI to define the next era of computing. As an NVIDIAN, you will address challenges spanning architecture, silicon, firmware, software, and production — and excellent judgment matters as much as technical depth! Every major NVIDIA silicon product family—from the chips powering AI and datacentre systems to gaming, professional, embedded, and automotive platforms—passes through our productization work on its way to production. NVIDIA’s Silicon Co-Design Productization team works from pre-silicon strategy and feature development through bring-up, characterization, correlation, and optimization. Our charter spans power & performance modelling, bring up & tuning of low-power features, power & thermal controllers , and system-level optimization that ultimately shape how NVIDIA products are configured, binned, specified, and shipped. What you'll be doing Drive silicon power productization from pre-silicon planning through bring-up and production, including test strategy, feature readiness, characterization, and optimization. Partner with architecture and design teams to identify improvements, validate features, and help translate them into production-ready solutions. Correlate measured silicon behaviour with pre-silicon expectations, investigate gaps, and drive complex issues to root cause. Build power and performance models and characterization methodologies that decide silicon binning, product specifications, productization decisions, and customer guidance. Use AI/ML and data-driven methods to analyse characterization & telemetry data, identify anomalies & trends , and accelerate issue debug across silicon, board, power delivery, firmware,

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Site Reliability Engineer (SRE) - Security and Data Systems Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering and systems administration, where you'll be responsible for building and maintaining highly reliable, scalable, and secure infrastructure. You will be a key contributor, applying your expertise to automate manual processes and proactively solve complex problems before they become incidents, handling incidents, and includes on-call shifts. Responsibilities Platform & Reliability: Design, build, and maintain the core infrastructure that underpins our security SaaS offerings, ensuring high availability, performance, and scalability. This includes building and operating the tooling for our Snowflake data systems. Automation: Develop robust automation using code to eliminate toil and ensure consistency across our environments. You'll be a key driver in automating everything from infrastructure provisioning to application deployment and incident response. Security & Compliance: Work closely with our security teams to embed a security-first mindset into all our processes and infrastructure. You will be responsible for ensuring our systems and data platforms are compliant with industry standards. Incident Response: Participate in on-call rotations and be a primary responder for critical inci

kubernetesmachine learningartificial intelligence
View job →
🔔

Get new ai systems engineer jobs by email

Daily job updates · Unsubscribe anytime