About the team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the role: We are seeking an experienced Optical Network Engineer to lead Laser related work within our optical interconnect efforts for large-scale compute systems. The role also requires broad, hands-on optical validation experience across IM/DD-based interconnects, working from lab characterization through production readiness and scaled deployment. In this role you will: Drive laser-focused requirements and technical direction within the broader optical interconnect roadmap. Lead evaluation and validation of optical components and subsystems, including laser-based elements, in lab and production-representative environments. Support end-to-end optical testing for IM/DD interconnects (e.g., module/system bring-up, characterization, debug, and readiness for scale). Work with external partners to align on development milestones, performance targets, and quality expectations. Own technical issue triage and resolution across performance, reliability, and manufacturability topics. Collaborate across internal teams to support integration, rollout, and operational success at scale. You might thrive in this role if you have: Strong experience in laser-focused optical engineering (development, validation, manufacturing readiness, or field support). Broad hands-on background with IM/DD optical technologies and optical test/debug workflows. Experience working with external suppliers/manufacturing partners and production-oriented execution. Demonstrated ability to debug complex t
Jobiba hiring network
Performance And Systems Engineer Jobs
6,482 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current performance and systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
We are seeking a Staff Engineer to join our growing team to provide technical direction and implement core parts of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Staff Engineer on this new team, you will be responsible for providing technical leadership to teams developing cutting edge technologies related to enabling deployment at scale of AI applications. You will take on challenging, high-visibility projects that improve and enhance the performance, scalability, and reliability of the distributed systems infrastructure for this new product. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day. We value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We're looking to speak with candidates based in the New York City area for our hybrid or in-office working models. Position Expectations Work closely with product management, product engineering, product design peers as well as other teams within the company to define the first version and future evolution of the service Design, build and deliver well-tested core pieces of the platform in collaboration with other vested parties Contribute to shaping architecture, code reviews and development practices, developer experience as the teams and product grow Mentor fellow engineers and assume ownership and accountability of projects Qualifications Strong background in building core components for high scale compute and data distributed systems 8+ years experience of building distributed systems, and/or foundational cloud services at scale and an interest in working with Python, Go and Java Proven success in designing, writing, testing, debugging, performance tuning, possessing a strong grip on the foundational materials of computer science and maintaining distributed and/or highly concurrent software s
About the Role At Jumio, the Software Development Engineer IV - QA (SDE-IV, QA) is a senior technical role focused on ensuring the quality, performance, and reliability of highly scalable web portals and distributed backend systems. Our platform spans multiple Java Spring Boot microservices and customer-facing web portals , deployed across AWS ECS, EKS, and Lambda , and integrated through event-driven messaging using SNS/SQS . In this role you will design and drive the test automation strategy for both UI (Playwright) and API/service layers, set the quality bar for the team, and act as a force multiplier — mentoring other engineers and embedding quality earlier in the development lifecycle. You will work closely with development, product, and DevOps teams to ensure our products meet the highest standards of quality, scalability, and security. This is a hands-on senior IC role: you will write code, but you will also influence architecture, own cross-service test strategy, and make build-vs-buy decisions for testing tooling. T-Shaped Engineering Expectation As part of Jumio's engineering culture, you will adopt a T-shaped engineering approach. Beyond deep expertise in test automation and quality engineering, you will contribute across the development lifecycle — understanding software architecture, participating in design and API-contract discussions, reviewing application code, and ensuring our distributed systems are testable, observable, and resilient by design. Role Value This role is critical to ensuring the reliability, scalability, and security of Jumio's products. By architecting and maintaining automated testing frameworks across web, API, and event-driven layers, you will enable faster, higher-confidence releases and reduce production risk in a complex microservices environment. What You'll Do Test Architecture & Strategy Define and own the end-to-end automated test strategy across web portals and backend microservices, balancing UI, API, contract, integ
At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Affirm’s Identity team is mission-critical to the customer checkout experience. When a customer chooses Affirm, one of the first steps is an identity check, and our ability to make the right decision directly impacts conversion, revenue, fraud exposure, and regulatory compliance. Our team owns Identity for all markets outside North America, playing a key role in Affirm’s international expansion. We are responsible for KYC, user lifecycle management, and identity decisioning across multiple regulatory environments. This work is about shaping how Affirm adapts its core Identity platform to new market dynamics, customer expectations, and compliance requirements. We are a team of six engineers based across Spain and Poland. We are looking for a Senior Software Engineer who can turn ambiguous business and customer problems into reliable, well-designed technical solutions. In this role, you will help shape our quarterly technical direction, translate team goals into concrete projects, and identify cross-cutting risks, tradeoffs, and opportunities. You should be comfortable designing across system boundaries, driving collaboration with partner teams, and advocating for technical investments that improve long-term execution. We’re looking for someone who takes ownership beyond shipping code: raising code quality and review standards, making performance, availability, and scale tradeoffs explicit, and improving the operational health of the systems they own. You’ll help reduce toil, create useful playbooks, mentor engineers, support new hires and interns, and contribute to high-signal hiring. Most importantly, we’re excited to work with someone who leaves things better than they found them, communicates clearly across audiences, uses customer feedback to influence technical plans, and help
At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Affirm’s Identity team is mission-critical to the customer checkout experience. When a customer chooses Affirm, one of the first steps is an identity check, and our ability to make the right decision directly impacts conversion, revenue, fraud exposure, and regulatory compliance. Our team owns Identity for all markets outside North America, playing a key role in Affirm’s international expansion. We are responsible for KYC, user lifecycle management, and identity decisioning across multiple regulatory environments. This work is about shaping how Affirm adapts its core Identity platform to new market dynamics, customer expectations, and compliance requirements. We are a team of six engineers based across Spain and Poland. We are looking for a Senior Software Engineer who can turn ambiguous business and customer problems into reliable, well-designed technical solutions. In this role, you will help shape our quarterly technical direction, translate team goals into concrete projects, and identify cross-cutting risks, tradeoffs, and opportunities. You should be comfortable designing across system boundaries, driving collaboration with partner teams, and advocating for technical investments that improve long-term execution. We’re looking for someone who takes ownership beyond shipping code: raising code quality and review standards, making performance, availability, and scale tradeoffs explicit, and improving the operational health of the systems they own. You’ll help reduce toil, create useful playbooks, mentor engineers, support new hires and interns, and contribute to high-signal hiring. Most importantly, we’re excited to work with someone who leaves things better than they found them, communicates clearly across audiences, uses customer feedback to influence technical plans, and help
About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform providing APIs for knowledge retrieval, inference, evaluation, and more. We are seeking a strong Senior Full-Stack Engineer to help us build, scale, and refine our rapidly growing product. The ideal candidate is deeply grounded in software engineering best practices and experienced in developing and scaling modern web applications end-to-end. You will work across the stack—from React/TypeScript frontends to Python-based backends—while integrating with LLMs and machine learning systems. You will solve complex challenges in scalability, reliability, and product experience while owning significant product areas in a fast-paced environment. What You’ll Do Own major full-stack product areas , driving features from design through production deployment. Build modern frontend experiences using React and TypeScript, ensuring performance, usability, and responsiveness. Develop reliable backend services in Python, working with distributed systems, data pipelines, and ML/LLM components. Integrate with LLMs, vector databases, and AI infrastructure to power intelligent product experiences. Deliver experiments and new features quickly , maintaining high quality and tight feedback loops with customers. Collaborate across product, ML, and infrastructure teams to shape the direction of Scale GP. Adapt quickly —learning new technologies, frameworks, and tools as needed across the stack. Ideal Experience 5+ years of full-time engineering experience , post-graduation. Strong experience developing full-stack applications using React, TypeScript, and Python . Experience scaling or shipping products at high-growth startups . Familiarity with LLMs, vector databases, embeddings, or other modern AI tooling (tinkering or production experience welcome). Proficiency with SQL and modern API development. Experience with Kubernetes , containerization, and microservice architectures. Experience working with at leas
About the role: We are seeking a Senior Backend Engineer with deep backend engineering expertise and proficiency in one or more major programming languages (e.g., Python, Java, Go, Rust, or Kotlin), along with a strong understanding of AI models and agents. As a core member of our AI Engineering team, you will collaborate with data scientists, ML engineers, and product managers to build scalable, production-ready infrastructure and APIs that power intelligent systems. What you'll be doing: As a Senior Backend Engineer in the AI Engineering team, you will: Build and maintain reliable, scalable backend services to support AI agent execution and orchestration. Develop AI agent systems for complex operational workflows using LangChain, LangGraph, LiteLLM, and Langfuse. Orchestrate a hybrid model stack that includes OpenAI and Google Gemini alongside self-hosted and fine-tuned LLMs like Gemma and Llama. Build and maintain integrations with clinical systems (FHIR, EMR). Drive observability and reliability using OpenTelemetry, Datadog, and Langfuse. Design APIs (GraphQL, REST), background workers, and event-driven systems that interface with AI inference engines and agent runtimes. Collaborate with Data Science, ML, and engineering teams to deploy AI features and improve the performance, scalability, and reliability of backend systems. Participate in code reviews, knowledge sharing, and mentoring to elevate the team’s technical capabilities. What we're looking for: 6+ years of backend engineering experience, with strong proficiency in more than one major programming language (such as Python, Java, Go, Rust, or Kotlin). Solid understanding of AI systems architecture and experience working in environments involving AI agents, LLMs, or inference pipelines. Proven experience in building and scaling backend APIs, microservices, and background jobs. Strong experience with relational and NoSQL databases (e.g., PostgreSQL, MySQL, MongoDB, Redis), including schema des
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration, where your code directly impacts the performance of state-of-the-art machine learning models. As a GPU Kernel Engineer, you'll craft the foundation that powers modern AI workloads, optimizing every microsecond of computation to enable breakthrough applications. You'll work in a fast-paced, intellectually stimulating environment where technical excellence is paramount and your contributions directly influence production systems serving millions of users across numerous products. This role offers exceptional growth potential for engineers passionate about low-level optimization and high-impact systems work. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Model Performance team: Baseten Embeddings Inference: The fastest embeddings solution available The Baseten Inference Stack Driving model performance optimization RESPONSIBILITIES Core Engineering Responsibilities Design and implement high-performance GPU kernels for key ML operations, including matrix multiplications, attention mechanisms, and mixture-of-experts routing Write and optimize code using CUDA, PTX assembly, and architecture-specific techniques Apply advanced performance optimization methods such as memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap Performance & Innovation Impl
As an Engineering Manager on Coder’s Agentic Engineering team, you’ll lead engineers building and evolving the systems behind our agentic development experience. You’ll help make agents more capable, reliable, and useful across real development environments. You’ll guide technical direction, grow the team, and keep execution sharp. You’ll work closely with Engineering, Product, and Design across the agent harness, integrations, and developer workflows. What you’ll do here Lead and grow a team within our Agentic Engineering organization. Set technical direction across the agent harness, integrations, and workflows. Stay close to the code and contribute to architecture and implementation decisions. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve reliability, performance, and operability across agentic systems. Coach engineers, raise the technical bar, and create clarity around priorities and tradeoffs. What we’re looking for Experience managing and growing software engineering teams. Strong hands-on engineering experience with Go. Experience with React and TypeScript. Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS. Strong technical judgment and comfort working through ambiguity. A track record of helping engineers grow while maintaining a high execution bar. Our tech stack Backend: Go, Postgres Frontend: TypeScript, React Infrastructure: AWS, Kubernetes Observability: Prometheus, Grafana CI/CD: GitHub Actions Bonus tacos if you have (Tacos? If you need an ice-breaker, ask how we say thanks by giving tacos!) Experience building coding agents, developer tools, or cloud development enviro
About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale with trillions of data points per day, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams for tens of thousands of companies globally. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team The Datadog Security Libraries team owns the customer-side integrations behind our run-time security products App & API Protection , Workload Protection , and Code Security . Our libraries let customers automatically manage application security risk with continuous, real-time monitoring of vulnerabilities and threats against their web applications, serverless applications, and APIs, in production. Automatically integrated with Application Performance Monitoring (APM) distributed tracing and code-level context, our software empowers development, operations, and security teams to build and run secure applications. As a polyglot team we ship and maintain the security capabilities of Datadog's tracing libraries across .NET , Java , Go , Node.js , Python , Ruby , and PHP , on top of a shared C++ core and a set of HTTP proxy integrations (primarily Envoy, NGINX, and HAProxy). Our code runs inside thousands of production applications around the world. Recent work spans exploit prevention (RASP) and WAF detections, API Security, code security (IAST and SCA), and AI-assisted ("agentic") onboarding, always measured by real product outcomes and operational telemetry. The Opportunity We're looking for a senior, polyglot engineer to contribute across several of our security libraries, with .NET or Java expertise. You'll design and build security integrations and detection features, take them from prototype to production-hardened, and own them operationally as they instrument thousands of applications. As a se
About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale with trillions of data points per day, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams for tens of thousands of companies globally. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team The Datadog Security Libraries team owns the customer-side integrations behind our run-time security products App & API Protection , Workload Protection , and Code Security . Our libraries let customers automatically manage application security risk with continuous, real-time monitoring of vulnerabilities and threats against their web applications, serverless applications, and APIs, in production. Automatically integrated with Application Performance Monitoring (APM) distributed tracing and code-level context, our software empowers development, operations, and security teams to build and run secure applications. As a polyglot team we ship and maintain the security capabilities of Datadog's tracing libraries across .NET , Java , Go , Node.js , Python , Ruby , and PHP , on top of a shared C++ core and a set of HTTP proxy integrations (primarily Envoy, NGINX, and HAProxy). Our code runs inside thousands of production applications around the world. Recent work spans exploit prevention (RASP) and WAF detections, API Security, code security (IAST and SCA), and AI-assisted ("agentic") onboarding, always measured by real product outcomes and operational telemetry. The Opportunity We're looking for a senior, polyglot engineer to contribute across several of our security libraries, with .NET or Java expertise. You'll design and build security integrations and detection features, take them from prototype to production-hardened, and own them operationally as they instrument thousands of applications. As a se
At Datadog, we're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. APM at Datadog is on its way to redefine how users interact with their telemetry. We are integrating intelligence directly into troubleshooting workflows to help engineers find root causes faster, navigate complex distributed systems seamlessly, and optimize application performance with minimal cognitive load. APM provides deep visibility from end-user interactions to backend services and we are now expanding this foundation with new AI-driven insights, guidance, and automation. As a Product Manager II for APM, you will work with world-class engineers, designers, and partner product teams to shape the future of Distributed Tracing, Performance Analysis, and Intelligent Troubleshooting. You will help build advanced capabilities that scale to thousands of customers and make sophisticated observability workflows accessible to every engineer, from experts to beginners. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What you will do: Develop a deep understanding of APM customers, their performance challenges, telemetry workflows, and competitors Lead conversations with design partners and strategic customers to uncover real-world performance issues, validate product assumptions, and guide solutions from early prototypes through General Availability Define and deliver the next generation of APM features with engineering and design, especially agentic on
At Datadog, we're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. APM at Datadog is on its way to redefine how users interact with their telemetry. We are integrating intelligence directly into troubleshooting workflows to help engineers find root causes faster, navigate complex distributed systems seamlessly, and optimize application performance with minimal cognitive load. APM provides deep visibility from end-user interactions to backend services and we are now expanding this foundation with new AI-driven insights, guidance, and automation. As a Product Manager II for APM, you will work with world-class engineers, designers, and partner product teams to shape the future of Distributed Tracing, Performance Analysis, and Intelligent Troubleshooting. You will help build advanced capabilities that scale to thousands of customers and make sophisticated observability workflows accessible to every engineer, from experts to beginners. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What you will Do: Develop a deep understanding of APM customers, their performance challenges, telemetry workflows, and competitors Lead conversations with design partners and strategic customers to uncover real-world performance issues, validate product assumptions, and guide solutions from early prototypes through General Availability Define and deliver the next generation of APM features with engineering and design, especially age
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Data team within Plaid’s Fraud organization builds the machine learning systems that power Plaid’s fraud detection products, leveraging Plaid’s unique network data to identify and stop fraud before it happens. The team owns the full ML lifecycle—from feature pipelines and model training to production serving and monitoring—building reliable, scalable systems that deliver high-quality fraud detection as we grow to support hundreds of customers. As a Senior Machine Learning Engineer, you will own the development of high-performance feature computation and online inference pipelines that power production machine learning systems at scale. You’ll build robust observability, monitoring, and automated debugging capabilities, while leveraging AI-assisted tools to investigate complex system behavior and maintain high reliability. You’ll partner closely with ML Infrastructure, Data Science, and Product teams to execute critical technical initiatives and deliver scalable, high-impact ML solutions. Responsibilities: Build and scale machine learning systems that power a rapidly growing fraud detection product in a fast-paced environment. Solve complex technical challenges at the intersect
Get new performance and systems engineer jobs by email
Daily job updates · Unsubscribe anytime