A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure, primarily across on-prem environments for the US Government. Forward Deployed Site Reliability Engineers combine engineering experience and an innate drive to improve existing systems and processes, with the creativity to develop novel solutions to evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. You’ll travel to various locations where you will be the expert for Palantir’s infrastructure, helping partner teams build & configure their hardware and network for software to operate reliably within. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate and participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.
Jobiba hiring network
Performance And Systems Engineer Jobs
6,482 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current performance and systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The Code Gen team is tasked with building AI-powered code transformation tools that transform rigid, legacy applications that suffer from poor scalability and high operating costs into modern, microservices-based architectures that are built on top of MongoDB. Join our team and be at the forefront of innovation and creativity. We are looking for a Staff Engineer with domain expertise and years of experience in modernizing legacy applications that are based on traditional database systems. A significant advantage is profound prior experience in leveraging AI, particularly LLMs and GenAI capabilities, to enable reliable, self-driving automation of the code transformation, iterative build, and test processes. In this role, you will be instrumental in initiating technical strategies and ideas, lead the Code Gen team in designing, building, and optimizing our code transformation workflow and tools. You will work on critical components that ensure the scalability, efficiency, and reliability of our services. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation, build and test. This role will be based remotely in North America. A strong candidate for this position will have Extensive experience (8+ years) in software development and operations, with a proven track record of delivering high performance, correctness, and architectural excellence in fast-paced environments Experience using Relational Databases such as Oracle, MySQL, Microsoft SQL Server or PostgreSQL Experience with tools and methodologies for code analysis, refactoring, and automated testing Experience in designing and implementing complex software systems, collaborating effectively with engineers of all experience levels to achieve high reliability and performance Practical knowledge of integrating GenAI into large-scale, complex systems, including a clear unde
The Code Gen team is tasked with building AI-powered code transformation tools that transform rigid, legacy applications that suffer from poor scalability and high operating costs into modern, microservices-based architectures that are built on top of MongoDB. Join our team and be at the forefront of innovation and creativity. We are looking for a Staff Engineer with domain expertise and years of experience in modernizing legacy applications that are based on traditional database systems. A significant advantage is profound prior experience in leveraging AI, particularly LLMs and GenAI capabilities, to enable reliable, self-driving automation of the code transformation, iterative build, and test processes. In this role, you will be instrumental in initiating technical strategies and ideas, lead the Code Gen team in designing, building, and optimizing our code transformation workflow and tools. You will work on critical components that ensure the scalability, efficiency, and reliability of our services. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation, build and test. This role will be based remotely in North America. A strong candidate for this position will have Extensive experience (8+ years) in software development and operations, with a proven track record of delivering high performance, correctness, and architectural excellence in fast-paced environments Experience using Relational Databases such as Oracle, MySQL, Microsoft SQL Server or PostgreSQL Experience with tools and methodologies for code analysis, refactoring, and automated testing Experience in designing and implementing complex software systems, collaborating effectively with engineers of all experience levels to achieve high reliability and performance Practical knowledge of integrating GenAI into large-scale, complex systems, including a clear unde
We're looking for a Senior Engineer with a strong background in computer science fundamentals, systems design, experience in the Java ecosystem, streaming systems, and data-intensive applications to join our engineering team. In this role, you will be instrumental in designing, building, and optimizing the underlying data structures, algorithms, and database interactions that power our generative AI platform, code generation and migration tools. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation and building a sophisticated data migration suite using a modern technology stack, which includes Java, Spring Boot, Kafka, Debezium, and React.You will work on critical components that ensure the scalability, efficiency, and reliability of our services, collaborating closely with AI researchers, product management and other engineers to design and implement cutting-edge products that solve complex customer challenges.. We are looking to speak to candidates who are based in Sydney for our hybrid working model. The ideal candidate for this role will have 6+ years of engineering experience in backend systems, distributed systems, or core platform development. Proficiency in one or several of Java, Rust, C/C++, and/or Python, with a strong understanding of systems-level programming, memory management, and performance tuning. Extensive experience with streaming data platforms such as Apache Kafka and Change Data Capture (CDC) tools like Debezium Extensive experience with relational data modeling and hands-on experience with at least one SQL database (Postgres, MySQL, etc) Exposure to client-side technologies such as JavaScript and React is a plus Good understanding of algorithms, data structures and their time and space complexity Curiosity, a positive attitude, and a drive to continue learning Excellent verbal and wri
MongoDB Atlas is the premier multi-cloud database-as-a-service built and operated by the makers of MongoDB. The Cloud Operations Engineering team at MongoDB is a worldwide team responsible for the consistent operational success of every MongoDB Atlas customer. As a Cloud Operations Engineer, you will help ensure the success of our Atlas customers, whether they are early startups or large multinational companies, cloud-native or just getting started with a digital transformation to the cloud. You are excited about the core mission of MongoDB, and the opportunity to join the team responsible for operating Atlas, the fastest-growing multi-cloud database-as-a-service in the world. You are prepared to be one of the early members of a 24/7/365 global cloud operations team. Cloud Operations Engineers will be responsible for day-to-day duties such as creating and monitoring system’s alert dashboards, reviewing critical events and system logs, accessing customer instances that underpin their production databases and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace. At MongoDB you will grow your career and skills, wear multiple hats, and be part of an operations team that works at the frontier of Cloud services and database systems. We are looking to speak to candidates who are interested in working out of our Dublin or Cork office under our in-office working model Monday to Friday. Due to the 24/7 nature of our support organization, certain events throughout the year will require volunteering for coverage outside one’s normal work days or work hours (i.e. regional offsites, regional holidays, etc). These are typically announced weeks in advance with a sign-up system that considers equitability. Responsibilities Successfully coordinate and collaborate with a global team of Cloud Operations Engineers wh
MongoDB Atlas is the premier multi-cloud database-as-a-service built and operated by the makers of MongoDB. The Cloud Operations Engineering team at MongoDB is a worldwide team responsible for the consistent operational success of every MongoDB Atlas customer. As a Cloud Operations Engineer, you will help ensure the success of our Atlas customers, whether they are early startups or large multinational companies, cloud-native or just getting started with a digital transformation to the cloud. You are excited about the core mission of MongoDB, and the opportunity to join the team responsible for operating Atlas, the fastest-growing multi-cloud database-as-a-service in the world. You are prepared to be one of the early members of a 24/7/365 global cloud operations team. Cloud Operations Engineers will be responsible for day-to-day duties such as creating and monitoring system’s alert dashboards, reviewing critical events and system logs, accessing customer instances that underpin their production databases and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace. At MongoDB you will grow your career and skills, wear multiple hats, and be part of an operations team that works at the frontier of Cloud services and database systems. We are looking to speak to candidates who are interested in working out of our Dublin or Cork office under our in-office working model Monday to Friday. Due to the 24/7 nature of our support organization, certain events throughout the year will require volunteering for coverage outside one’s normal work days or work hours (i.e. regional offsites, regional holidays, etc). These are typically announced weeks in advance with a sign-up system that considers equitability. Responsibilities Successfully coordinate and collaborate with a global team of Cloud Operations Engineers wh
MongoDB Atlas is the premier multi-cloud database-as-a-service built and operated by the makers of MongoDB. The Cloud Operations Engineering team at MongoDB is a worldwide team responsible for the consistent operational success of every MongoDB Atlas customer. As a Cloud Operations Engineer, you will help ensure the success of our Atlas customers, whether they are early startups or large multinational companies, cloud-native or just getting started with a digital transformation to the cloud. You are excited about the core mission of MongoDB, and the opportunity to join the team responsible for operating Atlas, the fastest-growing multi-cloud database-as-a-service in the world. You are prepared to be one of the early members of a 24/7/365 global cloud operations team. Cloud Operations Engineers will be responsible for day-to-day duties such as creating and monitoring system’s alert dashboards, reviewing critical events and system logs, accessing customer instances that underpin their production databases and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace. At MongoDB you will grow your career and skills, wear multiple hats, and be part of an operations team that works at the frontier of Cloud services and database systems. We are looking to speak to candidates who will be based remotely in Ireland. Due to the 24/7 nature of our support organization, certain events throughout the year will require volunteering for coverage outside one’s normal work days or work hours (i.e. regional offsites, regional holidays, etc). These are typically announced weeks in advance with a sign-up system that considers equitability. Responsibilities Successfully coordinate and collaborate with a global team of Cloud Operations Engineers who are tasked with ensuring our uptime guarantees to our Atlas customer base
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The Build Systems team within Figma’s Developer Experience organization owns Figma’s build and CI infrastructure, enabling engineers to ship changes to production quickly and safely. We build and operate core platforms across our polyglot monorepo, including build systems, artifact repositories, merge queues, test frameworks, and CI pipelines. We’re looking for an experienced technical leader to help shape these platforms, uplevel the team, and deliver high-impact platforms that accelerate engineering velocity. The ideal candidate has deep experience with large monolithic codebases, builds durable and scalable systems, and is motivated by solving high-leverage problems that amplify productivity across the engineering organization. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Drive technical roadmap and strategy for the Build Systems team Partner with cross-functional teams and leadership to identify developer pain points and design elegant, scalable solutions Lead complex, multi-quarter initiatives reducing build/test times and improving CI reliability, all while balancing technical excellence with pragmatic delivery Design, build, and maintain modern developer tools including scalable build systems, distributed CI platforms, and test frameworks that serve thousands of engineers Architect and implement large-scale infrastructure on AWS that powers our entire build pipeline to ensure reliability, performance, and cost
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! Application Platform is part of Product Platform within Figma’s Infrastructure organization. We build shared foundations that help engineers ship backend product work quickly, safely, and reliably. The team owns a large Ruby application that powers Figma’s REST APIs, asynchronous jobs, and workflow orchestration. Our systems sit on critical production paths and shape the day-to-day experience of backend engineers across Figma. We’re looking for an experienced backend engineer with meaningful production Ruby experience who enjoys building for other engineers. You don’t need to be a Ruby language specialist. You should be comfortable making informed tradeoffs in a substantial shared codebase and turning recurring problems into durable systems, tools, and paved paths that improve engineering velocity across the company. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you’ll do at Figma: Design and evolve shared Ruby frameworks for REST APIs, asynchronous jobs, workflow orchestration, and data access Improve the availability, reliability, scalability, and performance of backend systems that support critical product functionality Modernize a large Ruby codebase through safe architectural changes, including moving REST APIs from Sinatra toward Rails and introducing stronger endpoint abstractions and type safety Make backend development faster by improving hot reloading, local workflows, test infrastructure, CI reliability, debugging tools,
About the Role The Mechanical Commissioning Project Engineer owns mechanical commissioning planning, readiness, quality coordination, and test execution oversight for the project. This role ensures mechanical systems are installed, inspected, started, balanced, controlled, and tested in a way that supports reliable integrated facility performance. Reports to the Commissioning Project Lead and partners closely with mechanical contractors, equipment vendors, design/engineering teams, the electrical commissioning lead, controls stakeholders, and vendor field/test engineers. Key Responsibilities Develop and maintain the mechanical commissioning scope, readiness criteria, inspection strategy, and discipline test execution plan. Review mechanical design packages, specifications, submittals, method statements, controls narratives, sequence assumptions, and testing requirements for commissionability and risk. Coordinate mechanical QA/QC inspections with contractors and vendor field/test engineers, including installation checks, pre-functional readiness, deficiency capture, and closeout tracking. Own mechanical commissioning procedure development and review, including equipment startup, functional testing, controls verification, balancing prerequisites, failure mode validation, and integrated systems testing inputs. Coordinate with equipment vendors on factory/site acceptance requirements, startup support, test prerequisites, documentation packages, and vendor participation during critical tests. Support readiness and execution for cooling, ventilation, hydronic, pumping, heat rejection, controls, and other project-specific mechanical systems. Lead discipline-level review of mechanical test results, deficiencies, corrective actions, retest requirements, trend logs, and acceptance evidence. Maintain mechanical commissioning dashboards and status inputs for the Commissioning Project Lead, including risk items, resource needs, test readiness, and issue aging. Partner with the E
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
About the Team The Software Engineering team is responsible for designing and building the scalable, performant, and secure backend systems that power our products—from early prototypes to large-scale deployments. We collaborate closely with product, hardware, and full-stack teams to ensure our infrastructure enables fast iteration while setting a strong foundation for long-term growth. About the Role As a Backend Engineer , you will design and build services, APIs, and infrastructure that support evolving product needs. You’ll apply a deep understanding of backend systems and maintain enough end-to-end context—from hardware to cloud—to guide technical decisions that best serve the product and team. We’re looking for engineers who thrive in fast-paced, collaborative environments and care deeply about building robust systems that scale. This role is based in San Francisco, CA . We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Architect, build, and maintain high-performance, secure backend systems. Design APIs, data models, and infrastructure to support evolving product needs. Balance near-term development velocity with long-term maintainability and scalability. Collaborate with cross-functional teams to ensure cohesive, end-to-end solutions. You might thrive in this role if you: Have 7+ years of professional software engineering experience, with a focus on backend systems. Have a proven track record of building and scaling systems from early stage to large scale. Are proficient with Python and Go, and familiar with a range of server-side technologies. Have a strong grasp of system design, performance optimization, and security best practices. Can reason about full-stack tradeoffs from hardware through cloud infrastructure. (Nice to have) Have experience with distributed systems and cloud architectures. (Nice to have) Bring a background in instrumentation, analytics, and performanc
OpenAI’s charter calls on us to ensure the benefits of AI are distributed broadly and safely. Our Health AI team focuses on expanding access to high-quality medical expertise and aims to set a high standard for deploying AI responsibly in high-stakes domains. Improving health will be one of the defining impacts of AGI. Today, millions of people lack access to reliable medical information, and clinicians around the world face increasing time and resource constraints. We are building AI systems that support patients, clinicians, and health workers, while meeting the highest standards for safety, reliability, and privacy. We are seeking full stack software engineers to help build and scale products used by consumers and care providers globally. You will work closely with product, design, and research teams to ship real systems in a fast-moving, high-impact environment. In this role, you will: Design and build scalable fullstack systems for consumer and enterprise health. Own end-to-end feature development—from early design and implementation through deployment, monitoring, and iteration. Build and maintain data pipelines and services that meet strict privacy, security, and compliance requirements (e.g., HIPAA). Collaborate closely with researchers and safety teams to integrate reliability, evaluation, and guardrails into production systems. Debug, optimize, and harden systems to support high availability, performance, and global scale. Take ownership of ambiguous problems and drive them to practical, high-quality solutions. You might thrive in this role if you: Are deeply motivated by improving health outcomes and expanding access to medical expertise. Are a strong engineer who enjoys building durable, well-designed systems. Have 5+ years of experience writing maintainable, production-quality code. Can operate with high agency—owning problems end-to-end with minimal supervision. Enjoy working in fast-moving, cross-functional teams with engineers, product managers, desi
Principal Engineer - Backend About Us: Paytm is India’s leading digital payments and financial services company, which is focused on driving consumers and merchants to its platform by offering them a variety of payment use cases. To merchants, Paytm offers acquiring devices like Soundbox, EDC, QR and Payment Gateway where payment aggregation is done through PPI and also other banks’ financial instruments. To further enhance merchants’ business, Paytm offers merchants commerce services through advertising and Paytm Mini app store. Operating on this platform leverage, the company then offers credit services such as merchant loans, personal loans and BNPL, sourced by its financial partners. About the role: As a Principal Engineer, you will help define the technical design and implementation roadmap across multiple solutions and will work with engineering leadership to ensure we resource and equip our squads with the right expertise to deliver those solutions. Requirements: 8 to 12 years of strong software design/development experience in building massively large scale distributed internet systems and products Hands on experience in Advance Java, Spring boot, AWS, Node, Agentic AI, LLM, RAG, Cursor, Copilot Experience and knowledge of open source tools & frameworks, broader cutting edge technologies around server side development Should be an active contributor to developer communities like Stack overflow, Top coder, Git hub, Google Developer Groups (GDGs). Superior organization, communication, interpersonal and leadership skills. Must be a self-starter who can work well with minimal guidance and in fluid environment. Preferred Qualifications : Bachelor's/Master's Degree in Computer Science or equivalent Skills that will help you succeed in this role: Expertise in Java, DB: RDBMS, Messaging: Kafka/RabbitMQ, Caching: Redis/Aerospike, Micro services, AWS Strong experience in scaling, performance tuning & optimization at both API and storage layers Problem
Get new performance and systems engineer jobs by email
Daily job updates · Unsubscribe anytime