Jobiba hiring network

Performance And Systems Engineer Jobs

6,482 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current performance and systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Firmware Test Engineer at Micron Technology, Inc., you will build groundbreaking high-performance controller firmware test framework for volatile and non-volatile memory systems. You will assist in the evaluation, creation, build, bench testing, debugging, and failure analyzes of firmware for new high-performance memory controllers and Solid State Drives (SSD) that will improve performance, while reducing power, latency and SoC (System on Chip) complexity for the target sectors. You can expect to partner multi-disciplinary Engineers seek multi-functional product development issues. You will triage failures, file bug reports, and help the development teams with isolating issues. Experience / Skills: 8 to 12 years of experience in managing the test development team within the storage domain. In depth knowledge and extensive experience with embedded firmware development Expertise in the use of scripting languages, programming tools and environments Extensive experience programming in Python Technical Expertise in the storage industry in SSD, HDD, storage systems, or a related technology Understanding of storage interfaces including ideally PCIe/NVMe, SATA, or SAS Experience with NAND flash and other non-volatile storage Ability to work independently with a minimum of day-to-day supervision Experience with team leadership and/or supervising junior engineers and technicians Ability to work in a multi-functional team and under

pythonartificial intelligenceai
View job →

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Staff DB SRE to build the runtime foundation for NVIDIA’s enterprise AI platforms — with a strong emphasis on database infrastructure at scale. This role blends large-scale database transformation with the building and development of GPU-accelerated platforms. You'll develop the software systems, automation frameworks, and high-performance database services that power NVIDIA’s AI workloads at scale. What you'll be doing: Design and operate highly available database clusters (MySQL, MSSQL, Oracle) with automated replication, failover, point-in-time recovery, and disaster-recovery strategies at enterprise scale. Drive database performance engineering — own query optimization, indexing strategies, connection pooling, lock-contention analysis, and storage-engine tuning for production systems handling millions of transactions. Build self-service database lifecycle automation — from one-click cluster provisioning and schema migrations to zero-downtime upgrades, blue-green deployments, and automated capacity scaling. Bridge relational and AI-native data infrastructure — extend traditional database exper

pythonmysqlkubernetes
View job →
B
Biohub
📍 Redwood City• Full-time• Hybrid• $241K – $331K/yr
1mo ago

Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery, integrating frontier AI models, biological foundation models, and lab capabilities, with the ultimate goal of curing disease. Our technology powers scientists around the world, translating AI capabilities into tools that accelerate research everywhere. The Team The AI Cluster Production Engineering team is part of the AI Compute Platform organization at Biohub, a non-profit research lab committed to open science and open-source AI. We own the design, operation, and reliability of large-scale multi-GPU AI clusters that power frontier AI biology research: protein language models, genomic foundation models, and scientific reasoning systems built to be shared, not monetized. Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the intersection of AI tooling, distributed systems, HPC, and frontier AI, debugging deep AI infrastructure problems and building AI systems critical to the entire AI organization. The Opportunity CZ Biohub's mission is to cure or prevent all human disease. Achieving that requires training frontier-scale AI biology models, and that demands reliable, high-performance compute infrastructure. This is production engineering work at a frontier AI lab, with the twist that the mission is biology and the science is open. You'll keep GPU clusters running at high utilization, debug the toughest distributed systems failures, and build the operational foundations for scaling to multi-thousand GPU hero runs. The technical problems are genuinely hard (e.g., multi-node distributed training, InfiniBand fabrics, large-scale storage, Slurm at scale) inside an organization where the work is aimed at helping peop

pythonkubernetesgit
View job →

What you’ll do Partner with medical image reconstruction scientists / engineers to build ML components that improve reconstruction quality, speed, robustness, or quantitative accuracy. Define training/evaluation pipelines, datasets, and metrics that map to user needs and design requirements. Productionize models: inference performance, reproducibility, monitoring for drift/regressions, and safe fallbacks. Collaborate on hybrid algorithms, incorporating physics and learned priors, denoisers, learned regularizers, and quality estimation. Help build tooling for rapid experimentation as well as rigorous verification of algorithm changes. What we’re looking for Strong applied ML experience plus comfort with signal processing / imaging or adjacent domains. Ability to move fluidly between research prototypes and production-quality systems. Strong evaluation discipline: metrics, ablations, data leakage avoidance, and reproducibility. A demonstrated track record of applying ML to physics-based or inverse problems (i.e., shipped projects, a portfolio, or publications.) Useful experience ML for imaging/inverse problems (or adjacent) with strong evaluation discipline and comfort with GPU performance constraints. Pragmatic production mindset: reproducible training/inference, regression testing, and safe deployment in high-stakes contexts. A background in computational physics or scientific computing. Leverage ML-based methods such as PiNNs and Neural Operators to solve partial differential equations arising in ultrasound simulation and imaging. Experience in Agentic-SciML is a plus. Hands-on experience with data curation for ML: building datasets from messy, real-world sources, defining ground truth, and managing labeling or simulation pipelines. Background in data assimilation: combining observations with physics-based models (Kalman filtering, variational methods, ensemble approaches, or learned variants).

R
Roblox
📍 San Mateo• Full-time• From $295.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox Ads business growing at a rapid rate, we are building large scale ads machine learning infrastructure to deliver effective performance ads to our users, and more business values to our advertisers. We’re looking for an EM to lead a team of exceptional ML infrastructure engineers, build scalable, reliable, and high-performance infrastructure that powers ML systems across our organization. You’ll operate at the scales of hundreds of billions of engagements, and redefine how we deliver performance ads to hundreds of millions of users. You Will: Lead strategic planning and roadmap execution of scalable production-ready ML systems including model training, data pipelines, feature engineering and model inference. Own the architecture, establish engineering best practices of scalability, reliability, and cost-effectiveness of ML infrastructure (e.g., training, serving, feature). Work closely with data scientists, ML engineers, platform teams, and product stakeholders to design, implement, and operate robust ML platforms that accelerate model development and deployment. Recruit, mentor, and grow a high-performing team of ML infrastructure engineers. You Have: 5+ years of experienc

awsgitmachine learning
View job →
M
Mongodb
📍 United States• Full-time• From $151K/yr
1mo ago

We’re looking for a Senior Engineering Manager who is ready to lead through ambiguity and improve how software gets built at MongoDB. This role leads teams focused on developer productivity, with an emphasis on measurable improvements to the software development lifecycle. This role can be based remotely in the United States. The Team The AXIS team (AI, X-functional tools, Insights, and Signals) sits within Developer Productivity and is responsible for overseeing the metrics and observability infrastructure of our expansive developer environment to help build a strong data-driven culture. You’ll also be a key partner in building the agentic ecosystem for AI-driven development across engineering. Candidate Profile We’re looking for an experienced leader with a passion for solving the big challenge of measuring developer productivity and providing the actionable signals that help teams improve their performance. They should be comfortable working collaboratively with other leaders and partners across our Engineering and Data teams in maximizing the use of data for insights and AI enablement. The right candidate for this role will have 4+ years of experience managing software engineers, including hiring, performance management, growth planning, and compensation; required for external candidates and preferred for internal candidates 8+ years of hands-on software engineering experience building and operating production systems; experience in developer tooling, platform engineering, observability, or data engineering is a strong plus Demonstrated the ability to lead through ambiguity, work across team boundaries, and deliver outcomes without close supervision Strong customer orientation and sound judgment in finding practical, high-leverage solutions Experience working with systems involving analytics, data pipelines, and metrics platforms Experience with AI tools development and enablement efforts Strong technical judgment, including the ability to evaluate t

mongodbawsazure
View job →
M
Mongodb
📍 Alberta• Full-time• From C$191K/yr
1mo ago

We’re looking for a Senior Engineering Manager who is ready to lead through ambiguity and improve how software gets built at MongoDB. This role leads teams focused on developer productivity, with an emphasis on measurable improvements to the software development lifecycle. This role can be based remotely in Canada. The Team The AXIS team (AI, X-functional tools, Insights, and Signals) sits within Developer Productivity and is responsible for overseeing the metrics and observability infrastructure of our expansive developer environment to help build a strong data-driven culture. You’ll also be a key partner in building the agentic ecosystem for AI-driven development across engineering. Candidate Profile We’re looking for an experienced leader with a passion for solving the big challenge of measuring developer productivity and providing the actionable signals that help teams improve their performance. They should be comfortable working collaboratively with other leaders and partners across our Engineering and Data teams in maximizing the use of data for insights and AI enablement. The right candidate for this role will have 4+ years of experience managing software engineers, including hiring, performance management, growth planning, and compensation; required for external candidates and preferred for internal candidates 8+ years of hands-on software engineering experience building and operating production systems; experience in developer tooling, platform engineering, observability, or data engineering is a strong plus Demonstrated the ability to lead through ambiguity, work across team boundaries, and deliver outcomes without close supervision Strong customer orientation and sound judgment in finding practical, high-leverage solutions Experience working with systems involving analytics, data pipelines, and metrics platforms Experience with AI tools development and enablement efforts Strong technical judgment, including the ability to evaluate tradeoffs, i

mongodbawsazure
View job →
O
1mo ago

About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t

awskubernetesrest
View job →

Manufacturing Engineer (Associate or Experienced) Company: The Boeing Company Boeing Defense, Space & Security (BDS) Laser and Electro-Optical Systems (LEOS) site is seeking an Experienced Manufacturing Engineer for our advanced development and production manufacturing team located in Albuquerque, NM. The Boeing Laser and Electro-Optical systems (LEOS) group works to develop and implement next-generation technologies and products serving a variety of commercial and military customers. Our team focuses on rapid prototyping and precision production of advanced laser and electro-optical systems that satisfy challenging product performance requirements built in accordance with aerospace (AS9100) quality standards. Position Responsibilities: Design, develop and optimize manufacturing processes and tooling approaches for complex electro-optical aerospace parts and assemblies Collaborate with design engineers to ensure manufacturability and cost-effective production Review engineering drawings and provide comments to responsible engineer Participate in Integrated Product Teams (IPTs) to integrate technical solutions across multiple disciplines Participate in supplier selection and evaluation to ensure adherence to quality and delivery standards Write content for work instructions, travelers, and standard shop procedures Utilize electronic work instruction, material control and non-conformance management systems as needed to support production Ensure compliance with aerospace industry quality standards (e.g., AS9100, J-STD) Ensure the manufacturing work instructions and special processes follow safety and environmental regulations Assist in shop layout and

supply chainrecruitment
View job →
N
11 days ago

We are now looking for a Senior Deep Learning Software Engineer, PyTorch. NVIDIA is hiring software engineers to design and build tools used by AI engineers across the world to design, develop, and deploy AI applications scalable across thousands of GPUs. This position will embed you in an ambitious and diverse team that influences all areas of NVIDIA's AI platform as well as directly contributes to PyTorch, a premiere deep learning framework. In this role you will work with multiple teams at NVIDIA across fields, as well as collaborate internationally with the PyTorch community to develop the best AI platform in the world. What you will be doing: Design and build PyTorch components that run efficiently on supercomputers with 1000s-100ks of GPUs. Collaborate with NVIDIA’s hardware and software teams to improve the overall GPU performance in PyTorch. Design, build and support production AI solutions used by enterprise customers and partners. Work with internal applied researchers to improve their AI tools. What we need to see: BS in Computer Science or Engineering (or equivalent experience). 3+ years professional experience in deep learning. Proficient with C++ programming. Strong understanding of systems software and interfaces. Demonstrated experience with Thread and Distributed Parallel Programming Demonstrated background developing large software projects. Strong verbal and written communication skills Ways to stand out from the crowd: Contributions and participation in the open source community. Familiarity with deep learning compilers. Familiarity with deep learning modeling trends. Background with CUDA Programming as well as Python.

pythonartificial intelligenceai
View job →

Experienced Embedded Software Engineer Company: The Boeing Company The Boeing Company has an exciting opportunity for an Experienced Embedded Software Engineer to support the ADaPS organization in Huntsville, AL. ADaPS (an organization of over 100 software engineers) supports programs across the entire portfolio of Boeing products including Space Launch Systems (SLS) Upper Stage, Satellite Efforts, and Advanced Missile Defense efforts including Patriot Advanced Capability (PAC3). ADaPS specializes in implementing rapid prototyping and development programs using small, adaptable agile teams. Position Responsibilities: Develops and maintains tactical software for the PAC3 RF Seeker system Develops, analyzes, and tests software requirements, algorithms, and designs Develops, maintains, executes, and documents software tests to verify software system requirements. Supports software project management Establishes and maintains a modern software development environment, with formalized processes for development, implementation, integration, testing, documenting, and deploying software changes Communicates technical content to stakeholders Develops and deploys processes and tools Tracks and evaluates software team and supplier performance to ensure product and process conformance to project plans and industry standards Performs software research and trade studies. Troubleshoots software issues Assists in leading and mentoring junior software engineers Security Clearance: This position requires the ability to obtain a U.S. Security Clearance for which the U.S. Government requires U.S. Citizenship. An interim and/or final U.

linuxproject managementrecruitment
View job →
D
DevRev
📍 Bengaluru• Full-time
19 days ago

About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. About the Role: As a Forward Deployed Engineer at DevRev, you will work closely with our pre- and post-sales Customer Experience team by developing solutions over the core platform via applications written in TypeScript, JavaScript, Python, etc. One of the key functions of this role is, building the Knowledge Graph that includes integration of DevRev with other SaaS and non-SaaS systems through webhooks and APIs and/ or real-time communication architectures. You will be expected to employ your creativity and expertise to ideate and shape how AI, analytics and workflows are used in customers' processes. What You'll Do: Customer-facing: 30% working directly with customers. Coding & Integration: 70-80% hands-on technical implementation Build & Deploy Solutions: Design, develop, and launch AI agents, integrations, and automations that connect DevRev with customers' existing tech stacks and workflows. Integrate Systems: Connect DevRev with SaaS and non-SaaS platforms through APIs, webhooks, and real-time communication architectures for seamless data flow. Optimize AI Performance: Apply prompt engineering, fine-tune semantic search engines, and leverage generative AI techniques to enhance agent ac

javascripttypescriptpython
View job →
MI
Mitsogo Inc
📍 Atlanta• Full-time
19 days ago

About Hexnode Hexnode, the Enterprise software division of Mitsogo Inc., was founded with a mission to simplify the way people work. Operating in over 100 countries, Hexnode UEM empowers organizations in diverse sectors. Fueling the transformation to a seamless ecosystem of connected tools, Hexnode is revolutionizing the enterprise software and cybersecurity landscape. Role Overview We are seeking a AWS Operations Specialist to manage and maintain our cloud infrastructure and device ecosystems. This is a highly operational, execution-focused role—not an architecture position. The ideal candidate has 2 to 4 years of experience executing infrastructure as code, monitoring environments, and following documented playbooks to keep our systems secure and resilient. Because this role handles secure environments, candidates must be US Citizens and capable of passing a comprehensive federal background check. Key Responsibilities Infrastructure Execution: Run, maintain, and execute existing Terraform and Ansible scripts to deploy and update infrastructure. GovCloud Monitoring: Actively monitor our AWS GovCloud dashboards, keeping a close eye on system health, performance metrics, and security baselines. Mobile Device Management: Manage Android Enterprise kiosk configurations, ensuring secure deployments and smooth device operations. Incident Response & Triage: Respond swiftly to operational alerts by strictly following our documented team playbooks. Escalation: Identify anomalies or issues that fall outside established, documented procedures and escalate them accurately to the engineering team. Required Qualifications & Profile Citizenship: Must be a US Citizen (required for GovCloud infrastructure management). Background: Must be able to successfully clear a rigorous federal background investigation. Experience: 2 to 4 years of hands-on experience in a technical operations, DevOps, or SysAdmin role. Technical Familiarity: Comfort executing/running Terraform and

awsaiswift
View job →
E
19 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Drive the mission-critical quality strategy for the industry’s most innovative, high-performance storage array platform. In this pivotal engineering leadership role, you will scale systems testing, feature interoperability, and test automation to ensure zero-downtime resilience for global enterprise applications. Partnering directly with cross-functional development, support, and escalation engineering teams, you will champion a customer-first quality model across physical hardware and cloud-native environments (Cloud Block Store, CloudSnap). This position elevates product reliability and shapes how cutting-edge software resilience is delivered at scale. WHAT YOU'LL DO Define & Execute Quality Strategy: Ownership of end-to-end system test designs, focusing on feature interoperability at scale to guarantee zero-downtime performance across enterprise and cloud environments. Build High-Impact Automation & Tooling: Design and deploy automated test workflows and triage tooling to accelerate defect detection, drastically reducing execution friction across thousands of automated test suites. Real-World Customer Simulation: Replicate complex customer deployment architectures and enterprise application workflows to validate real-world resilience, fault tolerance, and resilience against failure domains. Root-Cause Resolution & Continuous Improvement: Partner directly with escalation and support teams to reproduc

javaawsrest
View job →
F
19 days ago

About Forma.ai: Forma.ai is a Series B startup that's revolutionizing how sales compensation is designed, managed and optimized. We handle billions in annual managed commissions for market leaders like Edmentum, Stryker, and Autodesk. Our growth has been fuelled by our passion for fundamentally changing and shaping how companies use sales intelligence to drive business strategy. We’re welcoming equally driven individuals who are excited about creating something big! Senior Staff Backend Engineer About the Team We build enterprise software that helps organizations optimize sales performance, enabling go-to-market agility. Our engineering organization includes multiple product application teams responsible for delivering core customer-facing capabilities. We are seeking a Senior Staff Backend Engineer to join our application teams and help set technical direction across multiple domains within engineering. You'll work alongside staff, senior, and early-career engineers, and partner closely with engineering leadership to define, evolve, and scale the systems that power enterprise-grade product workflows. This is an opportunity to own complex, multi-domain technical problems and shape product direction beyond a single team. We are low on meetings, high on accountability. Most of the teams are in the EST time zone, but we have a few located in AST, PST, and Central as well. What you'll be doing You will play a pivotal role in shaping the technical direction of our application stack across multiple domains. You will lead development efforts for our most complex initiatives, the kind that span two or more teams or product areas, and serve as a technical benchmark for system design, code quality, and long-term maintainability. You'll operate at the intersection of data modelling, business logic, and enterprise-scale reliability, and your work will often set standards that neighboring teams adopt. This remains a hands-on

javascripttypescriptpython
View job →
🔔

Get new performance and systems engineer jobs by email

Daily job updates · Unsubscribe anytime