Jobiba hiring network

Performance And Systems Engineer Jobs

6,348 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current performance and systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

C
19 days ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? This role is for people who love building tools for their coworkers. The Internal Applications team creates tools that help us create better models. In this role you will collaborate with internal stakeholders, which include annotators, ML researchers, product managers and more. Join our team of builders who create tooling that will pave the way for the next generation of large language models! As a Full-Stack Software Engineer on the Internal Applications team, you will: Work with a small talented and enthusiastic team of software engineers Contribute to delightful experiences for our user-facing products, meticulously crafting code for browsers and servers Collaborate and grow with your engineering colleagues of all levels through direct pairing sessions, architectural designs, documentation and talks Identify and remove roadblocks to enable your team to increase its engineering velocity. Build resilient systems that are mission-critical Keep up with the cutting edge and adopt new technologies to improve performance and reliability You may be a good fit if: You have experience shipping products with a large numb

REMOTEtypescriptpythonreact
View job →
F
Fin
📍 London• Full-time
29 days ago

Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What’s the opportunity? Fin's websites, including fin.ai , are strategic growth levers for the company — powering storytelling, brand, experimentation, and product discovery. As our Staff Engineer, you will shape the technical direction and architectural evolution of our web platform and systems. You’ll act as the technical leader across our marketing and growth web surfaces. This is a cross-functional engineering leadership role, partnering with marketing, analytics, design, data science, and engineering teams — to ensure our web stack is modern, performant, measurable, and delightful to build on. You’ll operate with a high degree of autonomy and will be accountable for setting the long-term technical strategy for the team and executing against it. What will I be doing? As a senior technical leader, you will: Own and evolve the architecture of Fin's web platform. Define the long-term technical strategy for the web team, focusing on scalability, performance, developer productivity, o

reactawsai
View job →
N
Nvidia
📍 Santa Clara, United States
1mo ago

NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel

pythonkuberneteslinux
View job →
C
Clickup
📍 Canada• Full-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview We’re looking for a Staff Frontend Engineer to lead the design and development of major frontend systems and product initiatives at ClickUp. You’ll work across teams to solve complex technical challenges, improve engineering velocity, and help shape the direction of our frontend architecture. This role is for someone who thrives in fast-paced environments, drives clarity in ambiguity, and can influence both technical decisions and execution across the organization. What You’ll Do Lead development of complex product features and frontend systems in Angular 2+ and React. Partner with backend, integrations, product, design, and QA to deliver high-quality user experiences at speed Architect scalable, reusable frontend patterns that improve product quality and developer velocity Identify and address performance bottlenecks, UI architecture issues, and scalability risks Drive engineering best practices across testing, observability, code quality, and maintainability Help teams make strong technical decisions under tight timelines and evolving priorities Own delivery across large initiatives, balancing immediate product needs with long-term technical health Mentor engineers and elevate frontend craftsmanship across the team Contribute to improving how frontend engineers work together across domains Qualifications 7+ years of frontend engineering experience, with deep expertise in Angular 2+ and React Strong command of TypeScript, RxJS, NgRx , and modern frontend architecture patterns Experience building reusable component systems and scalable frontend application structures Deep knowledge of per

typescriptreactangular
View job →
C
Clickup
📍 United States• Full-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview We’re looking for a Staff Frontend Engineer to lead the design and development of major frontend systems and product initiatives at ClickUp. You’ll work across teams to solve complex technical challenges, improve engineering velocity, and help shape the direction of our frontend architecture. This role is for someone who thrives in fast-paced environments, drives clarity in ambiguity, and can influence both technical decisions and execution across the organization. What You’ll Do Lead development of complex product features and frontend systems in Angular 2+ and React. Partner with backend, integrations, product, design, and QA to deliver high-quality user experiences at speed Architect scalable, reusable frontend patterns that improve product quality and developer velocity Identify and address performance bottlenecks, UI architecture issues, and scalability risks Drive engineering best practices across testing, observability, code quality, and maintainability Help teams make strong technical decisions under tight timelines and evolving priorities Own delivery across large initiatives, balancing immediate product needs with long-term technical health Mentor engineers and elevate frontend craftsmanship across the team Contribute to improving how frontend engineers work together across domains Qualifications 7+ years of frontend engineering experience, with deep expertise in Angular 2+ and React Strong command of TypeScript, RxJS, NgRx , and modern frontend architecture patterns Experience building reusable component systems and scalable frontend application structures Deep knowledge of per

typescriptreactangular
View job →
D
1mo ago

We are building the best platform in the world for engineers to understand, scale, and protect their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, application tracing, and security insights for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way . At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Solve a scaling bottleneck in a critical service Deploy a new feature to production, progressively rolling it out with feature flags Investigate and fix a production issue from a service your team owns Design a way to scale up a service for more traffic With your team, plan the most important projects to work on next Who You Are: You have significant experience in one or more languages You value code simplicity and performance You can design architecture to solve problems at high scale You have a BS/MS/PhD in a scientific field or equivalent experience You want to work in a fast, high-growth startup environment that respects its engineers and customers You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how You have demonstrated ability to use AI coding tools in day-to-day workflows and build, validate, and refine AI-generated output in products You can design AI Backend systems, with awareness of quality, cost, and latency tradeoffs Bonus: You’re motivated to push the boundaries of how AI can improve software engineering best practices and contribute to building AI-enabled products Datadog values people from all walks of life. We understand not everyone will meet all

aigorust
View job →
M
1mo ago

The MongoDB Query Execution Team is hiring software engineers who want to join us in developing a high performing, reliable and modular distributed query system. Our engineers work on implementing and maintaining execution algorithms, building new query language features, tuning database performance, and more to power our customers' critical workloads. This role can be based out of our Dublin office or remotely in Ireland. Relocation can be supported. Position Expectations Understand and improve current functionality of the MongoDB query engine Contribute high quality C++ code and give and solicit feedback in code reviews Identify, design, implement, test, and support new features related to query performance and robustness, query language enhancements, diagnostics for query performance problems, and integration with other products and tools Work constructively with peers to deliver excellent technical solutions Candidate Profile 5+ years of experience in systems programming Experience in databases and/or data management systems is a huge plus, but not a requirement Hands-on experience building industrial-strength software Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Experience with large code bases, preferably in C++, C, Rust or a similar compiled language B.Sc in Computer Science or similar field, or equivalent practical experience Interest in the theory and practice of database query engines. Hands-on experience or M.Sc./Ph.D in the domain is a plus Success Measures In three months you’ll have contributed to the development of a project slated for the next major version, as well as fixed a few bugs in a minor version of our latest stable release series In six months, you’ll have taken on code review responsibilities and are independently delivering complex functionality and squashing bugs independently In twelve months, you’re leading the development of a new major feature and are h

mongodbawsazure
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workflows Pr

mongodbawsazure
View job →
M
Mongodb
📍 New York City• Full-time• From $157K/yr
1mo ago

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in New York City for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workf

mongodbawsazure
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Cork for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workflows Pref

mongodbawsazure
View job →
O
1mo ago

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

pythonawslinux
View job →
O
1mo ago

Location: San Francisco, CA (Hybrid: 4 days onsite/week). Relocation assistance available. About the Team: We build foundational platform software that enables reliable, secure, and performant products. The team works across system layers and partners closely with adjacent engineering groups to deliver robust capabilities from concept through launch. About the Role: We’re seeking a System Software Engineer to design, implement, and debug core platform components and the pipelines that build and update system images. You’ll work across operating system layers, focusing on performance, security, and deep system debugging to ship production‑grade systems. In this role, you will: Design, implement, and debug system‑level components and services across kernel and user space. Configure and maintain OS platform services (init, services, networking, security policies) and related tooling. Build and operate image and update pipelines, ensuring reliability, reproducibility, and rollback safety. Instrument and analyze performance using profiling and tracing; optimize CPU, memory, I/O, and power usage. Own platform observability and reliability: logging, crash capture, watchdogs, and diagnostics. Collaborate with cross‑functional teams to define interfaces and deliver end‑to‑end features. Establish strong engineering practices: code review, CI, reproducible builds, and release management. Partner with external suppliers to support builds and deployments. You might thrive in this role if you: Have shipped production systems software on modern operating systems. Are proficient in C/C++ and a scripting language, and comfortable with OS internals (concurrency, memory management, filesystems, networking, power management). Bring strong systems debugging skills using debuggers, tracers, profilers, and logs across kernel/user‑space boundaries. Understand configuration of platform services and interfaces, and can translate requirements into stable, well‑documented APIs. Are fluent in u

awsrestai
View job →
O
1mo ago

About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun

awsazuregcp
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

pythonlinuxartificial intelligence
View job →
C
12 days ago

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary CVS Health's Adjudication & Client Experience Engineering organization is seeking a motivated and highly skilled Senior Analyst - Software Development Engineering to join our Application Production Support team. This role will support critical business applications by providing production support, troubleshooting technical issues, and contributing to ongoing application enhancements and stability improvements. As a Sr. Analyst, you will work closely with Lead Engineers, Software Development Engineers, Product Owners, QA teams, and business stakeholders to investigate and resolve production incidents, perform root cause analysis, and implement code fixes. You will be responsible for supporting enterprise applications built on Java, Angular, APIs, and Cloud platforms while ensuring the reliability and performance of systems that serve our PBM (Pharmacy Benefit Management) business. This role is ideal for a hands-on engineer who enjoys solving production challenges, developing software solutions, and collaborating within a fast-paced environment. The successful candidate will contribute to application support activities, system enhancements, and continuous improvement initiatives while growing their technical and business domain expertise. Required Qualifications 5-8 years of experience in software development, application s

REMOTEtypescriptjavaangular
View job →
🔔

Get new performance and systems engineer jobs by email

Daily job updates · Unsubscribe anytime