Jobiba hiring network

Performance And Systems Engineer Jobs

6,348 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current performance and systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

D
Datadog
📍 New York• Full-time• From $244K/yr
1mo ago

About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Opportunity: Datadog’s Staff Engineers are our technical leaders operating at the forefront of technology, building solutions that take us through at least our next five years of growth. They do this in three major ways: As individual contributors, they bring world class technical abilities to deliver industry leading systems in areas such as data visualization, virtual runtime profiling, and planet scale streaming. As technical leaders they bring experienced technical breadth and communication skills to tackling design and architectural problems spanning the organization, charting the right course, then leading delivery. In both roles they participate in the staff engineering community and help us learn from what the industry is doing and what we've built before, and so improve company wide standards around software and systems engineering. Some examples of projects a staff engineer may own include designing and building a new data storage engine handling hundreds of millions of records per second, being the lead engineer building a new product like synthetics or profiling, or rebuilding a critical service to handle the next two orders of magnitude of scale. What You'll Do: Be the technical owner of multiple pieces of critical architecture in your area of the business Own delivery of the systems you architect from beginning-to-end, doing what it takes to get things shipped and at full scale in production Dive deep into performance of systems; inventing new approaches that bring efficiency at scale Who You Are: You have a BS/MS/P

aigorust
View job →
L
Lyft
📍 Toronto• Full-time• From C$108K/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As an Observability team member, you are responsible for the operation and maintenance of our logging and metrics infrastructure. You ensure all teams at Lyft are aware of the operational health of their products by monitoring system availability and take a holistic view of our platform performance. You build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you find opportunities to improve our systems in order to push our platform forward. You provide our partners with the support they need to help them build robust large scale distributed systems. We count on the reliability of our infrastructure to empower Lyft teams to provide our customers rich experiences that are highly available with rock solid performance to ensure our transportation platform continues to connect people and places. As we grow our team, we are seeking experienced Infrastructure Engineer to ensure that as our Infrastructure continues to scale, our platform continues to provide an essential and dependable service that transports millions of people every day. Specifically we are searching for someone who brings fresh perspectives, enjoys collaborating with cross-functional teams in order to continually improve our products and services for our customers. Responsibilities: Maintain, improve, and develop tooling and systems that enhance the reliability, scalability, and efficiency of our platform. Assist engineering teams in defining service-level objectives (SLOs) and provide the necessary toolin

pythonawskubernetes
View job →
T-
14 days ago

About the Role: The Machine Learning team at Tubi drives the innovation behind personalized user experiences. With the largest inventory in the industry and hundreds of millions of viewers, we tackle problems in the space of recommendations, search, content understanding and ads optimization that shape the future of streaming. We are seeking a highly skilled Senior Machine Learning Engineer to contribute to transformative projects in video personalization. In this role, you will design and implement advanced algorithms and systems to improve our personalization strategy. As a senior technical expert, you will tackle complex problems in machine learning at scale, collaborating closely with cross-functional teams to develop and optimize machine learning-driven solutions. This is a hybrid role in our Toronto office. What You'll Do: Design, develop, and implement recommendation systems and algorithms for a global audience Conduct deep dives into algorithmic components and systems, ensuring that models are optimized for both performance and scalability across multiple regions and product areas Build and deploy high-impact robust ML pipelines, including data extraction, feature development, model training, testing, and deployment Continuously monitor, evaluate, and optimize the performance of deployed models, ensuring they meet business goals and provide high-quality user experiences. Work closely with Product, Engineering, and Data Science teams to align on product requirements, set expectations, and deliver machine learning-driven solutions that improve user engagement Your Background: 3+ years of industry experience building production Machine Learning systems BS, MSc, or Ph.D. in Computer Science, Machine Learning, Statistics, Mathematics, or a related field Experience with deep learning technologies for recommendation systems, including TensorFlow, PyTorch, or similar frameworks Proficiency in building and deploying full-stack machine learning pipelines: data e

machine learningaigo
View job →
O
OpenAI
📍 New York• Full-time• Remote
21 days ago

About the Team API Frontiers turns OpenAI’s frontier models into production APIs that developers can use to build reliable products and agents. We own the core path connecting models to developers through the Responses API, with a focus on safety, reliability, and speed. Working closely with Research, Safety, Codex, and other API teams, we bring new model capabilities into production and improve them through developer feedback. About the Role We are looking for a backend software engineer to build and operate the services behind the Responses API. You will shape API behavior, bring new capabilities from research into production, and make long-running agent workflows dependable and fast. The work combines distributed systems engineering with product judgment: designing useful developer interfaces, managing staged rollouts, and following production issues through to durable fixes. In this role, you will: Design, build, and operate APIs and backend services that bring frontier model capabilities to developers. Partner with Research, Safety, Codex, and API teams to define API behavior and deliver safe, staged launches. Build API capabilities for agent workflows, including task delegation, context sharing, and parallel execution. Strengthen long-running request reliability across timeouts, cancellation, streaming, and background execution. Improve request-processing performance and tail latency through profiling, efficient systems code, and persistent connections. Turn developer feedback and production failures into better observability, diagnostics, and lasting product improvements. Your background might look something like: 5+ years of experience building and operating backend services or developer-facing APIs in production. Strong software engineering fundamentals, with practical knowledge of distributed systems, concurrency, and asynchronous execution. Ability to diagnose production failures and performance bottlenecks using observability data and profiling. Product

REMOTEawsrestai
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role We’re a team of high-output generalists where ML and systems engineering converge. This is not a "run the models" role. We reason from first principles about why a perception model learns what it learns, close the gaps that cap its performance, and raise the bar on the data and evaluation loop that drives autonomy. Your work will directly impact how autonomous systems understand rare scenarios, adapt to global geographies, and scale safely. About the work You’ll solve autonomy’s hardest data challenges through applied ML and systems rigor: Diagnose why perception models underperform on the long tail, and turn that into targeted data and training priorities. Design eval metrics and regression detection that tell us whether a model is ready. Curate and clean training data for segmentation and occupancy; hunt the data problems that silently cap performance. Run controlled ML experiments and ablations; cleanly sepa

pythonrestai
View job →
M
Mongodb
📍 Ireland• Full-time
1mo ago

The data management software market is transforming how organisations build and run applications. MongoDB is the leading developer data platform and the first database provider to IPO in more than 20 years. Join us at the forefront of data and application development. MongoDB Technical Services Engineers combine deep technical expertise with exceptional problem-solving and customer-service skills. You’ll advise customers and resolve complex challenges across MongoDB Core, drivers, Atlas, Cloud Manager, cloud platforms, and infrastructure. We’re looking for candidates based in Dublin to join our vibrant office and collaborative in-office team. This is a five-day role with one of the following schedules: Tuesday–Saturday, Sunday–Thursday, or a five-day pattern covering both Saturday and Sunday. Under our hybrid model, employees on weekend schedules are expected to work from the office two days per week. Cool things you’ll do You’ll help customers troubleshoot complex issues and run critical MongoDB workloads with confidence. You’ll: Solve customer challenges across architecture, performance, recovery, and security Lead investigations from diagnosis to resolution, providing clear, actionable guidance Partner with Product Management and Engineering to advocate for customers and improve MongoDB Build tools, documentation, and training while mentoring peers and raising technical excellence What you need We value curiosity, adaptability, strong technical foundations, and a genuine desire to help customers. You should bring many of the following: 5–6 years of experience in technical support, systems engineering, database administration, SRE, or a related field Experience running complex, mission-critical production database systems Strong Linux and systems engineering skills, including performance, memory, I/O, storage, networking, security, clustering, and troubleshooting A solid understanding of networking concepts and protocols, including DNS, TCP/IP, and SSL/TLS Ability

javascriptpythonjava
View job →
I
Intel
📍 San Jose, Costa Rica
9 days ago

Job Details: Job Description: Join Intel's Silicon and Platform Engineering Group (SPE) as an Exempt Tech Contract Employee, The Role and Impact :As a Power Integrity Engineer, you will play a pivotal role in ensuring Intel's platforms and products meet the highest standards of performance and reliability. This position offers you the opportunity to work on cutting-edge technologies, designing and optimizing power delivery networks for complex high-speed platforms, boards, packages, and silicon. Your contributions will directly impact Intel's ability to deliver industry-leading solutions to its customers, driving innovation in power integrity and enhancing the performance and functionality of our silicon and systems. Join us to tackle exciting challenges and create impactful solutions that push technological boundaries. The primary responsibilities for this role will include, but are not limited to: Develop and analyze power delivery networks, including 2D and 3D model extraction and noise analysis across die/C4 bumps, silicon, package, sockets, and boards. Define power grid specifications and power and area targets to achieve optimal balance between power integrity and performance. Design test structures, electrical analysis methodologies, and verification plans to address power integrity challenges. Perform measurements to characterize power noise profiles across frequency, ground bounce, and other critical metrics, verifying power delivery networks post-design. Correlate measurements back to presilicon analysis and estimations for improved accuracy. Apply expertise in power integrity design and tradeoffs to guide simulations of power networks and physical implementations for packages and platforms. Ensure on-die power noise meets system-on-chip (SoC) functionality an

recruitment
View job →

About the Role: The Machine Learning team at Tubi drives the innovation behind personalized user experiences. With the largest inventory in the industry and hundreds of millions of viewers, we tackle problems in the space of recommendations, search, content understanding and ads optimization that shape the future of streaming. We are seeking a highly skilled Staff Machine Learning Engineer to contribute to transformative projects in video personalization. In this role, you will design and implement advanced algorithms and systems to improve our personalization strategy. As a senior technical expert, you will tackle complex problems in machine learning at scale, collaborating closely with cross-functional teams to develop and optimize machine learning-driven solutions. This is a hybrid role in our Toronto office. What You'll Do: Lead the design, development, and implementation of advanced recommendation systems and algorithms for a global audience Conduct deep dives into algorithmic components and systems, ensuring that models are optimized for both performance and scalability across multiple regions and product areas Build and deploy high-impact robust ML pipelines, including data extraction, feature development, model training, testing, and deployment Continuously monitor, evaluate, and optimize the performance of deployed models, ensuring they meet business goals and provide high-quality user experiences. Work closely with Product, Engineering, and Data Science teams to align on product requirements, set expectations, and deliver machine learning-driven solutions that improve user engagement Your Background: 8+ years of industry experience building production Machine Learning systems MSc or Ph.D. in Computer Science, Machine Learning, Statistics, Mathematics, or a related field Experience with deep learning technologies for recommendation systems, including TensorFlow, PyTorch, or similar frameworks Proficiency in building and deploying full-stack machine le

machine learningaigo
View job →
G
14 days ago

Job Summary Reporting to the Memory Validation leadership team, the Senior Silicon DDR/HBM Validation Engineer will be responsible for the bring-up, validation, characterization and debug of advanced memory subsystems used in next-generation AI compute platforms. The role will focus on DDR and HBM technologies, working closely with silicon design, firmware, characterization, platform and systems teams to ensure robust memory subsystem functionality, performance and reliability. The successful candidate will take ownership of significant validation activities, contribute to debug and root-cause analysis efforts, and help improve validation methodologies, automation and infrastructure. The Team The Memory Validation team sits within the Validation organisation and is responsible for the bring-up, validation, characterization and debug of memory subsystems across Graphcore silicon and platform products. The team supports DDR and HBM validation activities throughout the product lifecycle, from first silicon through production readiness. Engineers work closely with architecture, RTL, firmware, characterization, systems and platform teams to ensure memory technologies meet functionality, performance, reliability and performance objectives. Responsibilities and Duties Execute validation and bring-up activities for DDR and HBM memory subsystems Verify memory bring-up software, firmware and scripts against defined project requirements Debug firmware, hardware and system-level issues and contribute to root-cause analysis activities Analyse system logs, validation data and characterization results to identify failures and performance issues Perform PHY characterization and analog-level analysis during stress testing and validation activities Develop and execute functional, stress, performance and corner-case validation tests Perform signal integrity, voltage, frequency and timing measurements using laboratory instrumentation Char

pythonaigo
View job →
C
Cloudflare
📍 Hybrid• Full-time• Hybrid
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX About the role Reporting to the Director of Vulnerability Management, the Lead Vulnerability Management Engineer will serve as the technical authority for the program, focusing on the architecture, systems design, and long-term technology strategy for Cloudflare's vulnerability management capabilities. You will lead the evaluation and integration of advanced security technologies and systems across global infr

pythonawsai
View job →
C
Cloudflare
📍 Distributed; Hybrid• Full-time• Hybrid• $185K – $275K/yr
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations Atlanta, US Austin, US Denver, US New York, US Toronto, Canada Washington DC, US Seattle, WA Remote candidates within North America will also be considered. About the Role The Security Platform team is an infrastructure/developer tools group tasked with building and operating powerful, resilient, and secure infrastructure and systems that enable other engineering teams to deliver products to our customers efficiently and secure

pythonawslinux
View job →
C
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX About the team Cloudflare’s mission is to help build a better internet and the Threat Application Services (TAS) team lives at the core of that effort. The team builds the automated services and systems that transform Cloudflare’s massive network telemetry into actionable intelligence. We empower customers with products to combat advanced threats, enabling them to query, investigate, and automatically mitigate attac

javascriptjavasql
View job →
C
Cloudflare
📍 Hybrid• Full-time• Hybrid
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, Atlanta Software Engineer, Network Firewall Role Summary As a Software Engineer on our team, you will work across a wide range of technologies and systems to deliver new features, improve performance, and increase the scalability of our Network Services products. You will help evolve our Network Firewall and related security capabilities across both north-south and east-west network architectures, with demand

awslinuxrest
View job →
T
Tenstorrent
📍 Toronto• $100K – $500K/yr
9 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As a Datacenter Liquid Cooling Architect, you will define, design, and architect next-generation liquid cooling infrastructure for Tenstorrent’s large-scale AI training and inference clusters. You will partner with systems engineering, mechanical engineering, software, and cross-functional design teams to develop chassis-, rack-, and cluster-scale cooling solutions, including CDU integration, telemetry and control, leak detection, and resilient operating strategies. This role will help shape reliable AI datacenter architectures and deployments for both internal and external customers. This role is on-site, based out of Toronto, Canada, Austin, Texas or Santa Clara, California. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A datacenter and system thermal design professional with 10+ years of experience architecting cooling infrastructure for complex computing environments. An experienced liquid cooling architect who can design chassis- and rack-scale solutions for large AI training and inference clusters. A systems thinker who understands how mechanical, electrical, software, facility, and systems engineering decisions come toge

T
14 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. This role sits at the center of cutting-edge AI hardware development, keeping the servers, PCIe systems, and engineering infrastructure running that power next-generation compute. You’ll be hands-on with rapidly evolving prototype and production systems, installing, maintaining, and troubleshooting hardware in fast-paced R&D and data center environments. Acting as a critical bridge between hardware engineers, software teams, and IT, you’ll help ensure seamless access to the platforms that turn ideas into working silicon and systems. This role is onsite, based out of Toronto, Canada or Austin, Texas. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A hands-on hardware professional who enjoys building, maintaining, and troubleshooting complex computing systems. Comfortable working in fast-paced R&D environments where hardware, firmware, and software are constantly evolving. Knowledgeable in computer architecture, operating systems, and hardware diagnostics, with strong problem-solving skills. Collaborative, detail-oriented, and motivated to improve processes through documentation, scripting, and automation. What We Need Inst

awsaisem
View job →
🔔

Get new performance and systems engineer jobs by email

Daily job updates · Unsubscribe anytime