About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role As a Staff Machine Learning Engineer on Sentry’s AI/ML team, you’ll be directly responsible for developing the models and agents used to make our product smarter and more capable. This role is crucial; you will be at the forefront of integrating AI and machine learning into our core products, from issue triage and resolution to predictive analytics for application performance monitoring. Your work will help companies around the globe gain actionable insights into their software, enabling them to build better products, faster. In this role you will Build state-of-the-art agentic AI systems to triage, debug, and solve real production issues Leverage Sentry’s novel (and massive) dataset of errors, spans, and profiles Own the development of major initiatives in the AI/ML space You'll love this job if you Are driven by impact and enjoy working on high-stakes, high-visibility projects Enjoy building things. You will have the opportunity to join the AI/ML team as one of its foundational members Thrive in cross-functional teams and enjoy building features alongside developers and product teams Qualifications Minimum 4+ years of professional experience with a MS/PhD degree in computer science, machine learning, or a related field Minimum 6+ years of professional experience with Bachelor’s degree in computer science, machine learning, or a related field Demonstrated expertise building production-grade agentic systems and tools You are comfortable writing production quality code (we use Python) Expertise with deep learning frameworks (we use PyTorch) Familiarity in deploying machine learning models at scale in production
Jobiba hiring network
Performance And Systems Engineer Jobs
6,348 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current performance and systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX - Hybrid About the team The Data Intelligence & Analytics organization builds the core data platform and internal products that power decision-making across the company. We design and operate large-scale data systems, own the company’s data lake, ingestion infrastructure, and platform tooling, and develop end-to-end applications that transform complex datasets into fast, reliable, business-critical products used
Executive - Building Management Systems (BMS) is responsible for overseeing the engineering, maintenance, and operational management of all BMS systems within the airport. This includes ensuring the proper functioning, reliability, and performance of systems such as HVAC, lighting, fire alarm systems, access control, and other integrated mechanical, electrical, and civil systems that are controlled through BMS. The role ensures preventive and corrective maintenance, and manage troubleshooting for BMS-related issues across airport facilities. Source: Adani Group | Job ID: 46328
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role AI and machine learning are reshaping how developers debug, monitor, and ship software, and Sentry is uniquely positioned to lead that shift. We sit on a novel and massive dataset of real production errors, spans, and logs from tens of thousands of engineering organizations — the kind of signal that makes ML genuinely useful, whether it's a clustering model that groups related issues, a ranking system that surfaces the right alert at the right time, or an agent that proposes a fix. We're looking for an Engineering Manager to lead and grow our Machine Learning Engineering team. This team owns the full spectrum of ML at Sentry: classical techniques like clustering, ranking, anomaly detection, and embeddings that quietly power core product surfaces today, alongside the LLM-based and agentic systems shaping where the product is headed. You'll partner closely with product, design, and engineering leaders to decide where ML belongs in our products, what kind of ML actually fits the problem, and how we translate that work into experiences millions of developers rely on every day. In this role you will Set technical direction across the team's full ML surface area — from classical models for clustering, ranking, and anomaly detection to LLM-based and agentic systems — and make sharp calls about which approach fits each problem Define how the team evaluates and monitors ML systems in production, from offline metrics to online experimentation to model and agent observability Stay hands-on enough to review code and model designs, contribute to architecture discussions, and unblock engineers on complex ML problems Define
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides tools that help developers find and fix issues in their applications. The Developer Infrastructure team owns the systems that make every engineer at Sentry more effective at shipping high quality software: developer environments, CI/CD for our open source and closed source codebases, the golden path for building new services, our SDK and library publishing tooling, and the metrics and dashboards that keep it all healthy. Our goal is straightforward: developers should spend their time thinking about the software they build for customers, not the software they use to build it. That means everything from local and cloud development environments, to the CI/CD pipeline that gets code out safely and quickly, to the tooling that catches flaky tests before they cost someone a day. As AI coding agents become part of how engineers work, we're also investing in making sure our environments, CI, and deployment systems support that shift. As a Software Engineer on Dev Infra, you'll help build and scale this infrastructure end to end. In this role, you will Build and maintain the tooling that powers local and cloud-based developer environments Improve CI for our codebases and CD for our deployment pipeline, keeping both fast and reliable as the org and its infrastructure grow Define and evolve the golden path for building new services and libraries at Sentry Contribute to our SDK and library publishing tooling and release processes Build the metrics, dashboards, and flaky test detection that give the org visibility into deployment health and CI reliability Build tooling that gives engineers, and the AI codin
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role This isn’t a typical engineering role. You won’t be embedded in a single product team or siloed in one product area. Instead, you’ll sit within Platform Engineering, own the AI-assisted coding domain, and work across all of engineering at Sentry, focused specifically on how AI coding agents participate in our software development lifecycle. For AI coding agents to work well in our repo, the internal systems they depend on need to be accessible via API, not locked behind UIs that require human interaction. Right now, many of those systems aren’t agent-ready. You’ll audit and prioritize that gap, expose those systems programmatically, and build the connections that let tools like Claude Code operate on them end-to-end. From there, the scope expands to improving the quality of AI-generated pull requests and automating the engineering work that’s important but consistently deprioritized. You will look from context engineering standpoint to see what to send to our model; you will look from harness engineering standpoint to see the tools it can use, the permissions it has, the state it carries forward, the tests it has to pass, the logs you capture, the retries, checkpoints, guardrails, and evals. You’ll work closely with the dev infrastructure team as your home base, then collaborate across every product team coding in our repo once the tooling foundation is in place. It’s a broad role with real impact, and the work you do will directly change how Sentry engineers ship software. What You’ll Do Audit Sentry’s internal developer systems and make them API-ready for AI agents. You’ll prioritize and drive the work of ex
Our mission: Creating the freedom for SMEs to succeed in business and beyond, by delivering Europe's leading finance workspace. Our journey: Founded by Alexandre and Steve in July 2017, Qonto has rapidly gained trust, serving over 600,000 customers with a team of 1,600+ Qontoers across Europe. Our beliefs: We evaluate applicants based on skills and potential — no boxes to tick. Our team is 55% international, 44% women, and 20% parents. Discover the steps we took to create a discrimination-free hiring process. ⭐ Mission Join us as a Freelance Full-Stack Engineer to build one of Qonto's first internal applications: a modern, high-quality People platform that transforms the daily experience of 1,700+ Qontoers - from compensation cycles and performance reviews to everyday HR workflows. You'll work hand-in-hand with Mélodie, our Head of People Systems, on a high-ownership mission: co-define the product vision, design the architecture, and ship to production. This is not a staff-augmentation engagement - it's a mandate for someone who builds like an owner. 👩💻🧑💻 As a Freelance Full-Stack Engineer, you will • Ship the People platform MVP by December 2026: migrate key Workday modules into a custom internal app (React for the front, Go for the back) that handles comp cycles, performance reviews, and daily HR workflows for 1,700+ Qontoers - on time and with zero compromise on quality. • Co-define the product vision and roadmap with Mélodie: run user discovery with People teams, challenge specs, and help decide what gets built first - not just how to build it. • Design for security and compliance from day one: architect systems that handle sensitive People data (GDPR, access controls, data residency) with the rigor the context demands. • Ship faster with AI: leverage AI tools (Cursor, Claude Code, Copilot…) in your daily workflow and build AI-enabled features where they create real leverage for the People team. • Enable long-term adoption: document your work, r
About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. Role Overview We are seeking a Package Reliability Engineer to lead reliability engineering for advanced packages used in high-performance AI and computing systems. The primary focus of this role is to assess package level mechanical and thermal reliability risks and apply thermal and mechanical modeling to optimize package design, material selection, and assembly processes. The engineer will also develop reliability test plans with external partners, identify failure mechanisms, perform root-cause analysis, and recommend practical corrective actions. In this role, you will assess package reliability risks from early architecture development through product qualification and high-volume manufacturing. You will work closely with package design, silicon design, system engineering, manufacturing, and ASIC partners to predict package behavior, develop qualification strategies, resolve reliability issues, and improve overall package robustness and lifetime. In this role you will: Lead reliability test plan and assessments for advanced HPC packages, including risk identification, potential failure-mechanism analysis, root-cause investigation, mitigation planning, and corrective-action development. Drive reliability-focused package design optimization based on thermo-mechanical modeling to improve package reliability, power integrity, thermal performance, mechanical robustness, and platform scalability. Develop, validate, and apply package reliability models and lifetime-prediction
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity We are a global team of innovators shaping the future of observability. Our intelligent platform gives customers real-time insight into complex systems so they can innovate faster and operate reliably in an AI-first world. If you’re excited by high-throughput distributed systems and want to contribute to one of the largest and fastest-growing observability platforms, we’d love to hear from you. Join a backend engineering team focused on building and operating JVM-based services that ingest, process, and serve massive volumes of telemetry data. You’ll work on high-scale, low-latency systems that power mission-critical observability features used by engineers worldwide. What you'll do Design, build, and operate JVM-based microservices (primarily Java and Kotlin) with a focus on performance, scalability, and reliability. Own services end-to-end: architecture, implementation, deployment, monitoring, on-call participation, and continuous improvement. Apply strong concurrency and performance practices: asynchronous programming, backpressure, efficient I/O, memory management, and GC tuning. Build and evolve event-driven systems; work with Kafka for streaming, partitioning, consumer groups, and schema evolution.Instrument services for deep observability (metrics, logs, traces), define SLIs/SLOs, and use e Experience with Kafka or similar streaming technologies (topic/partition strategy, consumer lag, idempotency, schema compatibility) strongly preferred. Proficiency w
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role As a Senior Software Engineer on Sentry’s AI team, you’ll be directly responsible for developing the platform used by our debugging agents. This role is crucial; you will be at the forefront of integrating AI and machine learning into our core products, from issue triage and resolution to predictive analytics for application performance monitoring. Your work will help companies around the globe gain actionable insights into their software, enabling them to build better products, faster. In this role you will Build state-of-the-art agentic AI platforms to triage, debug, and solve real production issues Leverage Sentry’s novel (and massive) dataset of errors, spans, and profiles Own the development of major initiatives in the AI/ML space You'll love this job if you Are driven by impact and enjoy working on high-stakes, high-visibility projects Enjoy building things. You will have the opportunity to join the AI/ML team as one of its foundational members Thrive in cross-functional teams and enjoy building features alongside developers and product teams Qualifications Minimum 5+ years of professional experience with Bachelor’s degree in computer science, machine learning, or a related field Demonstrated expertise building production-grade agentic systems and tools You are comfortable writing production quality code (we use Python and Typescript) Familiarity with deep learning frameworks (we use PyTorch) Familiarity in deploying machine learning models at scale in production environments The base salary range (or hourly wage range, if applicable) that Sentry reasonably expects to pay for this position is $155,000 to
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. As an Engineering Manager, you’ll lead a team of engineers owning critical workflows such as checkout and invoicing, while also developing new features to help customers manage their spend growth. In this role, you’ll partner across the organization to ensure our customers redeem everything Sentry has to offer and budget for future expansion. In this role you will Strategic Planning & Roadmap: Define and drive the team's roadmap. Align team goals with organizational objectives and contribute to the overall platform strategy. Technical Guidance & Operational Excellence: Provide technical leadership and guidance on complex distributed systems and design. Ensure the team is proactively identifying areas for improvement. Cross-functional Collaboration: Partner closely with business and technical teams to translate business goals into actionable objectives and scalable solutions. Team Leadership & Development: Lead, mentor, and grow a team of talented engineers, including Staff-level engineers. Build a culture of technical excellence, collaboration, continuous
About the Role: The Machine Learning team at Tubi drives the innovation behind personalized user experiences. With the largest inventory in the industry and hundreds of millions of viewers, we tackle problems in the space of recommendations, search, content understanding, and ads optimization that shape the future of streaming. We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction and execution for scaling our machine learning capabilities while ensuring our distributed systems and infrastructure can support innovation at massive scale. You will combine technical depth with leadership excellence to guide teams that deliver both foundational ML systems and high-performance distributed services. This is a hybrid role for our Toronto office. What You'll Do: Lead and manage high-performing teams across ML engineering and ML infrastructure, fostering a culture of innovation, collaboration, and growth. Define and execute the strategic roadmap for ML systems, including recommendation, personalization, and ads optimization. Oversee the design, development, and deployment of scalable ML pipelines: data ingestion, feature engineering, model training, evaluation, and serving. Architect distributed systems to support ML workloads at scale, ensuring reliability, observability, and operational excellence. Partner closely with Product, Engineering, and Content teams to align on business goals and deliver impactful ML-driven experiences. Support best practices in experimentation, evaluation, and ML system monitoring. Ensure cost efficiency, scalability, and performance in ML infrastructure investments. Your Background: 10+ years of industry experience spanning machine learning engineering and distributed systems. 3+ years of leadership and management experience, with a proven ability to build and lead strong t
Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Advanced Engineer, Power System Controls plays a key role in the design, development, and implementation of control software for Hyliion’s Karno Power Module. This position focuses on systems involving combustion, thermal management, pressure regulation, high-voltage, and power management. The engineer will be responsible for developing and validating control algorithms, tuning system parameters, and analyzing data to ensure performance meets engineering specifications. Additional responsibilities include preparing technical documentation, supporting root cause analysis, and ensuring timely, high-quality software delivery. The role requires cross-functional collaboration and occasional travel to support system testing and troubleshooting. Duties and Responsibilities Design, develop, and implement high-quality control software for Hyliion’s Karno Power Module, which includes combustion, thermal, pressure, high-voltage and power management systems. Define and conduct tests to verify software and tune control parameters to meet key performance indicators. Prepare reports and technical documentation related to system performance, control strategies, and compliance. Process and analyze data to verify software against engineering specifications, support root cause analysis and for optimizing performance. Ensure on time delivery with quality. Assist product team in defining customer requirements and generate corresponding engineering specifications. Qualifications Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. Qualifications include: Education, Experience and Cert
About the Team OpenAI develops models that can reason through complex problems and hardware designed for the demands of advanced AI. AI for Chips connects these efforts: applying increasingly capable AI systems to the work of semiconductor engineering. Our goal is to help engineers develop better chips and shorten design cycles. This work brings research, model training, and hardware expertise together to build tools that engineers can use on real designs, with correctness and measurable performance at the center. About the Role We’re hiring a Software Engineer to build the research infrastructure and tooling that help OpenAI models design silicon. You’ll turn chip-design workflows into reliable environments for reinforcement learning and evaluation, and make it easier for researchers to run experiments and iterate on new ideas. You’ll move between software engineering, tool integration, and open research problems. We value strong coding fundamentals, clear technical judgment, and independent execution. Prior chip-design experience is helpful, but you can learn the domain alongside the team’s hardware specialists. In this role, you will: Build and maintain infrastructure for reinforcement learning environments, evaluations, and long-running experiments. Integrate electronic design automation (EDA) tools into workflows for RTL generation, verification, and physical design optimization. Improve experiment reliability, reproducibility, observability, and performance; debug failures across tools, services, and infrastructure. Develop tooling and model harnesses that let researchers test ideas quickly and measure correctness and power, performance, and area (PPA). Collaborate with researchers and engineers to turn successful experiments into reusable systems and training workflows. Own ambiguous projects end to end, communicate progress, and use results to guide the next iteration. You might thrive in this role if you: Have strong software engineering fundamentals, with
About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. Our team builds the software, tooling, and operational systems that help manage this fleet at scale. We work across production engineering, distributed systems, capacity management, and operational automation to improve reliability, reduce manual work, and make better use of available compute. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will develop the systems that help manage the GPU fleet powering ChatGPT, including tooling for fleet health, capacity planning, operational automation, and incident response. You will work closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and compute utilization. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference. Develop systems for capacity planning, fleet health monitoring, and resource utilization. Automate operational workflows, including incident detection, diagnosis, and response. Identify and address bottlenecks affecting fleet reliability, scalability, and performance. Partner with infrastructure, research, and product engineering teams to improve the compute platform. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems. Have a background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering. Have built software that automates operational workflows and reduces manual work. Have worked with distributed infrastructure, cluster orchestration, or large-scale int
Get new performance and systems engineer jobs by email
Daily job updates · Unsubscribe anytime