Jobiba hiring network

Workload Porting And Performance Engineer Jobs

751 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current workload porting and performance engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

M
Mongodb
📍 Chicago• Full-time• From $60K/yr
1mo ago

MongoDB’s mission is to empower innovators to create, transform, and disrupt industries by unleashing the power of software and data. We enable organizations of all sizes to easily build, scale, and run modern applications by helping them modernize legacy workloads, embrace innovation, and unleash AI. Our industry-leading developer data platform, MongoDB Atlas, is the only globally distributed, multi-cloud database and is available in more than 115 regions across AWS, Google Cloud, and Microsoft Azure. Atlas allows customers to build anywhere—on the edge, on premises, or across cloud providers. The Inside Account Executive role focuses exclusively on formulating and executing a sales strategy within MongoDB’s most strategic accounts, closing net new workloads and expanding MongoDB’s footprint. We are looking to speak to candidates who are based in Chicago for our hybrid working model. What you will be doing Proactively identify, qualify and close a sales pipeline within MongoDB’s most Strategic Accounts Drive product adoption in LOPs/BUs you own in the account, upsells / cross sells by cultivating strategic relationships with executives and multi-level champions, aligning to their key initiatives and long term goals Collaborate cross-functionally with Customer Success, Professional Services, Marketing, Product, and the internal sales ecosystem to drive customer adoption and satisfaction Meet and exceed quarterly quotas on NWLs, NARR, and PS Invest in your self-development, focusing on the skills and attributes that will make you successful in your core role and set you up for future success What you will bring to the table Min. 1 year B2B sales experience in a quota carrying closing role, or Strategic Account-based selling experience Experience selling complex Saas and/or Cloud products/services A proven track record of overachievement through generating your own pipeline and hitting sales targets Energetic, upbeat, entrepreneurial, tenacious teammate. Possess a str

mongodbawsazure
View job →
M
Mongodb
📍 New York City• Full-time• From $157K/yr
1mo ago

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in New York City for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workf

mongodbawsazure
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Cork for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workflows Pref

mongodbawsazure
View job →

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the Role + Team The Security Engineering team builds the identity and access control plane governing how employees, services, and agentic workloads access Robinhood’s critical infrastructure. We are a platform-building engineering team focused on creating secure, frictionless, and automated self-service access systems at scale. As a Staff Software Engineer on this team, you will shape Robinhood’s foundational identity architecture, define technical strategy, and build Tier-0 access infrastructure used across the entire company! From solving complex authentication and authorization challenges to setting secure guardrails for emerging AI and agentic workflows, your technical leadership will directly impact every product team at Robinhood! What You'll Do Access Control Plane Architecture: Define and build the core technical architecture for managing employee, service, non-human, and agentic access across company-wide resources Tier-0 Platform Engineering: Design, build, and operate highly available, resilient, and observable backend identity infrastructure using Go, Python, Rust, or comparable systems languages. AuthN & AuthZ Solutions: Implement modern identity protocols and access frameworks (OAuth 2.0, OpenID Connect, RBAC, ABAC, and policy-based access controls). AI & Agentic Access Governance: Develop secure access patterns, identity frameworks, governance, and observability guardrails for AI agents, LLM applications, and Model Context Protocol (MCP) systems. Self-Service Developer Experience: Create automated, developer-friendly platforms that allow teams to request, approve, provision, and audit access without ma

pythonvueaws
View job →
L
Lyft
📍 Toronto• Full-time• From C$172K/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. As the Engineering Manager for the Lakehouse Foundation team, you will lead a group of engineers responsible for the foundational data layer that all of Lyft's data systems and emerging AI workloads are built on. The team owns catalog and metadata management, table formats and storage, and the access patterns and gateways through which other engineering teams interact with Lyft's data. As Lyft converges on a unified lakehouse architecture, this team builds and operates the single source of truth that powers analytics, machine learning, experimentation, and every business decision made from data. You will play a key role in shaping the team's technical direction, partnering with peer Data Platform teams on a multi-year platform evolution, and developing engineers who operate with autonomy on systems of significant scale and complexity. Lyft's Infrastructure teams build the foundational systems that the rest of engineering depends on to move fast, ship reliably, and scale efficiently. These are high-leverage roles where the work you and your team do has a multiplicative effect across the company. We're looking for experienced leaders who can balance the discipline of operating critical infrastructure with the curiosity to keep evolving how Lyft builds. Engineering at Lyft is a place where managers and engineers operate with high ownership and strong technical judgment. Our engineers expect their managers to be honest, available, and focused on the work that matters: developing their teams, removing obstacles, and giving people the support they need to do their best work. We build teams that are inclusive, technically rigorous, and have a strong sense of ownership for what they build. Responsibilities: Lead a team responsible for Lyft's foundational data layer, including catalog and metadata management,

restmachine learningai
View job →

Integration Reliability Engineer, Technical Operations About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team APAC Payins TechOps is a newly formed team based in Singapore. Our charter is to improve the health and resilience of our payment acquiring volume, making it easier for Stripe to build and operate the integrations that support hundreds of billions of dollars in payments annually. Our team partners closely with payments and platform engineering teams to assess systems or processes that create high operational workloads, then develops durable solutions via code, tooling, data, and process improvements. We are responsible for the financial data quality of key systems at Stripe, ensuring that data is quarantined without impacting downstream systems while implementing the right changes upstream to prevent recurring issues. The team's work has a direct impact on Stripe's ability to expand into new markets and offer more sophisticated payment features to merchants. What you’ll do Responsibilities Scope and lead technical initiatives end-to-end: identify problems worth solving, propose the right solution approach, and deliver on that solution — not just execute on a pre-defined plan. Investigate problems in systems by tracing problems through Stripe’s stack. You’ll examine code, write SQL queries, read logs, and inspect data pipelines to understand system behavior, then make changes to address. Examine updates being made by Stripe’s financial partners to understand impact, and make the changes within Strip

sqlgitrest
View job →

About the Team The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure. Key Responsibilities Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. Develop operationa

awsazuregcp
View job →
O
1mo ago

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

pythonawsazure
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s API Multicloud team is responsible for extending OpenAI’s API platform into strategic cloud environments, starting with AWS . The team’s mission is to distribute OpenAI’s API broadly and safely by enabling key API technologies in cloud-native environments, in close partnership with Amazon and internal teams across Codex, Research, Safety Systems, and Applied. The team is focused on bringing core developer and enterprise capabilities into cloud-native environments, including cloud-hosted Codex, model customization / post-training as a service, and new stateful runtime environments for agentic workloads. This work sits at the intersection of production ML systems, developer platforms, model behavior, and large-scale infrastructure. About the Role We’re looking for a backend engineer who can quickly understand OpenAI’s models, products, and systems, then adapt first-party deployments for other cloud platforms. You’ll build backend services, APIs, SDK integrations, authentication flows, and cloud service infrastructure that let developers use OpenAI capabilities in the cloud environments where they already build. This role involves working across teams, sometimes embedded with partner product groups, to ship products quickly and across multiple platforms at the same time. It’s a strong fit for engineers who have built developer tools, especially AI-powered tools, communicate clearly across technical boundaries, and can shape architectures that support different deployment models; experience building cloud services is a strong plus. In this role, you will: Build backend and infrastructure systems that extend OpenAI’s API platform into cloud-native environments, like AWS. Design and ship cloud-contained products that allow customers to use OpenAI capabilities while keeping workloads and data within cloud environments. Help stand up cloud-hosted Codex experiences powered by the OpenAI Responses API. Build the infrastructure and runtime abstractions

typescriptpythonaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Cooperative AI team is scaling to devices and embedded operations and user experiences. Our model-powered scaled workforce and knowledge system are moving on to the edge and powering our devices and edge experiences. By leveraging OpenAI’s state-of-the-art models and technologies, in production and in the lab, we develop systems that reason and work autonomously with customers and with our workforce responsible for operational work. We carry real workloads for critical systems across finance, sales, customer support, integrity, product insights, internal operations, and now devices to drive insights into product and industry. We partner closely with internal teams and external customers globally, operating in a hyper-fast feedback loop where many of our users are just a few steps away. This proximity allows us to iterate quickly, validate impact in real time, and accelerate industry impacting learnings and systems builds. We are a highly multidisciplinary, self-contained team focused on transforming the workplace via smart systems, knowledge, scalable and reliable primitives that apply world-class AI capabilities across domains. Our mission is to learn fast and transform how humans collaborate with AI at scale. About the Role We are looking for a Technical Lead Manager to lead a team of engineers building AI-native embedded experiences and operations-forward systems. In this role, you will perform both hands-on technical leadership and small team management. You will drive business outcomes, architecture and technical strategy for complex systems, contribute directly to implementation, and help grow a high-performing team. You will work closely with internal stakeholders to understand operational challenges, identify high-leverage opportunities for automation, and deliver solutions that create measurable impact. This role is ideal for someone who enjoys moving between technical design, coding, mentoring engineers, and working directly with users t

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and

awsrestai
View job →
O
1mo ago

About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p

pythonawsci/cd
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Infrastructure team is central to this mission, setting the core strategy and implementing the vision. From site selection to deployment to operations, this team sits at the intersection of commercial, technical, and operational domains, interacting with experts and executives inside and outside of OpenAI. We design and operate mission-critical facilities that support cutting-edge AI workloads at scale. About the Role We are seeking a Facilities Operations Lead to support the commissioning, deployment, and long-term operation of our next-generation AI data centers. This role bridges the interface between data center construction and hardware landing, ensuring seamless integration of mission-critical infrastructure with hardware deployment timelines. You will define and execute commissioning plans, support infrastructure bring-up, and take ownership of operations and maintenance for cutting-edge, large-scale, AI data centers. You will collaborate closely with design, construction, and hardware teams to define repeatable processes for new data center builds and lead hands-on operations to uphold the performance and reliability of our deployed infrastructure. Key Responsibilities Define and execute sequences of operations, commissioning steps, and bring-up processes for mission-critical data center facilities. Interface with the design and hardware teams to define deployment procedures tailored to each data center and hardware configuration. Oversee installation, commissioning, and operational readiness of large-scale data center campuses. Manage monitoring, maintenance, and quality control of the data center infrastructure, including high-performance liquid cooling systems. Develop on-site operations staffing strategy. Develop and enforce procedures for planed and unplanned downtime and SLAs for critical

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our infrastructure team helps deliver OpenAI’s most capable models and products to the world by scaling infrastructure and turning demand into useful FLOPS. We collaborate across research, engineering, design, and business to turn cutting-edge AI advancements into impactful, real-world applications. Our team ensures the right compute is available—at the right time and place—to support some of the world’s most demanding workloads. We empower all of OpenAI’s products and research by scaling the infrastructure behind them. Our work makes it possible to launch new models and products reliably and at scale. About the Role As a Data Scientist on the Infra team, you will play a key role in shaping how we scale the infrastructure that powers OpenAI’s products and research. This is critical as we operate one of the largest and most advanced compute fleets in the world, supporting millions of users and businesses globally. We focus on aligning infrastructure measurement, planning, scaling, allocation, and efficiency to drive measurable impact across the company. You should expect to guide the definition of foundational datasets for infrastructure resources, develop metrics that inform key decisions, build forecasting and optimization models, and establish source of truth dashboards and analyses that enable teams to understand and improve infra usage. Most importantly, you should expect to be a core partner to engineering, research, and product teams in shaping the infrastructure that powers everything OpenAI builds. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build and maintain foundational datasets and metrics that reflect infrastructure usage, efficiency, and scaling. Develop forecasting and optimization models to support infra planning and resource allocation. Partner with engineering, research, and product teams to shape infrast

pythonsqlaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

awslinuxrest
View job →
🔔

Get new workload porting and performance engineer jobs by email

Daily job updates · Unsubscribe anytime