Senior Product Manager, Cluster Scalability Location: Dublin About the Role MongoDB is hiring a Senior Product Manager to own scalability for the Atlas data platform. Atlas is how the world’s most demanding applications run on MongoDB — from real-time gaming and financial transactions to AI workloads with unpredictable traffic patterns. When those applications grow, slow down, spike, or contract, scalability is the layer that decides whether the customer trusts us to keep running their business. This role owns the end-to-end product experience for how customers scale up, scale down, and horizontally scale their Atlas deployments. That includes vertical scaling (instance tier changes), autoscaling, horizontal scaling (sharding and shard key strategy), and the operational workflows around all of them. You will set the strategy, drive execution with engineering, and be accountable for measurable year-over-year improvements in how reliably and predictably customers scale. We care less about whether you’ve worked on databases before and more about whether you’re the kind of PM who can ramp fast on a hard technical domain, form a clear opinion about what needs to happen, and defend that opinion to senior engineers, field partners, and enterprise CTOs, especially when they disagree with you. Database and scaling concepts are teachable. Judgment, conviction, and crispness are not. Responsibilities Set a clear, opinionated vision and strategy for Atlas cluster scalability and translate it into concrete, prioritized bets each release. Own the end-to-end roadmap for vertical scaling, autoscaling, and horizontal scaling, including the operational and observability surfaces customers rely on during scaling events. Make hard prioritization calls, including saying no to things with real customer or leadership support, and own those decisions and their consequences. Build deep technical understandin
Jobiba hiring network
Workload Porting And Performance Engineer Jobs
751 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current workload porting and performance engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
MongoDB is helping define the next era of application development as organizations modernize legacy systems, build AI-native experiences, and power mission critical workloads across cloud, hybrid, and on-premises environments. From MongoDB Atlas, our AI ready developer data platform, to our industry leading server technology and enterprise software offerings, MongoDB provides a flexible and unified software stack that helps customers build faster, scale with confidence, and support a wide range of modern application needs. MongoDB’s standing in the market reflects that momentum: the company continues to be recognized as a leader in cloud database management while also serving organizations that require the performance, flexibility, and control of self-managed deployments. For candidates, that means the opportunity to join a company with strong market relevance, a builder-first culture, and a clear strategic focus on shaping the future of modern, intelligent applications wherever they run. We’re looking to speak to candidates based in Palo Alto, California for our hybrid working model. Life in MongoDB Technical Services As part of the Support organization, you will be directly interfacing with our largest customers and their most difficult problems. There will almost certainly be sweat. But don’t worry – you’ll have time to prepare before you take on your first escalation. The limits of your understanding will constantly be pushed, and you’ll be challenged to grow. Underlying it all is the fuel that drives our team: customer obsession. We keep mission-critical deployments online, recover from high-pressure incidents with grace and humility, become trusted advisors, and work so closely with our customers that they would swear we were part of their team. These are just a few of the things you can expect to do as a Technical Services Engineer: Become well-versed in core aspects of MongoDB Gain deeper expertise in specialized areas of the product Combine technical
We’re looking for a Software Engineer 3 to help bring Voyage’s embedding models - used for semantic search, retrieval, and AI-native experiences; to the platforms and environments where customers already run their workloads, beyond first-party MongoDB Atlas. You’ll join the broader Search and AI Platform organization and collaborate closely with the engineers building Voyage’s first-party inference. Together, we’re extending that platform across cloud marketplaces, third-party inference providers, and self-managed deployments so customers get the same Voyage models, behaving consistently, wherever they choose to run them. As a Software Engineer 3, you'll focus on building the systems, tooling, and deployment workflows that power third-party model delivery. You'll own key components of how Voyage models are packaged, validated, and deployed, work across teams to ensure tight integration with the core inference platform, and contribute to delivery surfaces designed for reliability, observability, and ease of use. We are looking to speak to candidates who are based in Sydney for our hybrid working model. What you'll do Port and tune the model server that runs Voyage embedding and reranking models: improving inference performance, consistency, and runtime behavior across environments Productionize new Voyage models for delivery beyond first-party Atlas, owning the packaging, configuration, and deployment workflows that get them running on AWS, Azure, GCP and more Design correctness, correlation, and performance validation that proves third-party deployments match first-party behavior Build operability into every surface: structured logging, metrics, diagnostics, and health checks with tools like Prometheus and OpenTelemetry Debug problems that span model servers, containers, deployment configuration, and partner cloud environments Work alongside Voyage's model-serving teams, and partner with GTM, SAs, TSEs, and strategic customers on the hardest external deployments Who
The Opportunity We are searching for an ambitious Associate Commercial Growth Account Executive. Your role will primarily be identifying new workloads to expand MongoDB's usage into accounts that are already leveraging our Database as a Service (DBaaS), Atlas, in a self-serve manner. You'll focus on helping customers understand the continuous evolution of our developer data platform and its additional features that could be leveraged for their specific use cases. You will support customers at any stage of their journey, fostering greater adoption, retention, and overall satisfaction. This role provides an excellent opportunity for growth in a highly dynamic sales team. This role will be based in our Dublin or Cork office on our hybrid working model. Day to Day Identify new workloads and potential use cases within existing accounts to drive increased Atlas consumption via your own pipeline generation and outbound activities Articulate the value of MongoDB's expanded feature set and evolving developer data platform Support customers throughout their journey, addressing any queries and ensuring optimal product usage Develop relationships within customer accounts to promote greater visibility and deepen the partnership between MongoDB and your account's decision-makers Strategically navigate diverse industries with a data-driven discovery and sales process to help land and expand customer workloads What you bring to the table At least 1 years of Sales experience, with a preference for database or software Consistent record of meeting or exceeding targets in a highly competitive environment (top 10% performer) Dynamic, enthusiastic, tenacious, team-oriented mindset Excellent verbal and written communication skills Experience working within a quota and commission structure College degree or equivalent work experience Native Arabic Speaker High EQ, coachability, intellectual curiosity, and motivation to build a career at MongoDB Things we love Familiarity with database, de
The Opportunity We are searching for an ambitious Commercial Growth Account Executive. Your role will primarily be identifying new workloads to expand MongoDB's usage into accounts that are already leveraging our Database as a Service (DBaaS), Atlas, in a self-serve manner. You'll focus on helping customers understand the continuous evolution of our developer data platform and its additional features that could be leveraged for their specific use cases. You will support customers at any stage of their journey, fostering greater adoption, retention, and overall satisfaction. This role provides an excellent opportunity for growth in a highly dynamic sales team. We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Day to Day Identify new workloads and potential use cases within existing accounts to drive increased Atlas consumption via your own pipeline generation and outbound activities Articulate the value of MongoDB's expanded feature set and evolving developer data platform Support customers throughout their journey, addressing any queries and ensuring optimal product usage Develop relationships within customer accounts to promote greater visibility and deepen the partnership between MongoDB and your account's decision-makers Strategically navigate diverse industries with a data-driven discovery and sales process to help land and expand customer workloads What you bring to the table At least 1-2 years of Sales experience, with a preference for database or software Consistent record of meeting or exceeding targets in a highly competitive environment (top 10% performer) Dynamic, enthusiastic, tenacious, team-oriented mindset Excellent verbal and written communication skills Experience working within a quota and commission structure College degree or equivalent work experience High EQ, coachability, intellectual curiosity, and motivation to build a career at MongoDB Native German speaking Things we love Familiarity with d
The Inside Account Executive role focuses exclusively on formulating and executing a sales strategy within MongoDB’s most strategic accounts, closing net new workloads and expanding MongoDB’s footprint. What you will be doing Proactively identify, qualify and close a sales pipeline within MongoDB’s most Strategic Accounts Drive product adoption in LOPs/BUs you own in the account, upsells / cross sells by cultivating strategic relationships with executives and multi-level champions, aligning to their key initiatives and long term goals. Collaborate cross-functionally with Customer Success, Professional Services, Marketing, Product, and the internal sales ecosystem to drive customer adoption and satisfaction Meet and exceed quarterly quotas on NWLs, NARR, and PS. Invest in your self-development, focusing on the skills and attributes that will make you successful in your core role and set you up for future success What you will bring to the table Min. 3 year B2B sales experience in a quota carrying closing role, or Strategic Account-based selling experience Experience selling complex Saas and/or Cloud products/services. A proven track record of overachievement through generating your own pipeline and hitting sales targets. Energetic, upbeat, entrepreneurial, tenacious teammate. Possess a strong desire to be successful. Ability to articulate the business value of complex enterprise technology both in verbal and written forms. Passionate about growing your career in the largest and fastest growing market in software (database) through constant development of sales and tech skills What we will bring to the table Opportunity to work with and learn from MongoDB’s most talented and experienced Account Executives. Internal mentor and buddy program cross-departmentally Clear and defined career path to EAE Uncapped commissions’ About MongoDB MongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined
The Inside Account Executive role focuses exclusively on formulating and executing a sales strategy within MongoDB’s most strategic accounts, closing net new workloads and expanding MongoDB’s footprint. We're looking to speak with candidates based in Toronto for our hybrid working model. What you will be doing Proactively identify, qualify and close a sales pipeline within MongoDB’s most Strategic Accounts Drive product adoption in LOPs/BUs you own in the account, upsells / cross sells by cultivating strategic relationships with executives and multi-level champions, aligning to their key initiatives and long term goals Collaborate cross-functionally with Customer Success, Professional Services, Marketing, Product, and the internal sales ecosystem to drive customer adoption and satisfaction Meet and exceed quarterly quotas on NWLs, NARR, and PS Invest in your self-development, focusing on the skills and attributes that will make you successful in your core role and set you up for future success What you will bring to the table Min. 1 year B2B sales experience in a quota carrying closing role, or Strategic Account-based selling experience Experience selling complex Saas and/or Cloud products/services A proven track record of overachievement through generating your own pipeline and hitting sales targets Energetic, upbeat, entrepreneurial, tenacious teammate. Possess a strong desire to be successful Ability to articulate the business value of complex enterprise technology both in verbal and written forms Passionate about growing your career in the largest and fastest growing market in software (database) through constant development of sales and tech skills What we will bring to the table Opportunity to work with and learn from MongoDB’s most talented and experienced Account Executives. Internal mentor and buddy program cross-departmentally Clear and defined career path to EAE Uncapped commissions’ About MongoDB MongoDB is built for change, empowering our customers and
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation
About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
About the Team Training Runtime designs the core distributed runtime that powers everything from early research experiments to frontier-scale model runs. We work on building robust, scalable, high performance components to support our distributed training workloads. Our priorities are to maximize the productivity of our researchers and our hardware, with the goal of accelerating progress towards AGI. Within Training Runtime, the Process Management team develops the distributed OS responsible for launching, coordinating, and supervising the large numbers of processes that make up modern training workloads. Our runtime sits beneath training frameworks and on top of research infrastructure, ensuring jobs run reliably across massive clusters while maintaining performance, stability, and observability. Success for us is measured by both system reliability and researcher velocity - enabling ideas to scale from experiments to production training runs. About the Role As a Training Runtime: Process Management Engineer , you will work on the software that ties thousands of computers together and exposes them as a unified system. This system has to serve individual researchers running multiple parallel experiments, as well as our largest training runs spanning 100’s of thousands and even millions of machines and accelerators. This requires easy to use, introspectable systems that can promote a fast debugging and development cycle, as well as relentless optimization for scale while maintaining stability and performance throughout. You will work primarily in Rust , building high-performance asynchronous systems with a strong emphasis on performance, correctness, and scalability. Working at this scale and at the frontier of AI development poses novel challenges. Out-of-the-box approaches often don’t work. The problems you will be working on are highly ambiguous and require strong design judgment as well as proficient execution to advance the state of our infrastructure. We’re loo
About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling, and infrastructure that help improve model performance across OpenAI’s training and inference workloads at frontier scale. About the Role We’re looking for an autonomous, high-ownership developer productivity engineer who cares deeply about helping other engineers move faster, safer, and with more confidence. This role will sit within OpenAI’s Model Performance organization, contributing to developer infrastructure, CI systems, testing workflows, tooling, and broader performance infrastructure efforts. There is also a strong opportunity to contribute to the Triton project and help improve the systems that support performance-critical engineering work across OpenAI. In this role you will: Improve development workflows for engineers working on model performance infrastructure Design and improve CI/CD, release, validation, and testing pipelines Build and maintain tools that improve reliability, iteration speed, and engineering confidence Partner closely with engineers to identify friction in testing, debugging, deployment, and development workflows Contribute to infrastructure efforts that support performance-critical training and inference systems Help improve developer experience across Python-heavy codebases and performance-oriented infrastructure Work in a high-context, ambiguous environment where ownership and good judgment matter You might thrive in this role if: You are motivated by enabling the people around you and helping engineers do their best work You have strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build/release workflows You are highly collaborative, empathetic, and comfortable partnering deeply with technical teams You are strong in Python and enjoy building reliable, scalable developer tools and infrastructure You have experience improving large-scale engineering work
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary The Senior Data Engineer will be responsible for delivering high quality modern data solutions through collaboration with our engineering, analysts, data scientist, and product teams in a fast-paced, agile environment leveraging cutting-edge technology to reimagine how Healthcare is provided. You will be instrumental in designing, integrating, and implementing solutions on-premise as well supporting migrations of existing workloads to the cloud. The Senior Data Engineer is expected to have extensive knowledge of modern programming languages, designing and developing data solutions. The position is open in a data engineering team that is responsible for processing payer files into our Data Warehouse. Required Qualifications 6+ years of experience working with SQL and relational database management systems 3+ years of experience in Cloud Data Engineering Platforms such as AWS, GCP, Azure, Databricks, Snowflake etc. 3+ years of experience in on-prem Data Engineering Platforms such as Microsoft SQL Server, Oracle, Teradata etc. Programming and modifying code in languages like SQL, Python, and PySpark to support and implement Cloud based and on-prem data warehousing services.</spa
Get new workload porting and performance engineer jobs by email
Daily job updates · Unsubscribe anytime