Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Data Center Engineer , you'll help us scale our Core/Edge Data Centers and hardware infrastructure at a time of incredible growth for our business. At Roblox, you'll have boundless opportunities to shape the future of the Imagination Platform™ and demonstrate your passion for delivering thoughtful solutions in front of a global audience. If you know what it takes to build and operate hardware infrastructure that can sustain millions of concurrent players year-round and you take play as seriously as we do, you'll fit right into our highly experienced and ever-expanding engineering team. You will report to the Technical Lead Data Center Engineer. You will: Develop and maintain the Core/Edge Data Center and hardware infrastructure to meet the large scale and real-time requirements of our Imagination Platform™ to ensure our community has an awesome experience anywhere in the world. This includes all aspects of the server, network infrastructure, power, and environmental life cycles. Own efforts to track and mitigate systemic issues preventing hosts from returning to service. Identify and solve critical problems and prevent them from re-occurring via root cause analysis and giving rec
Jobiba hiring network
Lead Infrastructure Software Engineer Jobs
6,753 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current lead infrastructure software engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Engineering Manager (Player & Coach), you will lead and mentor a team of Forward Deployed Engineers focused on building, scaling, and optimizing LLM inference workloads for Baseten customers. Applying both hands-on technical ownership and managerial leadership, you will guide your team through the processes of designing, deploying, and managing high performance, low latency AI applications on Baseten’s platform. FDE at Baseten is not a sales function – we are a mix of engineering, product, and customer architects who contribute to the core Baseten codebase, drive large portions of our feature roadmap, and execute on complicated customer engagements. You will also partner with product, infrastructure, and other customer engineering teams to ensure that large language models (LLMs) and other generative AI systems deliver best-in-class performance, reliability, and cost efficiency in production environments. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Leadership & Team Management Lead, mentor, and grow a team of Forward Deployed Engineers, providing guidance on technical direction, project execution, and professional deve
About the Role Anyscale is seeking a Senior / Staff Product Manager to lead Ray Data, our scalable data processing library for ML and AI workloads. This is a uniquely challenging role that requires balancing open source growth with commercial differentiation - driving rapid adoption in the open source Ray Data ecosystem while building compelling proprietary features for Anyscale RunTime, our high-performance commercial engine. You'll own the entire Ray Data product roadmap in a competitive landscape, working closely with the engineering team, the field team, enterprise customers, and the open source community. Success requires: Deeply ingraining yourself into the end-user experience to understand the nature of the product and its gaps and tradeoffs Working closely with customers and open source users to draw the subtle line between growth and commercialization Strategic thinking about which parts of the ML/Data lifecycle to focus on, identifying opportunities where our architectural strengths create the most value. Thinking deeply about and clearly articulating the product strategy to stakeholders Key Responsibilities Drive the Ray Data product roadmap - Balance open source Ray Data feature development with Anyscale Runtime commercial differentiation to ensure that both Ray Data becomes the open source standard for AI data processing and Anyscale Runtime remains sufficiently compelling. Drive open source Ray Data adoption - Focus on community growth, developer experience, and ecosystem integrations Market Positioning & Enablement - Work closely with Product Marketing on strategic market positioning, field enablement, and competitive analysis to maintain differentiation. Customer engagement - Drive key customer engagements assisting sales and field engineering teams. Required Qualifications 4+ years of product management experience with technical products Strong technical background in distributed systems, ML infrastructure, or data processing Experience working
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Manager, Data Center Operations, you'll help us scale our Core Data Center and hardware infrastructure at a time of incredible growth for our business. At Roblox, you'll have boundless opportunities to shape the future of the Imagination Platform™ and demonstrate your passion for delivering thoughtful solutions in front of a global audience. If you know what it takes to build and operate hardware infrastructure that can sustain millions of concurrent players year-round and you take play as seriously as we do, you'll fit right into our highly experienced and ever-expanding engineering team. You will report to the Senior Manager of Data Center Operations. This will be a position based in Goodyear, AZ. You will: Develop and maintain the Core Data Center and hardware infrastructure to meet the large-scale and real-time requirements of our Imagination Platform™ to ensure our community has an awesome experience anywhere in the world. This includes all aspects of the server, network infrastructure, power, and environmental monitoring. Lead a growing team of data center engineers focusing on rack deployments, hardware troubleshooting and break-fix, and decommissioning. Identify and solve criti
Datadog is the global leader in observability and security for cloud applications. Our SaaS platform integrates infrastructure monitoring, application performance monitoring, log management, digital experience monitoring, cloud security, and AI-powered observability to help organisations accelerate digital transformation and operate at scale. As organizations across Japan continue to modernize their technology stacks and embrace cloud-native architectures, Datadog is uniquely positioned to help customers achieve greater visibility, reliability, security, and business performance. The Opportunity Following significant investment in Japan by Datadog, we’re now seeking a senior executive to lead the next phase of growth and scale the business rapidly. This executive will be responsible for defining and executing Datadog's strategy in Japan, driving sustainable, predictable revenue growth, strengthening customer and partner relationships, developing high-performing teams, and firmly establishing Datadog as the market leader for Japan. Reporting to senior APJ leadership, the Country Leader will serve as the executive face of Datadog in Japan and will work cross-functionally with Sales, Customer Success, Marketing, Partnerships, Solutions Engineering, Product, and Corporate functions to accelerate market penetration and customer success. At Datadog, we work hard to maintain an excellent culture of trust, high standards, and great outcomes for customers, partners and employees About the Role Datadog is seeking an experienced and respected country executive to serve as Country Manager, Japan. As the most senior leader in the market, you will be responsible for owning Japan’s overall business health across acquisition, usage, retention, ecosystem, culture, and cross-functional execution. The Country Manager owns the country plan and operating cadence, ensuring effective direct management of all three sales segments (Enterprise, Mid Market, Commercial); and cross-
Workplace Strategy and Operations Manager Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe's Workplace team is responsible for delivering welcoming, functional, and beautiful environments worldwide. We ensure that Stripe's real estate portfolio evolves to stay aligned with our mission and the needs of our users, while also ensuring efficient space utilization. As part of the team, you'll help scale Stripe's future workspaces, collaborating with stakeholders across the company. What you'll do Reporting to the Head of Workplace, the Strategy and Operations Manager will play a key role in driving cross-functional initiatives across our core Workplace pillars (Real Estate Strategy and Transactions, Environments (design and construction), and Experience (workplace services and programs), alongside high-priority 'interstitial' projects. You'll contribute to projects and programs that shape the future of Workplace at Stripe. With the ability to influence from a strategic and administrative perspective, you'll be instrumental in executing key cross-org initiatives and helping us move from strategy to execution while ensuring operational excellence. Responsibilities Assist the Head of Workplace in defining SOKRs and ensuring all pillars and cross-functional partners are aligned with the team's roadmap Lead cross-functional projects that don't fit neatly into one department. Support M&A integrations Identify bottlenecks in our current workflows and design scalable processes to remove them
About the Team OpenAI’s Stargate and 3P Engineering teams are responsible for building and scaling the external infrastructure ecosystem that powers advanced AI systems. We work across hyperscalers, colocation providers, cloud partners, and strategic third-party operators to turn contracted capacity into production-ready compute. Our scope spans the full lifecycle of external deployments: commercial alignment, technical readiness, network integration, hardware enablement, operational readiness, and long-range scaling strategy. As OpenAI’s infrastructure footprint expands globally, we need leaders who can convert complex partner environments into reliable, high-velocity capacity for training and inference workloads. About the Role We are seeking a Technical Program Manager, Token-as-a-Service (TaaS) to lead delivery of external compute capacity that directly serves OpenAI model workloads. In this role, you will own complex cross-functional programs that transform third-party infrastructure into usable tokens at scale. You will partner across engineering, capacity planning, networking, hardware, finance, product, and external providers to ensure that deployed capacity translates into real production throughput. This role sits at the intersection of infrastructure execution, systems readiness, and business impact. Success requires strong technical fluency, elite program management, and the ability to drive accountability across internal teams and external partners. This is a high-visibility role with direct impact on OpenAI’s ability to scale model training and inference globally. This role is based in San Francisco, CA, with a hybrid work model of 3 days in office per week. Relocation assistance is available. Key Responsibilities Lead end-to-end delivery programs that convert external infrastructure capacity into production-ready token supply. Own readiness across compute, storage, networking, security, and operational dependencies for third-party environments. Build
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Opportunity The Global Demand Center is building a signal-led demand engine — one that uses account, audience, behavioral, intent, and market signals to identify where we should focus and how we should activate. The GrowthX team, within the Global Demand Center, is taking the lead on this initiative and looking for a new Signal & Growth Insights Manager. The Signal & Growth Insights Manager will help interpret signal, audience, account, campaign, and growth-motion data to surface insights, measure performance, and recommend opportunities for action. This role will support GrowthX strategy by helping the team understand what signals are emerging, which audiences are responding, which motions are working, and where we should refine targeting, activation, or follow-up. The ideal candidate combines analytical capability, B2B demand understanding, curiosity about modern GTM platforms, and strong business judgment. They can build reports, interpret patterns, surface insights, and make practical recommendations that help GrowthX and its partners improve performance. What You’ll Do Interpret signal and audience performance Analyze outputs from GrowthX tools and data sources to identify patterns across accounts, audiences, segments, regions, products, and growth motions. Use platform data and reporting from tools such as Common Room, Clay, 6sense, Salesforce, Marketo, and related GTM systems to understand account and audience behavior. Surface insi
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta is seeking a world-class deal strategy leader to partner with the APJ GTM teams to grow the business and close the best deals possible. The APJ Deal Strategy lead will set the strategy and direction with APJ GTM leadership to help Okta scale to $4B+, driving innovative approaches and operational excellence to help the field win. This role requires strong cross functional executive collaboration and influence along with the ability to deal with ambiguity, come up with creative solutions and drive urgency with quarterly business cadence while transforming for scale. This leader will have team management responsibilities with the opportunity to further build out a team over time to drive scale in deal coverage and business impact. The person in this role will work across APJ GTM leadership and cross-functional support teams such as Legal and Finance to determine the best deal structures as we ramp up the business. This leader will drive and maintain the deal strategy and operational excellence in alignment with global best practices. In addition, this leader will work with the broader business planning team including P&P, Business Value and Compete along with IT to support successful rollout of new offers and programs to market. As Okta rapidly launches new innovations and expands into new markets, this role is critical to our success. A successful leader in this role will have a combination of skills across deal strategy & management,
NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data
About the team Our Automation team is the engine room of Meesho's supply chain — the ones who show up when packages need to move faster, smarter, and more reliably than the day before. From sortation systems that untangle millions of parcels a day to the material handling tech that keeps our network humming without missing a beat, we turn operational chaos into effortless scale. We're not just automating tasks — we're engineering the backbone that lets Meesho suply chain chart the future of commerce for Bharat About the role We are seeking an experienced Project Manager to lead the end-to-end execution of a new automated warehouse sorting facility. The role involves managing cross-functional vendors across automation machinery, infrastructure, and operations to deliver the project on time, within budget, and to specification. The Project Manager will serve as the single point of accountability from inception to commissioning of the facility. What you will need 6–10 years of experience in project management in warehouse/industrial automation, logistics infrastructure, or large-scale industrial projects. Proven track record of managing multi-disciplinary teams (automation, infra, operations).Experience with vendors/OEMs in material handling, conveyors, robotics, or sortation systems. Excellent vendor management, negotiation, and contract handling. Strong problem-solving, risk management, and execution mindset. Excellent communication and stakeholder management. What you will do Project Planning & Execution - Define project scope, timelines, and milestones in line with business goals. Develop detailed project plans covering automation systems, infrastructure, IT integration, electrical power, and operat
About the Team OpenAI’s Legal team plays a crucial role in advancing our mission by tackling innovative and fundamental legal issues in AI. The team includes professionals from diverse legal fields—technology, AI, infrastructure, privacy, IP, corporate, employment, tax, regulatory, and litigation—who collaborate closely with colleagues across the company. If you are passionate about being a technology lawyer working on cutting-edge challenges, you’ll thrive here. About the Role We are seeking an experienced attorney to serve as the commercial legal lead for Marketing. Based in San Francisco, you will be a primary legal partner to OpenAI’s rapidly growing global Marketing organization, supporting high-impact campaigns, creative production, talent and creator relationships, sponsorships, events, and the agreements and rights that make that work possible. You will work closely with Marketing, Communications, Partnerships, Procurement, Finance, Product, and colleagues across Legal, including Product, Privacy, Regulatory, and IP/Brand, to deliver practical advice at the pace of the business. This is a unique opportunity to help shape OpenAI’s Marketing efforts, negotiate sophisticated transactions, and build scalable legal frameworks for responsible global growth. We operate on a hybrid work model of three days per week in the office and offer relocation support for new employees. In this role, you will: Serve as the commercial legal lead for Marketing, partnering closely with Brand, Creative, Product Marketing, Design, Film and Photo, Performance Marketing, Partner Marketing, Communications, and regional teams from concept through launch. Draft and negotiate a wide range of marketing and entertainment agreements, including agency, production, talent and creator, sponsorship, event, media, content-licensing, marketing-technology, and vendor agreements. Structure and clear the rights needed for campaigns and content, including talent and appearance releases, publicity and
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... We’re in an unbelievably exciting area of tech, fundamentally reshaping the cloud-native, modern virtualization, and AI infrastructure landscape. Here, you’ll lead with innovative thinking, grow alongside us, and work with some of the smartest minds in the industry. We’re looking for engineers passionate about system testing, distributed systems, Kubernetes, storage, modern virtualization, AI workloads, and automation. You’ll work on complex, real-world scenarios involving HA, resiliency, disaster recovery, scalability, and failure testing across large-scale Kubernetes environments. As enterprises modernize their infrastructure, containers, virtual machines, and AI/ML workloads are increasingly converging on Kubernetes. From traditional enterprise applications and VMs to GPU-accelerated AI training, inference, and data-intensive workloads, Kubernetes is rapidly becoming the common platform for running the next generation of applications. This role gives you the opportunity to test and influence how Portworx delivers enterprise-grade storage, data protection, availability, and resiliency across these workloads—including containerized applications, KubeVirt/OpenShift Virtualization VMs, and demanding AI/ML workloads running on Kubernetes. WHAT YOU WILL DO: Own system-level quality for Portworx Enterprise across Kubernetes, storage, and modern virtualization environments. Develop comprehens
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Engineering Manager for Drive Qualification, you will lead a high-performing Bangalore team dedicated to validating Everpure-developed SSDs across performance, reliability, and firmware maturity. In this impactful leadership position, you will own the end-to-end validation strategy for hyperscale and datastore programs, establishing a center of excellence for system-level robustness. Partnering closely with cross-functional firmware, hardware, and analytics teams globally, your mission is to deliver comprehensive qualification coverage that ensures our enterprise storage platforms launch with ultimate confidence and quality. WHAT YOU’LL DO Lead and Scale the Team: Coach, mentor, and grow a multi-level validation engineering team, building a culture of ownership, clear domain expertise, and continuous career development. Drive Validation Strategy: Own the execution roadmap for core qualification domains (including PCIe, NVMe/OCP, and power-loss robustness) across critical milestone gates from engineering samples to final product release. Foster Cross-Functional Alignment: Partner with global hardware, firmware, and program management stakeholders to align on test coverage, coordinate issue triage, and deliver clear risk assessments that dictate release readiness. Advance Automation and Infrastructure: Champion the expansion of Python and Linux-based automation frameworks, regression infrastructure, and data
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. SHOULD YOU ACCEPT THIS CHALLENGE... We’re in an unbelievably exciting area of tech, fundamentally reshaping the cloud-native, modern virtualization, and AI infrastructure landscape. Here, you’ll lead with innovative thinking, grow alongside us, and work with some of the smartest minds in the industry. We’re looking for engineers passionate about system testing, distributed systems, Kubernetes, storage, modern virtualization, AI workloads, and automation. You’ll work on complex, real-world scenarios involving HA, resiliency, disaster recovery, scalability, and failure testing across large-scale Kubernetes environments. As enterprises modernize their infrastructure, containers, virtual machines, and AI/ML workloads are increasingly converging on Kubernetes. From traditional enterprise applications and VMs to GPU-accelerated AI training, inference, and data-intensive workloads, Kubernetes is rapidly becoming the common platform for running the next generation of applications. This role gives you the opportunity to test and influence how Portworx delivers enterprise-grade storage, data protection, availability, and resiliency across these workloads—including containerized applications, KubeVirt/OpenShift Virtualization VMs, and demanding AI/ML workloads running on Kubernetes. WHAT YOU WILL DO: Own system-level quality for Portworx Enterprise across Kubernetes, storage, and modern virtualization environments. Develop comprehens
Get new lead infrastructure software engineer jobs by email
Daily job updates · Unsubscribe anytime