ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a Delivery Director, Capacity programs for our on-premises data center builds and neo cloud (GPU cloud) delivery programs. This is a high-visibility, execution-critical role sitting at the intersection of infrastructure engineering, capacity planning, vendor/partner management, and customer delivery. You will own the end-to-end delivery lifecycle for large-scale compute infrastructure — from initial site/capacity commitments through power, networking, and hardware bring-up, to production-ready GPU/compute capacity landing in the hands of internal teams or customers. You'll be the person who turns ambitious infrastructure roadmaps into predictable, on-time, delivery. RESPONSIBILITIES Own delivery of on-prem infrastructure builds — colocation expansions, power/cooling readiness, rack-and-stack, network fabric bring-up, and hardware acceptance testing — coordinating across colo providers and partners, network engineering, hardware ops, and vendor teams. Drive neo cloud delivery programs — manage capacity delivery from GPU cloud and neo cloud partners (e.g., colocation/bare-metal/GPU cloud providers), including contract milestones, capacity ramps, SLAs, and go-live readiness. Build and maintain master delivery schedules across concurrent, multi-site, multi-vendor programs, integrating power/shell timelines, hardware lead times, logistics, and software/platform readiness into a single critical path.
Jobiba hiring network
Lead Cloud Infrastructure Engineer Jobs
6,876 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current lead cloud infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
We’re hiring a Customer Success Manager to turn new deployments into measurable outcomes for engineering teams using Coder. You guide customers from onboarding through renewal, clearing blockers, and aligning our platform to their goals. You lead rollouts, monitor health and usage, and keep stakeholders informed. You partner with Sales on expansions and renewals, and with Support and Product on issues and feedback. What you’ll do here Own onboarding and rollout plans for new customers; coordinate with Sales for a smooth, successful implementation of Coder’s platform Remove barriers to adoption so customers achieve their desired outcomes expediently Engage with customers to understand goals, challenges, and use cases, and provide tailored guidance and recommendations Identify expansion opportunities within existing accounts, and partner with Sales on upsell and cross‑sell motions Oversee renewals and forecasting, ensuring timely, successful commitments Serve as the primary point of contact for inquiries, issues, and escalations, and collaborate with Coder teams to resolve Monitor and report on customer health and usage Develop deep expertise in Coder’s products to provide global customer coverage What we’re looking for 3+ years in software/SaaS sales and/or customer success, with enterprise experience preferred Working knowledge of cloud infrastructure, DevOps, developer tools, coding agents, platform engineering, CI/CD, and the SDLC Hands-on experience with Salesforce and other industry-standard CS platforms History of building strong relationships across executive, business, and technical stakeholders Consistent internal advocacy for customers and the ability to provide actionable feedback for Product and cross‑functional teams Habit of staying current on industry trends, CS best practices, and the competitive landscape Startup experience High EQ with strong verbal communication and technical writing skills Self‑motivated with a creative and analytical approach to
We are now looking for a dynamic business leader to grow NVIDIA's Host Networking business for AI infrastructure with AI Labs and Hyperscalers! This leader will drive strategic direction, customer engagement, and multi-year growth for networking products such as NVIDIA DPUs, SuperNICs, and their associated software and ecosystem. Success in this role will be measured by the level of adoption and integration of our Host Networking products with our end customers' workflows and workloads. Success is contingent upon building trust with executives, architects, product leaders, and platform teams across NVIDIA and our largest customers. This leader will lead the go-to-market motion, connecting customer AI factory needs to NVIDIA's networking portfolio and aligning product, sales, engineering, architecture, marketing, and partner teams to secure design wins and scale deployments. What you'll be doing: Identify, develop and close strategic design wins for DPU and SuperNIC with top AI labs and Cloud Service Providers! Build and implement the segment sales growth strategy for host networking across hyperscaler and frontier model AI labs building large scale AI infrastructure. Define customer-specific DPU and SuperNIC value propositions and deployment motions, and lead a matrixed team across product, architects, engineering, sales and marketing teams. Promote NVIDIA host networking products externally and internally, positioning their value for AI workloads and other infrastructure products from NVIDIA, in a collection of use-cases in Networking, Security and Storage. Build a robust opportunity pipeline with segment sales and account teams, including account mapping, customer requirements, proof points, executive engagement, and partner alignment. Track and drive quarterly business reporting, forecast accuracy, design-win progress, roadmap asks, and
We’re hiring a Sr. Customer Success Manager to turn new deployments into measurable outcomes for engineering teams using Coder. You guide customers from onboarding through renewal, clearing blockers, and aligning our platform to their goals. You lead rollouts, monitor health and usage, and keep stakeholders informed. You partner with Sales on expansions and renewals, and with Support and Product on issues and feedback. What you’ll do here Own onboarding and rollout plans for new customers; coordinate with Sales for a smooth, successful implementation of Coder’s platform Remove barriers to adoption so customers achieve their desired outcomes expediently Engage with customers to understand goals, challenges, and use cases, and provide tailored guidance and recommendations Identify expansion opportunities within existing accounts, and partner with Sales on upsell and cross‑sell motions Oversee renewals and forecasting, ensuring timely, successful commitments Serve as the primary point of contact for inquiries, issues, and escalations, and collaborate with Coder teams to resolve Monitor and report on customer health and usage Develop deep expertise in Coder’s products to provide global customer coverage What we’re looking for 3+ years in software/SaaS sales and/or customer success, with enterprise experience preferred Working knowledge of cloud infrastructure, DevOps, developer tools, coding agents, platform engineering, CI/CD, and the SDLC Hands-on experience with Salesforce and other industry-standard CS platforms History of building strong relationships across executive, business, and technical stakeholders Consistent internal advocacy for customers and the ability to provide actionable feedback for Product and cross‑functional teams Habit of staying current on industry trends, CS best practices, and the competitive landscape Startup experience High EQ with strong verbal communication and technical writing skills Self‑motivated with a creative and analytical approach
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture. EXAMPLE INITIATIVES The B200 Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's first Blackwell GPU clusters. Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems to move customer workloads seamlessly across regions to optimize cost and latency. Precision GPU Triage: Developing automated Go-based operators to identify, cordon, and repair unhealthy H100 nodes in under an hour. The Supply Chain of Intelligence: Partnering with lead
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking a Sales Manager to help lead a team of Strategic Account Executives. This person will be responsible for building and coaching a team of high-performing sales reps, driving revenue growth, and partnering closely with product and engineering to bring our cutting-edge AI infrastructure to customers. RESPONSIBILITIES Lead and mentor a team of Strategic Account Executives to consistently exceed pipeline and revenue goals. Help define and execution on go-to-market strategy for our fastest growing customer segment. Collaborate cross-functionally with Marketing, Product, and Engineering to align customer needs with Baseten’s product roadmap. Be deeply engaged with the product, enabling reps to have highly technical conversations with prospects and customers. Foster a culture of accountability, learning, and collaboration within the sales team. REQUIREMENTS 6+ years of closing sales experience, with 2+ years in management leading high-performing teams. Strong technical acumen, ideally with background in AI infrastructure, cloud infrastructure, or developer platforms. Comfortable operating in the weeds with technical products and guiding reps through complex deals. Proven track record of success in Strategic/Enterprise sales environments. Based in San Francisco or New York and open to coming in office at least 3 days per week (Tuesday-Thursday). BENEFITS Competitive compensation, including meaningful equ
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking a Sales Manager to help lead a team of Account Executives within our Startups segment. This hire will be responsible for building and coaching a team of high-performing sales reps, driving revenue growth, and partnering closely with product and engineering to bring our cutting-edge AI infrastructure to customers. RESPONSIBILITIES Lead and mentor a team of Startups Account Executives to consistently exceed pipeline and revenue goals. Help define and execution on go-to-market strategy for our fastest growing customer segment. Hire and scale the team by recruiting, interviewing, and onboarding top talent Collaborate cross-functionally with Marketing, Product, and Engineering to align customer needs with Baseten’s product roadmap. Be deeply engaged with the product, enabling reps to have highly technical conversations with prospects and customers. Foster a culture of accountability, learning, and collaboration within the sales team. REQUIREMENTS 4+ years of closing sales experience, with 2+ years in management leading high-performing teams. Strong technical acumen, ideally with background in AI infrastructure, cloud infrastructure, or developer platforms. Comfortable operating in the weeds with technical products and guiding reps through complex deals. Proven track record of success in high velocity sales environments. Based in San Francisco or New York City and open to coming in office at least 3 d
The Documentation team creates engaging and informative technical content, particularly our public product documentation: Datadog Docs . This is an opportunity for a Documentation Manager to help us deliver high quality technical documentation and lead one of our growing documentation teams. Our team is hands-on with the technologies that Datadog monitors, and collaborates with technical and product teams to create technical documentation for our APIs, SDKs, developer and security tools, and community developed integrations — thousands of pages of content! This role is remote in North America. What You’ll Do: Manage 3-5 technical writer direct reports, overseeing their day-to-day, helping them make and manage effective content project plans, formally supporting their growth and development, and fostering an environment that encourages positive engagement and performance. Partner with our Product and Engineering teams to create and maintain public documentation for users, developers, and SREs, helping them succeed with Datadog products, features, SDKs, APIs, and developer tools. Experiment with the product, dig into source code, and interview subject matter experts to research technical details for documentation. Collaborate with our contributor community and greater Documentation team to develop and improve documentation standards and processes. Who You Are: You have 5-10 years experience researching and developing documentation for technical topics, like databases, APIs, cloud infrastructure, and performance and security monitoring. You have 2-5 years experience managing a small team of technical writers, coaching their performance, guiding their career journey, and breaking down complex projects into achievable plans. You have publicly available technical writing samples. You are comfortable reading at least two programming languages (e.g. Ruby, Python, Go, bash). You have written or managed documentation as code, and are familiar with modern infrastructure such a
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! The Opportunity Cohere seeks a Chief Information Security Officer who can help shape Cohere’s security strategy & the broader conversation around securing AI at scale. You know how to build trust across organizations and communicate trade-offs clearly in environments where speed, innovation, and security must coexist. You will build the playbook and lead the evolution of Cohere’s global security program across corporate systems, cloud infrastructure, AI development, and enterprise operations. As a visible leader for Cohere’s security vision internally and externally you will navigate risk, represent Cohere in industry discussions, and reinforcing our commitment to building AI responsibly. In this role you will: Define and Scale Cohere’s Security Strategy: Define and execute Cohere’s overarching information security strategy in alignment with business priorities and long-term company growth. Build a Modern Risk, Governance & Compliance Program: Lead Cohere in identifying, assessing, and mitigating security risk across all business functions. Secure AI Systems and Technical Infrastructure: Lead the security architecture st
About the team Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security that could scale to an extreme level of severity. Our work involves: Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework , and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. About the role As AI agents become more capable at software engineering, and automate more of our internal work, they could become a dangerous cyber threat. People in this role will help OpenAI prepare for security threats from advanced AI agent insiders. In this role, you will: Identify paths by which capable future internal AI agents could compromise OpenAI. Design security controls - focusing on measures with long lead times that benefit from advanced preparation. Stress-test defenses with AI agent evaluations and penetration tests You might thrive in this role if you: Are deeply technical across security and modern infrastructure, and are comfortable digging into the details of operating systems, cloud, containers, CI/CD, or distributed systems. Have strong software engineering skills and enjoy building prototypes yourself. Are interested in engaging with stakeholders and can do so effectively. Bonus: have experience securing cloud infrastructure, and are deeply familiar with core components of the AI stack. Compensation Range: $293K - $405K USD About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy
Datadog’s Cloud Networks team designs, builds, and maintains the production network infrastructure that powers everything built on top of our platform across AWS, GCP, Azure, and beyond. In this role, you’ll set technical direction for how we scale our multi-region, multi-cloud network footprint while keeping reliability and performance high. You’ll partner closely with internal teams and Cloud Service Providers to troubleshoot complex connectivity issues, integrate new networking capabilities, and improve the foundations our engineers and customers rely on. This is a high-impact opportunity to drive meaningful improvements in scale, resiliency, and cost efficiency. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Design, build, and operate cloud network infrastructure across AWS, GCP, Azure, and Neoclouds in a multi-region environment. Own connectivity between clouds, customers, and developers—ensuring scalable, secure, and reliable network paths. Set clear technical direction for expanding data centers and evolving the network while maintaining stability and performance. Improve cross-site and cross-region connectivity patterns to support Datadog’s growing platform needs. Lead deep investigations into latency, packet loss, and connectivity failures – from pcap and path analysis through to escalations with cloud providers that may originate from customer support Identify and deliver network-related efficiency and cost-saving opportunities that positively impact business health. Who You Are: You have deep networking expertise. You understand BGP, route policies, path selection, prefix advertisement, and what breaks in large-scale networking. You have substantial experience designing, building, and evolving large-scale Software-Defined Networks—inclu
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking a strategic and results-driven engineering leader who is passionate about cloud agnostic infrastructure, operational excellence, and enabling engineering teams to operate autonomously and build with confidence. As Head of Infrastructure, you'll lead a talented and geographically distributed team of engineers across the SF Bay Area, India, and Europe, fostering a culture of collaboration, ownership, and continuous improvement. You'll own the infrastructure that underpins one of the world's most widely used API platforms, an environment handling ~80,000 requests per second at the front door, and be responsible for its reliability, scalability, and evolution. In addition to infrastructure, you'll own the Site Reliability Engineering (SRE) function at Postman, setting the standards and practices that keep the platform reliable at scale. You'll work closely with engineering managers, product managers, and platform teams to drive the technical roadmap for our cloud agnostic infrastructure and reliability practices, ensuring we can support a large and rapidly growing engineering organization. If you're p
As the Engineering Manager for Commercial Audit, you will lead a high-performing team responsible for scaling Datadog’s security and compliance posture through automation, tooling, and engineering excellence. Our GRC (Governance, Risk, and Compliance) function is a critical partner to the broader Security and Engineering organizations, ensuring that Datadog not only meets rigorous global regulatory standards but does so in a way that is efficient, scalable, and integrated into our cloud-native infrastructure. You will manage a team of engineers and analysts who are transitioning to a GRC engineering direction to treat compliance as a software problem, leveraging AI, custom tooling, CI/CD pipelines, and cloud-native services to turn complex regulatory requirements into actionable, automated controls. You will lead the strategy, roadmap, and execution of Datadog’s Commercial Audit initiatives. This is a high-impact leadership role where you will grow a team of engineers and analysts responsible for directly maintaining our compliance programs and related audits (e.g., SOC2, PCI, HIPAA, ISO) while looking to improve efficiency and effectiveness through platforms and tooling. You will act as a bridge between technical engineering, legal, and compliance, enabling the organization to move fast while maintaining a secure and compliant environment. You will champion a culture of "compliance-as-code," identifying opportunities to automate evidence collection, streamline control testing, and reduce manual toil for both your team and our partner engineering teams. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the strategy, roadmap, and execution of Datadog’s commercial security compliance efforts, shifting from manual audit processes to automated, scalable
As Engineering Manager for Threat Detection, you will lead a high-performing team that powers Datadog's detection program. Threat Detection is the organization responsible for keeping Datadog ahead of an evolving threat environment: closing coverage gaps faster, raising the bar on signal quality, and shipping detections that hold up under the scale and complexity of cloud-native infrastructure. Your team will combine direct detection expertise, platform engineering, and applied AI to ship detections at a pace and scale traditional rule-writing alone cannot match. Examples of what your team will work on include detection-authoring agents, the detection platform that powers every rule in production, coverage analysis, alert triage and response automation, and the evaluation infrastructure that holds these systems to a high bar of fidelity. Detection authorship is a shared responsibility across the organization, and your team will contribute both by building the systems that scale our authoring capacity and by writing detections directly when their domain expertise is the right tool. You will partner closely with our Security Incident & Response Team (SIRT), Cyber Threat Intelligence (CTI), AI Engineering teams, and Datadog's broader Security organization. This is a high-impact leadership role: you will grow a team of security and software engineers responsible for building and executing our detection and AI strategy. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the strategy, roadmap, and execution of Datadog Security's shift to AI-accelerated detection and response. Drive development of high-fidelity detections as a shared responsibility across the organization, ensuring your team's systems and direct contributions raise the bar on coverage and
Get new lead cloud infrastructure engineer jobs by email
Daily job updates · Unsubscribe anytime