Jobs in United States

Hr Business Partner Director in United States

1,447 active opportunities · Updated October 2026

Explore current hr business partner director jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

NR
📍 Atlanta, Georgia, United States· Full-time
✓ High-confidence listingCompany trend -73.9%

From $98K/yr

Quick readStrong listing-quality and freshness signals

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity At New Relic, we provide our customers with real-time insights, so they can innovate faster. Our software provides deep observability across the stack, enabling software teams to solve their customer’s problems, accelerate digital transformation, and make DevOps work. You will be at the heart of the teams supporting New Relic’s infrastructure and will work on a team that provides global service mesh and load balancing solutions. We provide these services on-premises, as well as using our multi-cloud infrastructure. We support each other to do our best work through positive communication and continuous improvement. What You’ll Do As a key member of our Infrastructure team, you will design and operate a scalable, resilient ingress data plane that directly impacts the value we provide to our customers. By ensuring the stability and performance of our global service mesh and load balancing solutions, you drive the foundational reliability that the entire New Relic organization depends on to deliver real-time insights. You will leverage advanced automation and infrastructure-as-code to accelerate development speed, allowing our engineering teams to ship safe, incremental changes across a massive fleet with confidence. Your work in evolving our DNS and CDN infrastructure is not just about maintenance; it is about creating a seamless, high-performance environment that enables innovation at scale. Through deep collaboration with Product, Design, and partner platform t

PythonAWSAzureKubernetes
T
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -50%

$150K – $200K/yr

Quick readStrong listing-quality and freshness signals

About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! Prior to applying please note: This role is hybrid requiring 2 days in office at our San Francisco hub every Tuesday & Wednesday (located at 130 Sutter St, San Francisco, CA). About the Role: The Architecture Team drives Taskrabbit's Platform Modernization efforts, guiding engineering teams toward our Target State Architecture (TSA). As a Staff Software Architect, you'll serve as an embedded representative of the Architecture team, partnering with engineering teams across the platform to accelerate their modernization work. You'll operate as a subject matter expert on platform modernization within a small team of architects, applying AI-assisted and Spec-Driven Development (SDD) practices to help teams build new components using our Target State Architecture stack of technologies. This role is hands-on and results-oriented: pairing directly with engineers and technical leads, not just advising from the sidelines. You will be: Act

TypeScriptJavaNode.jsKubernetes
T
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -50%

From $170K/yr

Quick readStrong listing-quality and freshness signals

About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! This role is hybrid requiring 2 days in office at our San Francisco or NYC hub every Tuesday & Wednesday. About the Role Machine Learning is a cornerstone at Taskrabbit, and we're looking for a Staff Machine Learning Engineer to join our team and lead the next phase of our customer retention strategy. This is a critical, full-stack role for an individual who is passionate about the end-to-end lifecycle: from initial research and model development to building the robust systems that power repeat customer engagement and lifetime value growth at scale. Taskrabbit's greatest growth opportunity lies in deepening customer relationships and accelerating repeat purchases. Our most valuable customers are those who return frequently, discover new service categories, and increase their spending over time. There's significant untapped potential in the marketplace: repeat customers spend 3-5x more than one-time users, and category expan

PythonSQLDockerKubernetes
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Applied team works across research, engineering, product, and design to bring OpenAI’s technology to the world. We seek to learn from deployment and broadly distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. We aim to make our innovative tools globally accessible, transcending geographic, economic, or platform barriers. Our commitment is to facilitate the use of AI to enhance lives, fostered by rigorous insights into how people use our products. About the Role We are seeking Software Engineers (Emerging Talent) to join our Applied Engineering team. You’ll work in a highly iterative, collaborative, fast-paced environment to bring our technology to millions of users around the world, and ensure it’s delivered with safety and reliability in mind. We value engineers who are self-starters, care deeply about the end user experience, and take pride in building products to solve customer needs. In this role, you will: Own the development of new customer-facing ChatGPT and OpenAI API features and product experiences end-to-end Talk to users to understand their problems and design solutions to address them Collaborate with a cross-functional team of engineers, researchers, product managers, designers, and operations folks to create cutting-edge products Optimize applications for speed and scale Create a diverse and inclusive culture that makes all feel welcome. Your background looks something like: Bachelor's or Master’s degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 0-1 years of experience in software engineering or a relevant field Proficiency with JavaScript, React, and some backend languages (we use Python) Some experience with relational databases like Postgres/MySQL Interest in AI/ML (direct experience not required) Ability to move fast in an environment where things are sometimes loosely defined and may have competing priorities or deadl

JavaScriptPythonJavaReact
T
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -50%

From $170K/yr

Quick readStrong listing-quality and freshness signals

About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! This role is hybrid requiring 2 days in office at our San Francisco or NYC hub every Tuesday & Wednesday. About the Role Machine Learning is a cornerstone at Taskrabbit, and we're looking for a Staff Machine Learning Engineer to join our team and lead the next phase of our customer retention strategy. This is a critical, full-stack role for an individual who is passionate about the end-to-end lifecycle: from initial research and model development to building the robust systems that power repeat customer engagement and lifetime value growth at scale. Taskrabbit's greatest growth opportunity lies in deepening customer relationships and accelerating repeat purchases. Our most valuable customers are those who return frequently, discover new service categories, and increase their spending over time. There's significant untapped potential in the marketplace: repeat customers spend 3-5x more than one-time users, and category expan

PythonSQLDockerKubernetes
P
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Fraud Data team at Plaid builds the machine learning systems that power Plaid’s fraud detection products, leveraging insights from across Plaid’s network to help identify and stop fraud before it happens. Our team works across the full data science and machine learning lifecycle—from discovering new signals and experimenting with models to deploying and optimizing them in production. We continuously learn from real-world model performance and customer feedback to improve our systems and develop new ways to protect customers and consumers from evolving fraud threats. As a Senior Machine Learning Engineer on Plaid's Fraud Data team, you will develop models that improve fraud detection for our customers. You will identify predictive patterns in Plaid's network data and lead projects from initial experiments through model deployment and ongoing improvement. Investigate fraud patterns and model errors to identify new signals, improve detection, and expand coverage across customers and use cases. Develop training datasets and predictive features, addressing challenges such as incomplete labels, class imbalance, data leakage, and changing fraud behavior. Design, train, and tune model

PythonSQLAWSGit
P
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. Fraud Data is the data science and machine learning team within Plaid’s Fraud organization, responsible for using data and ML to improve and scale Plaid’s fraud products. Within Fraud Data, the Customer & Product Intelligence team focuses on understanding product performance, uncovering customer insights, and enabling go-to-market teams with data-driven solutions. The team partners closely with customers and GTM teams on fraud analyses and proofs of concept, turning customer learnings into scalable, reusable product capabilities. We also build the metrics, analytics, and data foundations that measure product health, identify opportunities for improvement, and guide product decisions across Plaid’s Fraud portfolio. As a Data Science Manager, you will lead a team responsible for customer-facing data science and Fraud product analytics. You will set the team's roadmap, develop its data scientists, and remain involved in analytical methods, technical reviews, and customer investigations. You will: Set a 6–12-month roadmap with Product, Engineering, and GTM, and assign priorities and responsibilities across the team. Define product metrics, their underlying data, and reporting and

PythonSQLAWSMachine Learning
P
📍 New York City, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.7%
Quick readStrong listing-quality and freshness signals

About Pinecone: Pinecone is the trusted AI knowledge company. Its trusted AI knowledge platform—including its Database, Nexus, and Marketplace products—power accurate, fast, and cost-effective AI applications for more than 10,000 customers and 1M developers worldwide. Pinecone's mission is to make AI knowledgeable. For more information, visit pinecone.io . About the Team/Role: This is intended to be an entry level role for new grads/early career professionals. As an Associate Field Engineer, you'll work amongst our Customer Success and Solutions Engineering teams, building deep technical expertise and tooling to directly impact the core product and customer experience. You'll engage with technical stakeholders across the customer lifecycle, from helping prospects understand how Pinecone fits their architecture to guiding existing customers through implementation, optimization, and expansion. You'll collaborate closely with Sales, Product, and Engineering to scope solutions, troubleshoot complex issues, and surface customer insights that shape our product direction. Technical fluency, strong communication, intellectual curiosity, and an AI-first mindset are essential for this role. Responsibilities: Serve as a technical point of contact for customers across their journey from evaluation to production, ensuring best practices and accelerating time-to-value. Deeply understand customer architectures, use cases, and requirements to influence our roadmap and devise solutions that unblock and advance their goals. Build resources to monitor customer health, proactively identifying opportunities to address risk and drive expansion. Troubleshoot issues and deliver high-quality technical support, developing a strong command of Pinecone's product and the broader AI/ML ecosystem. Build processes and capture knowledge to scale and automate support, success, and outreach initiatives. Partner with Sales on technical discovery, deliver product demonstrations, and contribute to proof

JavaScriptPythonJavaAWS
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui

PythonAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI's Industrial Compute organization builds and operates the infrastructure required to train and serve frontier AI models. The Capacity Planning team connects rapidly changing research and product demand with the compute, networking, storage, power, data center, hardware, and operational resources required to make that demand executable. About the Role We are seeking a Technical Program Manager to build and lead capacity planning across OpenAI's large-scale AI infrastructure. You will translate uncertain workload demand into clear infrastructure requirements, allocation decisions, supply commitments, activation priorities, and long-range capacity strategies. This role sits at the intersection of research, engineering, infrastructure, finance, sourcing, deployment, and operations. You will create the planning models, operating cadences, governance mechanisms, and source-of-truth systems that allow teams to understand what capacity is required, what is available, what is at risk, and what decisions must be made. This is not a finance-only forecasting or reporting role. Success requires technical fluency across the infrastructure stack, strong analytical judgment, and the ability to move consequential decisions forward when requirements, timelines, and supply conditions change quickly. Key Responsibilities Own capacity-planning processes across near-term workload allocation, quarterly execution, and longer-range infrastructure horizons. Translate research, training, inference, and product demand into compute, accelerator, cluster, networking, storage, rack, power, and site requirements. Develop scenarios that make assumptions, confidence levels, constraints, sensitivities, and decision points explicit. Reconcile requested demand against contracted, delivered, installed, activated, and workload-usable capacity. Partner with research and engineering teams to understand workload priorities, technical dependencies, utilization patterns, and changing req

PythonSQLAWSRest
NR
📍 Atlanta, Georgia, United States· Full-time
✓ High-confidence listingCompany trend -73.9%

From $138K/yr

Quick readStrong listing-quality and freshness signals

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we believe in a virtuous cycle where customer success fuels improved revenue performance . Our Account Manager role is a pivotal function of an updated pod structure that sits at the intersection of Customer Success, Sales, and Product. Focused squarely on securing and growing our existing revenue base. We are a passionate, upbeat, and highly motivated group committed to exceptional customer loyalty, support, and expansion. We are seeking an Account Manager to join our team and make a positive impact in the continued success and growth of our organization. In this critical role, you will be a driver of retention, adoption, and expansion. Strong leadership ensures positive relationships and alignment across Product, Marketing, and Sales, and Support. If you are passionate about building deep, strategic customer relationships and thrive in a quota-carrying environment, let's talk! What you'll do You will utilize your skills in customer success, renewals, and sales methodologies to exceed customer expectations and deliver value across multiple key performance indicators. Drive Revenue & Retention Pod Collaboration: In partnership with the Account Team, you will function as a secondary seller within a unified pod structure. This role is essential in supporting the Account Executive (Core Seller) through seamless cross-functional alignment and collaborative deal execution. You will review areas of opportunity (whitespace

GitRestAIGo
S
📍 Bellevue, Washington, United States· Full-time· Remote
✓ High-confidence listingCompany trend -92.9%
Quick readStrong listing-quality and freshness signals

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a Senior Software Engineer for our Anti-Abuse team within Security Foundations. This team is responsible for protecting Snowflake and our customers from abuse on the Snowflake platform — building foundational security primitives across a complex multi-cloud environment to combat fraud, account takeovers, data exfiltration, and AI security threats at the scale of Snowflake’s hyper-growth. AS A SENIOR SOFTWARE ENGINEER, SECURITY FOUNDATIONS AT SNOWFLAKE, YOU WILL: Design, build, and scale security-by-default capabilities that protect Snowflake products across a complex multi-cloud environment Develop intelligent security platforms and services that leverage AI/ML and risk scoring to identify, prioritize, and respond to security threats and policy violations Build real-time detection and prevention systems to defend against account compromise, AI abuse, and data exfiltration through adaptive enforcement and granular policy controls Create security intelligence and monitoring capabilities that detect suspicious account & user activity, generate actionable insights, and enable timely response Partner across Security, Engineering, and Product teams to define security strategy, influence product design, and drive adoption of secure engineering practices Raise the

JavaScriptPythonJavaSQL
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team builds next-generation AI-native silicon and systems while working closely with software, research, and manufacturing partners to co-design hardware tightly integrated with AI models. In addition to delivering systems for OpenAI’s supercomputing infrastructure, the team develops the tools, methodologies, and strategic partnerships needed to accelerate hardware innovation. About the Role We’re seeking an experienced Hardware Strategic Sourcing Manager to own sourcing strategy and supplier partnerships for fiber and optical interconnect components across OpenAI’s next-generation AI infrastructure. Reporting to the Head of Partnerships & Strategic Sourcing, you will lead sourcing across fiber cable assemblies, internal optical harnesses, fiber shuffles, optical backplane assemblies, connectorized and standalone passive optical assemblies, fiber-array units (FAUs), fiber-to-chip and coupling interfaces, detachable connectors, optical routing, and assigned optical packaging, assembly, and test services. You will work closely with electrical engineering, optical engineering, systems engineering, mechanical and packaging engineering, quality, rack integration, data-center deployment,manufacturing, supply chain, finance, legal, and program management teams to translate demanding bandwidth, signal integrity, reliability, and scale requirements into resilient supplier partnerships and scalable commercial strategies. Your work will directly support the performance, reliability, manufacturability, and scale of the high-speed optical connectivity required for OpenAI’s next-generation AI systems. In this role, you will: Develop and execute a comprehensive sourcing strategy for fiber and optical interconnect components supporting high-bandwidth AI systems and infrastructure. Own sourcing across optical fiber cable assembli

AWSRestAIGo
D
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -93.7%

$150K – $230K/yr

Quick readStrong listing-quality and freshness signals

Drata is building the trust layer between great companies - automating compliance, managing risk, and helping organizations prove trust continuously as they scale. We're Dratanauts: a global crew of 600+ professionals united by a culture that rewards integrity, ownership, and raising the bar, no matter where in the world we're working from. Why Join the Drata Team? At Drata, you're not maintaining legacy compliance software - you're building the agentic AI platform defining what trust looks like for the next generation of companies. Here's what makes the work itself worth showing up for: Problems without a playbook: You'll work at the edge of AI and security, building agentic governance, continuous compliance, and real-time trust verification to solve problems that don't have an established answer yet. You're writing it as you go. Real ownership, not just process: Our values center on owning outcomes and raising the bar, not checking boxes. You're expected to have opinions and back them. A seat at the table: Your perspective is unique and valued. Open debate and diverse viewpoints are built into how decisions actually get made here, at every level. Growth at rocketship speed: Drata is scaling fast, which means scope grows fast too. High performers get more ownership, visibility, and experience. A crew, not just coworkers: Dratanauts consistently describe a "come as you are" culture with sharp, curious people—the kind of team that makes hard problems genuinely fun to solve. See what they say here and follow us on LinkedIn for company news, employee stories, and career updates. Job Summary: Drata is reimagining compliance as an intelligent, always-on experience — and AI is at the center of that vision. We are seeking a Senior AI Product Engineer to own the full-stack development of customer-facing AI features, embedded directly within our product teams. This is not a platform or infrastructure role. You'll translate the capabilities of LLMs, agents, and RAG pipelines

TypeScriptPythonReactNode.js
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture. EXAMPLE INITIATIVES The B200 Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's first Blackwell GPU clusters. Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems to move customer workloads seamlessly across regions to optimize cost and latency. Precision GPU Triage: Developing automated Go-based operators to identify, cordon, and repair unhealthy H100 nodes in under an hour. The Supply Chain of Intelligence: Partnering with lead

PythonAWSAzureGCP
🔔

Get new hr business partner director jobs in United States by email

Daily job updates · Unsubscribe anytime