About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine
Market pay estimate
$153,120–$244,000 / year for comparable Software Engineer roles. Not employer-provided.
Role overview
Job description
Join us in building the future of finance.
Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading.
About the team + role
We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards.
The Capacity & Efficiency Engineering team builds the software that manages, governs, and reduces Robinhood's AWS cloud spend, operating at the intersection of cloud infrastructure, data engineering, and FinOps. The team owns the full lifecycle of cloud cost: the data platforms that make spend transparent and attributable, the anomaly detection and forecasting systems that make it predictable, the automation that continuously rightsizes infrastructure at fleet scale, and the capacity planning and commitment strategy that keep a bursty, latency-sensitive trading platform both reliable and cost-effective. Our work carries CEO-level visibility and has already driven millions of dollars in annualized savings. We partner closely with Data Science, Infrastructure, Finance, and product engineering teams across Robinhood to turn cost insight into real financial outcomes.
As a Staff Software Engineer on the team, you will serve as the technical lead for the Capacity arm: setting engineering direction, owning the most complex systems, and raising the bar for how the team builds and operates. You will design the platforms that attribute cloud costs to their real consumers, catch spend regressions before they compound, and forecast future demand and unit economics. Just as importantly, you will build the systems that act on those insights: automation that safely optimizes production infrastructure, and tooling that keeps our commitment portfolio matched to actual usage. You will work directly with senior leadership and partner organizations to translate complex cost data into clear accountability. This is a role for an engineer who leads from the front technically while shaping the direction of a team that operates at the highest levels of the company.
This role is based in our Bellevue, WA office, with in-person attendance expected at least 3 days per week.
At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams.
What you’ll do
- Lead capacity planning and demand forecasting for a latency-sensitive, bursty production environment, ensuring the platform scales to absorb peak demand without sustained over-provisioning
- Design and build scalable cost attribution and chargeback systems that accurately identify true resource consumers across platform and product teams, enabling fair and transparent cost accountability
- Architect and own automated anomaly detection and governance tooling that detects cost regressions, routes findings to the correct owning teams, and tracks remediation outcomes end-to-end
- Partner closely with Data Science to develop and maintain forecasting models, year-over-year cost projections, and unit economics frameworks (e.g., cloud cost per engaged user) that surface future spend risk proactively
- Drive Robinhood's AWS contracts and pricing strategy by partnering with Amazon on reserved capacity, discount models, and on-demand optimization to ensure cost-effective infrastructure investment
- Shape how the company measures and governs fast-growing AI/ML infrastructure spend, and apply AI-driven automation to cost operations such as anomaly triage, reporting, and remediation tracking
What you bring
- Strong data engineering skills, including distributed data processing (e.g., Spark), workflow orchestration (e.g., Airflow), and analytical data stores serving interactive workloads
- Strong proficiency with Kubernetes and a deep practical understanding of resource efficiency, capacity planning, and the relationship between infrastructure configuration and cloud spend
- Familiarity with cloud billing data and commercial constructs (usage and billing reports, Reserved Instance and Savings Plan amortization, enterprise agreements) is a strong plus
- Experience building automation that safely modifies production infrastructure, with guardrails, progressive rollout, and rollback built in, not just systems that observe and report
- Ability to reason deeply about cloud cost models, including tradeoffs between reserved vs. on-demand capacity, instance efficiency, and the financial impact of infrastructure decisions at scale
- Ability to communicate fluently with both engineering and finance audiences, translating infrastructure decisions into financial outcomes and presenting to senior leadership
What we offer
- Challenging, high-impact work to grow your career.
- Performance-driven compensation with multipliers for outsized impact, bonus programs, equity ownership, and 401(k) matching.
- Best-in-class benefits to fuel your work, including 100% paid health insurance for employees with 90% coverage for dependents.
- Lifestyle wallet — a highly flexible benefits spending account for wellness, learning, and more.
- Employer-paid life & disability insurance, fertility benefits, and mental health benefits.
- Time off to recharge including company holidays, paid time off, sick time, parental leave, and more!
- Exceptional office experience with catered meals, events, and comfortable workspaces.
In addition to the base pay range listed below, this role is also eligible for bonus opportunities + equity + benefits.
Base pay for the successful applicant will depend on a variety of job-related factors, which may include education, training, experience, location, business needs, or market demands. The expected base pay range for this role is based on the location where the work will be performed and is aligned to one of 3 compensation zones. For other locations not listed, compensation can be discussed with your recruiter during the interview process.
Base Pay Range:
Click here to learn more about our Total Rewards, which vary by region and entity.
If our mission energizes you and you’re ready to build the future of finance, we look forward to seeing your application.
Robinhood provides equal opportunity for all applicants, offers reasonable accommodations upon request, and complies with applicable equal employment and privacy laws. Inclusion is built into how we hire and work—welcoming different backgrounds, perspectives, and experiences so everyone can do their best. Please review the Privacy Policy for your country of application.
What they are looking for
Skills & requirements
Hiring company
Robinhood
Explore this employer's active roles, salary signals and company profile on Jobiba.
Keep exploring
Similar active roles
Fresh roles matched to this title and market.
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is seeking a Staff Software Engineer, Capacity Engineering. The team is responsible for efficiently managing one of the largest-scale cloud-native infrastructures in the world. This role is highly impactful, as efficiency is an ongoing strategic priority for Pinterest. The role has direct visibility across Pinterest Engineering and with Engineering and company leadership. The team is looking for a candidate with a strong background in implementing performance and efficiency projects on large scale distributed systems. In this individual-contributor role you will own and drive performance and efficiency for a core area of Capacity Engineering, partnering with the company-wide efficiency lead and collaborating with performance and efficiency leaders across the organization. What you’ll do: Drive efficiency in large-scale shared environme
As a Staff Engineer on Datadog's Compute – Disruption and Workload Placement team, you'll help define how our Kubernetes fleet scales to meet the demands of rapidly growing AI and cloud-native workloads. You'll work on the systems that ensure engineering teams have the right compute capacity, in the right region, at the right time across AWS, Google Cloud, and Azure. This is a highly technical, high-impact role where you'll shape the future of capacity orchestration, influence platform architecture, and solve infrastructure challenges that directly support Datadog's continued growth. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead the technical direction of capacity management and workload placement for Datadog's Kubernetes platform spanning 100,000+ virtual machines across multiple cloud providers. Design and build systems that optimize how engineering workloads are scheduled and deployed across regions while balancing capacity constraints, reliability, and performance. Partner across infrastructure teams to evolve multi-region and multi-cloud capacity orchestration as Datadog continues to scale. Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency. Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies. Who You Are: You have significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems. You are an experienced software engineer with strong programming skills, ideally in Go or a comparable systems programming language. You have hands-on experience with at least one major cloud provider (AWS, Google Cloud, or Azure) and understand distributed cloud infrastructure. Yo
The Team MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of either our Dublin or Cork office or remotely in Ireland. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Build for reliability, making services and infrastructure avail
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role: Join our Infrastructure Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Infrastructure Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff Infrastructure Engineers who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios. Optimize Performance and Infrastructure: Collaborate with core infrastructure and product teams to performance tune and optimize our cloud deployments (Kubernetes, Docker, GCP). Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions. Elevate Developer Experience: Design and implement improvements to our build, test, and deployment systems to make software delivery faster, safer,
Location : Come and join us in Hamburg. Freenow by Lyft empowers smarter mobility decisions helping people to move freely and cities to thrive. We are looking for an Engineering Manager to lead a high-performing team of software engineers focused on our Rider domain — the customer-facing app and platform that millions of riders use every day. Your team will play a pivotal role in bringing Lyft's product experience to Europe, adapting and building the systems that let riders book rides seamlessly across all markets This critical role ensures team’s alignment with product and business strategy, while driving efficient delivery of features. You will foster an environment of technical excellence, psychological safety, and continuous individual development, ultimately maximizing business impact across all Freenow markets through innovation. Be ready to work in a multinational, diverse, highly motivated and collaborative team of passionate colleagues who strive for excellence and enjoy their work. Are you ready for your next ride? YOUR DAILY ADVENTURES WILL INCLUDE: Team Performance & Delivery: Within the frame of quarterly goals and company strategy, lead the team in delivering high-quality, scalable software solutions. Focus on continuously improving engineering processes to ensure predictable, timely delivery while managing team capacity and focus. You inspire innovation to disrupt incumbents. People Development: Provide coaching, mentoring, and continuous feedback to team members. Identify development areas and growth opportunities to support individual career paths and enhance team competence in building innovative customer and driver support solutions. You get to build a team from the ground up. Technical Ownership: In collaboration with Staff Engineers and Principals, lead technical decision-making for the support systems. Ensure high standards of quality, reliability, maintainability, and observability in all solutions, particularly the large-scale syste
🔔 Get job alerts
New Staff Software Engineer, Capacity and Efficiency Engineering jobs in Bellevue, WA, straight to your inbox.
No spam · Unsubscribe anytime