Jobiba hiring network

Design And Cost Estimation Head Jobs

6,587 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current design and cost estimation head jobs. Use filters to narrow by work mode, employment type, experience and date posted.

P
Pendo
📍 New York• Full-time• $300K – $325K/yr
1mo ago

The Team + The Role Our Emerging Team is focused on building AI Products for our product experience (PX) platform. We build from the ground up to explore, prototype, and ship AI-native experiences that change how software teams understand and serve their users. This is not an AI layer added to existing product; it is a deliberate bet on what product intelligence looks like next. The team operates with high autonomy, moves quickly, and builds products without clear precedents. As a Staff Software Engineer (AI), you will sit at the intersection of deep technical capability and strong product judgment. You will design and build production-grade AI systems, including RAG pipelines, agentic workflows, and LLM-powered features, while making clear tradeoffs across prompting, fine-tuning, architecture, evaluation, and deployment. You will also partner closely with product, design, and engineering stakeholders to frame the right problems and communicate technical decisions clearly. This role is based in our New York office. What this looks like day-to-day Applied AI systems: Design and build AI-native systems, including RAG pipelines, agentic workflows, and LLM-powered product features. You will take ideas from prototype through production and ensure they can support real users. Model strategy: Make principled decisions about when to prompt, when to fine-tune, and when to use a different technical approach entirely. You will explain those tradeoffs clearly to engineers and non-engineers. Evaluation and guardrails: Instrument and evaluate model outputs rigorously by defining evaluation frameworks and identifying hallucinations early. You will implement guardrails that hold up under real-world usage and load. Productionize AI ownership: Own model deployment, monitoring, latency optimization, cost management, and reliability at scale. You will ensure AI systems are observable, performant, and production-ready. Full-stack delivery: Contribute across the stack when needed to get

P
Pendo
📍 New York• Full-time• $250K – $275K/yr
1mo ago

The Team + The Role Our Emerging Team is focused on building AI Products for our product experience (PX) platform. We build from the ground up to explore, prototype, and ship AI-native experiences that change how software teams understand and serve their users. This is not an AI layer added to existing product; it is a deliberate bet on what product intelligence looks like next. The team operates with high autonomy, moves quickly, and builds products without clear precedents. As a Sr. Software Engineer (AI), you will sit at the intersection of deep technical capability and strong product judgment. You will design and ship applied AI systems, including RAG pipelines, agentic workflows, and LLM-powered features, from prototype through production. You will make principled technical decisions, evaluate model behavior rigorously, and communicate tradeoffs clearly to engineers and non-engineers alike. This role is based in our New York office. What this looks like day-to-day Applied AI systems: Design and build AI systems including RAG pipelines, agentic workflows, and LLM-powered features. You will take work from prototype through production and ensure it can hold up in real customer environments. Technical decision-making: Make principled decisions on when to prompt, when to fine-tune, and when to use a different tool entirely. You will explain these tradeoffs clearly so the team can move quickly without sacrificing quality. Model evaluation: Instrument and evaluate model outputs rigorously by defining evaluation frameworks and catching hallucinations early. You will implement guardrails that can withstand real-world load and production use. Productionize AI ownership: Own model deployment, monitoring, latency optimization, cost management, and reliability at scale. You will help ensure AI systems are observable, efficient, and dependable in production. Full-stack product shipping: Contribute across the stack when needed because this team ships products, not just models

C
Coinbase
📍 - USA• Full-time• Remote• From $186.1K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Software Engineer on the Platform Core Automation team, you'll build the AI infrastructure and agentic systems that automate customer support and compliance operations at Coinbase. This team is reimagining what these processes look like when AI handles them end-to-end, replacing manual workflows with intelligent agents that resolve customer issues and execute compliance tasks autonomously. You'll own the design and delivery of LLM-powered systems, grounding techniques, and integration pipelines that directly reduce resolution times, cut costs, and improve accuracy across millions of customer interactions. What you’ll be doing (ie. job duties): Own the design and delivery of agentic AI systems that power Coinbase's customer support and compliance automation, from LLM orchestration through production deployment and monitoring. Build scalable, secure backend infrastructure in Python and Golang that serves AI workloads, including model integration pipelines, guardrails, grounding mechanisms, and measurement frameworks. Drive end-to-end project execution on complex AI initiatives, making technical trade-offs across latency, accuracy, cost, and reliability. Partner with Customer Experience, Compliance, and product engineering teams to identify high-impact automation opportunities and translate operational pain points into AI-powered solutions. Strengthen engine

REMOTEpythonmongodbaws
View job →
A
Amplitude
📍 Remote• Full-time• $198K – $299K/yr
1mo ago

Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Staff Platform Engineer, you'll set technical direction for the platform across teams, lead our highest-complexity and highest-leverage initiatives, and shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll operate across team boundaries — partnering with product engineering, fellow Staff+ engineers, and engineering leadership to make Kubernetes and cloud infrastructure effortless across the entire engineering org. You'll build the self-service automation, shared standards, and scalable AWS and GCP infrastructure that let dozens of product teams ship faster, safer, and with less cognitive load — and you'll multiply the engineers around you while you do it. Key Responsibilities Set technical direction — shape platform and domain-level technical strategy that improves developer experience, reliability, security, and cost, and lead the high-complexity, cross-cutting initiatives that deliver it with measurable impact for the organization. Drive clarity through ambiguity. Take on the most loosely-defined problems, validate the critical assumptions early, and create alignment with stakeholders across teams so others can move quickly and confidently — driving cross-team decisions to a timely close and escalating when needed. Build the AI-augmented platform. Design org-wide tooling, guardrails, and policy-as-code that help every engineer get more out of AI-assisted development — infra primitives an LLM can safely reason about and PR against, automated review, and standards that hold as AI changes how code gets written. Own Infrastructure-as-Code standards for Kubernetes, AWS, and GCP using Terraform, Helm, Kustomize, and emerging tooling — setti

pythonawsazure
View job →
F
Figma
📍 Ca New York• Full-time• From $185K/yr
1mo ago

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! Figma seeks an AI-Native Performance TPM to own performance prevention, diagnostics, and safe rollout for both our flagship products and next-gen AI features. This platform-level role spans desktop, browser and mobile; requires deep experience with performance testing (load, stress, endurance, interference), observability, and incident response. You’ll partner across Product, Platform, Performance Program and Product Support. Join us to keep Figma snappy quick! This is a full time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: This is a specialist Platform TPM owning horizontal, high-visibility performance programs that span: Flagship product performance (load times, FPS, memory) and AI features (model inference latency, throughput, cost, and reliability) Desktop, browser, native mobile (React Native / WebView), and WASM contexts Cross-org programs: observability/telemetry, regression prevention (performance CI), xfn performance forum, SEV mitigation, safe rollout of AI capabilities, and CE Planning (customer engineering / enterprise readiness) We’d love to hear from you if you have: 5+ years in performance engineering, performance TPM, platform TPM, or SRE with hands-on experience shipping performance programs for SaaS products Demonstrated experience with load, stress, performance or scalability testing and new-build comparisons Deep familiarity with web performance (FCP, LCP), rendering/FPS, WASM memory, mobile profiling (Xcode I

reactawsci/cd
View job →
FB
15 days ago

Location : Come and join us in London! Freenow by Lyft empowers smarter mobility decisions helping people to move freely and cities to thrive. Be ready to work in a multinational, diverse, highly motivated and collaborative team of passionate professionals who strive for excellence and like to have fun. Are you ready for your next ride? The Manager, AV Fleet Services owns fleet management execution for our London AV operation. This role builds and manages the vendor ecosystem that supports the London fleet, defines the process standards that govern how our depot team and external partners work together, and owns the operational and financial performance of the fleet services function in London. YOUR DAILY ADVENTURES WILL INCLUDE: Design and build the London vendor architecture across fleet repair, service, maintenance, recovery, and towing. Decide what to insource vs. outsource, which partners to consolidate vs. diversify, and how the vendor mix evolves as the fleet scales. Own sourcing, evaluation, and onboarding. Draft and negotiate MSAs and SOWs with support from Legal and Procurement. Protect labor rates, parts pricing, and SLA provisions in new agreements. Own quarterly business reviews and ongoing performance management. Escalate SLA misses, manage vendor scorecards, and drive corrective action when performance dips. Design and document process flows for how the depot team interacts with external vendors on repairs, service, and maintenance. Own the playbooks for vehicle intake, damage assessment, repair routing, and return-to-service. Establish operational standards for depot and vendor workflows that align with partner requirements. Continuously improve based on operational data. Own damage assessment protocols and repair authorization workflows. Interface with vendors on complex or high-cost repairs before approval. Own the London fleet services budget. Track and report on labor, parts, tooling, and vendor spend against targets. Maintain comprehensive repai

aigoexcel
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
15 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a Design Verification Engineer to contribute to the unit-level verification of our RISC-V CPU front end—including instruction fetch, branch prediction, and surrounding fetch control structures. As a key individual contributor within our front-end DV team, you will focus on building testbenches, generating stimulus, and developing checkers for complex microarchitectural scenarios. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Front-End Experience: Solid background verifying CPU front-end blocks—such as instruction fetch, branch predictors (BTB, TAGE, RAS), instruction caches/TLBs, or decode—using SystemVerilog, UVM, and C++. Unit-Level Focus: Hands-on experience building clean, controllable unit testbench environments with precise stimulus and checking. Reusable Design: Pragmatic approach to building transactors, predictors, and scoreboards that can be reused across verification levels. Collaborative Problem Solver: Comfortable working alongside RTL designers to trace and fix misspeculation, redirect, and edge-case fetch bugs. What We Need Unit-Level Verifica

awsaic++
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
15 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are seeking a GCC Compiler Engineer to design, develop, and optimize compilers for next-generation RISC-V and AI compute architectures. You will work across hardware and software teams to improve performance, programmability, and integration of our custom toolchains into real applications. This role is fully hands-on and central to how developers interact with Tenstorrent hardware across both traditional compute and advanced machine learning workloads. This role is Hybrid, based out of Santa Clara, CA, Austin, TX, or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced compiler engineer with deep knowledge of GCC and LLVM internals, comfortable optimizing for custom hardware targets. Strong C/C++ developer with a solid grasp of algorithms, data structures, and performance analysis. Collaborative and analytical, able to work across hardware and software domains to deliver efficient, high-performance toolchains. Passionate about enabling breakthrough compute architectures through compiler innovation and software-hardware co-design. What We Need Design, develop, and optimize GCC and/or LLVM compilers for Tenstorrent’s cust

awsmachine learningai
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
15 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We're looking for an Engineering Program Manager — or a technical engineer ready to move into program management — to help drive execution on our RISC-V CPU team. You'll work across architecture, design, verification, physical design, and DFT to keep a high-performance CPU program moving from spec through tapeout and post-silicon debug. Partnering with a senior program lead and engineering managers, you'll own schedules, track milestones, surface risks early, and keep cross-functional teams aligned. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experience in CPU, SoC, or silicon development — as an engineer, technical lead, or program manager. Comfortable working with cross-functional engineering teams and translating technical detail into clear status, dependencies, and trade-offs. Organized and proactive; you chase down open items rather than waiting for them to resolve. A clear communicator who can hold your own in engineering discussions and write updates leadership can act on. What We Need Support the full CPU development cycle — spec, design, verification, DFT, ta

awsaisem
View job →
A
Asana
📍 Warsaw• Full-time• $372K – $432K/yr
1mo ago

We're looking for a Senior Infrastructure Engineer who brings strong software engineering skills and a deep understanding of production systems. This role is a good fit for someone who enjoys building systems that make infrastructure more scalable, reliable, and easy to operate – using code, not runbooks. You'll work with a highly collaborative team to design and build the internal platforms that power all of Asana, from product features to AI systems to offline analytics. Our tech stack includes: AWS, Kubernetes (EKS), MySQL (RDS), OpenSearch, DynamoDB, Redis, Terraform, Datadog, TypeScript, Scala, Go, and Python. We’re especially interested in people who think like backend engineers but care deeply about systems – things like failure modes, operational cost, debuggability, and performance. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you’ll achieve: Design and build frameworks, tools, and services that improve the reliability, observability, and scalability of Asana’s infrastructure. Lead end-to-end projects, from scoping and design through to rollout, across multiple systems and teams. Improve the operability of stateful infrastructure like MySQL, OpenSearch, and DynamoDB – and help drive Asana’s long-term vision for storage reliability. Debug production issues across the stack. Yes, there’s an on-call rotation – but this isn’t a pager monkey role. You’re here to fix things properly and make sure they don’t break again. Partner with product teams to shape a service-oriented architecture that enables fast, reliable development. Share knowledge through code reviews, design discussions, and mentorship. Abou

typescriptpythonsql
View job →
O
Okta
📍 Washington• Full-time• From C$132K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About the Role: Okta is looking for a highly skilled Globalization Engineer to be the technical backbone of our global digital operations. You will be responsible for owning, building, and optimizing the automation and technology stack that powers our localization team. This is a hands-on role for an experienced technical expert who is passionate about building robust, scalable systems. You will leverage your deep expertise in APIs, AI, and internationalization to solve complex technical challenges and drive efficiency. As a subject matter expert, you will also act as a key technical consultant to internal teams, helping them prepare content for a global audience. What you'll be doing: Build and Own the Localization Automation Framework: Design, develop, and maintain the core automation workflows for our content lifecycle using agentic workflows, Python and APIs, connecting our TMS with content sources, repositories, and other internal tools. Spearhead AI and Technology Integration: Research, evaluate, and implement cutting-edge localization technologies, including AI/LLM-based solutions and advanced TMS features, to improve the quality, speed, and cost-effectiveness of our workflows. Provide Technical Oversight and Systems Integrity: Act as the authority for resolving high-impact architectural failures. Conduct deep-dive root cause analysis on complex integration issues, file parsing conflicts, and system-wide bottlenecks to ensure uninterrupte

pythonawsci/cd
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de

awsrestai
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
15 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As our TT-Distributed Software Engineer, you will develop and optimize distributed software systems that power the most efficient and highest-performing AI and HPC clusters. In this role, you'll work on distributed programming across multiple nodes, utilizing systems programming, inter-node communication, and Tenstorrent’s scalable architectures to advance the state-of-the-art distributed inference and training infrastructure. This role is hybrid, based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Strong C or C++ engineer with solid foundations in systems programming, operating systems, and distributed systems principles. Enthusiastic about distributed computing, including IPC, socket programming, and cluster resource coordination. Comfortable reasoning about scalability, fault tolerance, and performance across multi-node environments. Curious and first-principles thinker who challenges conventional approaches to distributed system design. Motivated to grow into a deep technical expert in large-scale distributed AI infrastructure. What We Need Architect, implement, and optim

awsaic++
View job →
T
15 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a skilled Software Engineer with a passion for building high-performance, low-level systems software. In this role, you’ll contribute to the development and optimization of the infrastructure that powers our cutting-edge processors, with a primary focus on C/C++ development and low-level programming. You'll work closely with large inference and training model development to further drive Scale Out software and hardware performance. This role is hybrid, based out of Toronto, ON. Who You Are Strong C or C++ systems engineer with a deep understanding of memory, threading, I/O, and low-level execution models. Experienced building low-level software, drivers, embedded systems, or performance-critical infrastructure. Comfortable working close to hardware and curious about how systems behave under the hood. Proficient with Linux systems programming and debugging tools such as gdb, strace, and perf. Structured problem solver who thrives in fast-paced, highly technical environments. What We Need Design, develop, and maintain core infrastructure software that interfaces directly with Tenstorrent hardware. Build low-level libraries and APIs for communication and synchronization across compute nodes. Optimize system-level software for performance, scalability, and reliability in distributed environments. Support hardware

awslinuxai
View job →
🔔

Get new design and cost estimation head jobs by email

Daily job updates · Unsubscribe anytime