ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto
Jobiba hiring network
Production Cleaning Specialist Jobs
3,233 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current production cleaning specialist jobs. Use filters to narrow by work mode, employment type, experience and date posted.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re hiring a Data Engineer to build and scale Baseten’s internal data platform. This role sits at the intersection of data engineering, analytics, and data science, transforming raw product and business data into reliable datasets that power decision-making. You’ll design the data models, pipelines, and analytics infrastructure that enable teams across Product, Engineering, Finance, Marketing, and Sales to understand usage and performance. This includes working with AI inference, infrastructure, and observability data to generate insights about the product, business operations and platform economics. You’ll partner closely with stakeholders to build robust, scalable pipelines, define company-wide metrics that inform strategy and planning. RESPONSIBILITIES Design and maintain core data models and semantic layers Develop and orchestrate batch and streaming data pipelines using technologies such as Apache Beam, Kafka, Airflow, or similar frameworks Analyze inference and infrastructure telemetry , including data from OpenTelemetry, Grafana, and other observability tools Define and maintain company-wide metrics across product usage, performance, and customer lifecycle Enable self-service analytics through agents and tools, with well-structured semantic layers and context Ensure data reliability and quality through testing, documentation, and governance PREFERRED QUALIFICATIONS Understanding of inference metrics s
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. At Baseten, we are building the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale, the network is the computer. We are looking for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building block in our infrastructure and unlocking the next generation of distributed inference optimizations. THE OPPORTUNITY Networking and compute are no longer separate disciplines; they are converging. The massive throughput of H100, B200, and NVL72 architectures enables and demands a new approach where communication is co-optimized alongside computation. We are entering an era where the network is an active accelerator, leveraging smart hardware offloads and direct interconnects to ensure that data movement operates at wire-speed. In this role, you will go beyond network configuration to architect the software fabric that unifies thousands of GPUs into a cohesive operating system. While you will leverage the best of the open-source ecosystem, you won't be limited by it. Where off-the-shelf solutions stop, you will build from scratch, engineering the primitives required to co-optimize communication and compute for Disaggregated Serving, Wide Expert Parallelism (WideEP), and lightening cold starts. WHAT YOU'LL DO Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack,
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE The largest, most demanding enterprises are starting to run on Baseten, and they arrive with a range of security, compliance, and procurement requirements. As a Senior Engineer on Baseten's enterprise engineering team, you'll build the capabilities that enable large organizations like Writer, HubSpot, and Notion to succeed on Baseten. Enterprise engineering authors the core building blocks, APIs, and user experiences powering the Baseten platform: identity and access management, billing, regional isolation, and self-hosted and single-tenant deployment options. This is deep product and systems work across the full stack, from designing authentication and authorization systems using standards like OAuth and OIDC to shipping the admin experiences enterprise IT teams use to manage their organization. EXAMPLE INITIATIVES Recent and upcoming work on the team: Fine-grained authorization for users, service accounts, and agentic workloads SSO and SCIM support, allowing customers to centralize and automate access to Baseten Expanding the billing platform to support evolving pricing models, advanced data exports, and controls to manage spend In-product management and enforcement of customer compliance requirements like data residency and HIPAA Securing network paths in and out of a customer's models with private connectivity and ingress and egress restrictions Allowing customers to run Baseten inside their own VPC, on-pr
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: As a Software Engineer at Baseten, you will own one of the most critical surfaces of our business: pricing, billing, and revenue infrastructure. As we launch more and more products— billing is no longer just operational plumbing. It is a strategic lever for growth. This role will establish clear ownership of billing as a function and create leverage for Finance, Sales, and GTM teams while maintaining a seamless customer experience. RESPONSIBILITIES: Own Baseten’s end-to-end billing and revenue infrastructure, including pricing, invoicing, metering, and reporting foundations. Build and evolve our billing platform and integrations (including Orb), ensuring correctness, auditability, and a high-trust experience for customers and internal teams. Partner closely with Finance, Sales, GTM, and Forward Deployed Engineering to turn real-world workflows into reliable internal tooling and automation (quoting, approvals, renewals, usage reconciliation, revenue reporting). Design systems that scale with new products, packaging, and go-to-market motions, making billing a strategic lever for growth. Drive reliability and operational excellence for revenue-critical workflows: monitoring, alerting, incident response, backfills, and clear runbooks. Lead from the front on high-impact projects: clarify requirements, propose crisp technical approaches, ship iteratively, and raise the bar on quality and velocity. Debug and resolve
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg
About Us What if your work could drive change in a globally established industry, shaping processes that touch every corner of the world? At Forto, we are at the forefront of change, harnessing the power of AI to revolutionise logistics. We want to reinvent digital supply chains to be transparent, frictionless and sustainable. From day one, our mission has been to simplify global trade – creating a seamless and efficient logistics process. Your role & Mission The Site Reliability Engineering team at Forto is responsible for reliability and developer experience. We enable our development teams to write complex business logic by providing best-in-class tooling and infrastructure. We have a production environment based on GCP, Kubernetes, Terraform, and Helm. On top of that, we have self-service tooling written in TypeScript. “You build it, you run it” - our job is to make that real. This is a high-ownership role on a lean team that directly shapes how 70+ engineers build and ship software. If you care about platform quality and want your work felt immediately across an engineering org, this is a great match for you. What you will do Build out our runtime platform as a self-service product that enables our engineering teams to write code, run workloads, and drive engineering culture forward. Bring software development skills and practices into platform engineering, such as code quality, domain-driven design, and test-driven development. Own the developer portal and internal platform roadmap, including leading this year's overhaul of our CI/CD pipelines in collaboration with all product teams. Ensure site reliability by building observability solutions, deployment, and disaster recovery capabilities. Own reliability standards end-to-end through SLOs and error budgets — shaping how teams balance velocity and risk. Drive infrastructure cost optimisation across Kubernetes, MongoDB, and Datadog at scale. Improve our security posture through tooling, compliance work, and
About the Role & Team We’re looking for an Engineering Manager to lead the Data Infrastructure team within Statsig Experiment at Amplitude. You will lead a multidisciplinary team of software engineers, data engineers, and data scientists responsible for the systems that power experimentation at scale. The team owns three critical areas: Data ingestion: Collecting and importing experiment exposures, custom events, OpenTelemetry data, and real user monitoring data across SDKs, streaming systems, cloud storage, and customer data warehouses. Data computation: Building distributed computation systems that transform raw data into accurate, timely experiment results. Stats engine: Developing and productionizing the statistical methods that help customers make trustworthy decisions from their experiments. This is not a traditional data engineering management role. We are looking for a leader with a solid data science and statistical foundation who can connect advances in experimentation methodology with scalable production systems. You will help set our technical and scientific direction, translating new statistical methods and machine learning research into capabilities that customers can use reliably at scale. You’ll partner closely with data scientists, engineers, product managers, and customers to advance the state of experimentation. The ideal candidate is equally comfortable discussing causal inference and statistical power with data scientists, distributed computation architectures with engineers, and experimentation strategy with customers. What You’ll Do Lead and grow the team responsible for Statsig’s data ingestion, experiment computation, and stats engine. Define the technical and scientific strategy for advancing experimentation across both Statsig Cloud and warehouse-native deployments. Partner with data scientists and engineers to turn new statistical and causal inference methods into scalable, reliable product capabilities. Evolve our data and computatio
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta’s Developer Experience team builds the tools engineers use every day to bring ideas to production rapidly and reliably. You’ll empower other Vanta engineers to leverage cutting-edge technologies and best practices to make Vanta more performant and scalable on a platform level. Example projects include modernizing our CI/CD pipelines, introducing new test frameworks, launching AI-powered dev tools, and scaling developer environments to support a growing engineering team. This team has a wide breadth of impact across all of product engineering. The work we do compounds in value by making it easier for engineers to diagnose and solve bugs, streamline workflows, and ship value to our customers quickly and safely. Vanta engineers design and develop new product functionality and infrastructure leveraging modern frameworks and tooling, including TypeScript, React, Node.js, MongoDB, Github Actions, and various AWS services such as Fargate and ECS. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. We’d love for you to join us! You will: Set direction for critical dev infrastructure, enabling us to stay ahead of continued rapid growth Design and build CI and build systems that ensure Vanta engineers can develop and ship robust products quickly and confidently Improve the efficiency and reliability of our deployment workflows, including tools for hotfixes, rollbacks, and incident mitigation Lead development of tools that accelerate feedback loops — from typechecking and linting to running tests and deploying changes Build and maintain scalable developmen
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. As a Sr. AI GTM Engineer on Vanta's GTM Engineering team, you will design and ship internal AI products and processes that completely transform how our go-to-market teams operate and win. You'll embed with Sales, Customer Success, and Revenue Operations to build novel applications and platforms that accelerate pipeline generation, improve win rates, and scale how Vanta engages customers. You will also own end-to-end outcomes, move quickly from concept to production, and directly shape how customers experience Vanta in the field. This role is ideal for AI-pilled builders who want to be close to users, solve complex business problems with elegant technical solutions, and define entirely new categories of enterprise software. Vanta's GTM Engineering team operates as an internal incubator: applying AI at scale to revolutionize customer engagement. We build high-impact tools that reshape conversations, learn from every customer interaction, and demonstrate the value of our platform in real-world scenarios. This team sits at the intersection of engineering and GTM, shipping solutions that range from rapid prototypes to production-grade systems. What you'll do as a Senior AI GTM Engineer at Vanta: Ideate and build products that transform how Vanta generates new business, interfaces with prospects and customers, and grows existing customer relationships Own the full product lifecycle: prototype, iterate, ship, and maintain solutions that solve real-world customer and internal workflow problems Partner closely with Vanta’s CRO and executive leadership teams to architect and execute plans that bring Vanta to the forefront of GTM orgs lev
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the role We're looking for an engineer to drive the evolution of OrioleDB and its collaboration with the upstream PostgreSQL community. OrioleDB is a next-generation storage engine for PostgreSQL, and this role sits at the intersection of database internals, open source community work, and Supabase's managed Postgres platform. This role requires strong working overlap with Americas time zones due to production support and on-call responsibilities. What you'll work on New OrioleDB features. Design and implement new capabilities in OrioleDB — for example, native index access methods (GiST, GIN, HNSW for pgvector), disaster recovery tooling, and other storage-engine-level features that expand what OrioleDB can do. Stability. Strengthen OrioleDB's reliability through deeper test coverage, fault injection, crash and recovery testing, and improvements to CI infrastructure. Ensure regressions are caught early and that OrioleDB behaves predictably under stress, replication, and failure scenarios. Upstream collaboration. Contribute changes directly to PostgreSQL core. Part of OrioleDB lives as a patch on top of PostgreSQL — moving the right pieces upstream shrinks what we maintain ourselves and benefits the wider community. This happens through PostgreSQL's open development process: the pgsql-hackers mailing list, public code review, and commitfests. Supabase integration. Work with Supabase's Postgres team to ensure OrioleDB fits naturally into Supabase's managed offering and roadmap. You Will: Design, implement, and test new OrioleDB features and integrate them cleanly with PostgreSQL's planner, executor, and surrounding subsystems. Build out and maintain test infrastructure: regression suites, f
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We're looking for a Release Engineer (SRE) to join our Release Engineering team (part of EngOps) — a production-operations expert who brings an SRE mindset to how Supabase ships and runs, making deploys safe, observable, and recoverable at scale. Release Engineering's scope has grown well beyond build-and-ship: we increasingly own the operational reliability of the systems that deploy and run Supabase. In this role you'll treat our deployment pipelines, pre-production signal, and the control plane itself as production systems — with SLOs, error budgets, and on-call ownership — and you'll be the person teams lean on when reliability is on the line. This is not a "gatekeeper" role. You'll make the reliable path the easy path: standardising how we deploy, instrumenting what we ship, and ensuring that when something breaks, we detect it quickly and recover quickly. What You'll Be Responsible For In this role, you'll: Own the reliability of Supabase's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies) Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load Participate in on-call, lead blameless po
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We're looking for engineers to join the team building Supabase Lite , a lightweight, TypeScript-native implementation of Supabase. It runs interchangeably on SQLite and Postgres and ships a PostgREST and Auth compatible API, so applications written against @supabase/supabase-js work as-is. It exists because AI builders and platform partners need a database they can give to every prototype without the cost or the wait: sub-second provisioning, a footprint small enough to run inside a sandbox, and an upgrade path to full Supabase when an app graduates to production. This is a role with a lot of agency. Small team, working product, patterns still to be set - you'll own large areas end to end and make real decisions from day one. What You'll Be Responsible For In this role, you'll: Grow API compatibility with the Supabase client, keeping behavior faithful to hosted Supabase and backed by conformance tests Bring more of the Supabase stack to SQLite Build and harden the path for upgrading a project from Supabase Lite to full Supabase Own the developer experience end to end: the CLI, local tooling, and docs Build and harden the hosted offering on a scale-to-zero, near-zero-cost footing, kept loosely coupled so the implementation can be swapped without a rewrite Work directly with design partners and turn real-world migration and integration friction into a sharper product Engage with the open-source community and contribute back to the broader Supabase stack. You Might Be a Good Fit If You Have substantial backend or full-stack experience with strong fluency in TypeScript, and have built or contributed to a backend system, framework, or developer platform before Are a strong generalist,
Team Frontend within Supabase owns tooling for the marketing/docs website and the core product called “Studio” Its used by millions of developers around the world building insanely cool applications. We live and die by user feedback and relentlessly address issues every day to make the product better! What You’ll Own Ship products 0→1 and then earn the next version from users. You’ll take features from a blank editor to real users in production, then use what you learn (support threads, usage data, direct conversations) to decide what to build next. We want someone who has done this before: launched something, watched real people use it, and improved it on the strength of that feedback rather than a spec handed down from above. Own the quality pipeline, not just the feature. In a world where an agent writes a large share of the code, fast, the CI is what keeps that speed safe. You’ll build and tighten the layered checks that let us move quickly without breaking things: type safety at the boundaries, lint and format, unit and component tests, contract tests against our APIs, accessibility, visual regression, and bundle-size budgets, so a red check tells you (or the agent) exactly what to fix before anything reaches users. Make sure a bug caught once is caught forever. When something slips through, you don’t just patch it. You add the check that would have caught it, at the cheapest stage that can, so the same regression can’t happen twice. Work end-to-end and in the open. Frontend UI, the API contracts it depends on, preview deploys, observability. You’ll operate autonomously across the stack and do it publicly, in a codebase thousands of developers read and contribute to. What You’ll Bring A track record of taking a product from 0→1. You’ve launched something real, put it in front of users, and iterated based on how they actually used it, ideally something you can point us to. Deep frontend fundamentals with strong React. TypeScript and React, plus a real understand
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We’re looking for an EngProd Engineer to join our Engineering Operations team and own the engineering experience from local setup to production deploy. You’ll build and integrate the tools our engineers rely on every day: local development, code search, testing, code review, CI, AI-assisted workflows, and the metrics that show where time is being lost. You’ll work closely with other engineering teams at Supabase to eliminate friction, optimize cycle times, and make reliably shipping to production effortless. This role is ideal for someone who thrives in async, fast-paced environments and is excited about building developer tools that engineers love and that empower them to do their best work. What You’ll Be Responsible for In this role, you’ll: Improve local development and the daily workflow Own local development environment stability; eliminate “works on my machine,” keep environments reproducible across teams, and make onboarding and setup self-service for every engineer Improve the daily editing experience, including IDE and editor tooling, extensions, and code search, indexing, and navigation that stays fast as the codebase grows Build paved roads: project scaffolding, templates, and golden-path workflows that make the right thing the easy thing Speed up build, test, and CI Improve our testing tooling and make sure it works well for engineers, including test orchestration, ephemeral end-to-end environments, and techniques like mutation and chaos testing that catch what ordinary tests miss Profile build and test performance, optimize CI pipelines, parallelize and shard tests, cache aggressively, and hunt down the bottlenecks that leave engineers waiting Integrate and optimize
Get new production cleaning specialist jobs by email
Daily job updates · Unsubscribe anytime