Jobiba hiring network

Edge Infrastructure Engineer Jobs

1,807 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current edge infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto

pythonci/cdgit
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re hiring a Data Engineer to build and scale Baseten’s internal data platform. This role sits at the intersection of data engineering, analytics, and data science, transforming raw product and business data into reliable datasets that power decision-making. You’ll design the data models, pipelines, and analytics infrastructure that enable teams across Product, Engineering, Finance, Marketing, and Sales to understand usage and performance. This includes working with AI inference, infrastructure, and observability data to generate insights about the product, business operations and platform economics. You’ll partner closely with stakeholders to build robust, scalable pipelines, define company-wide metrics that inform strategy and planning. RESPONSIBILITIES Design and maintain core data models and semantic layers Develop and orchestrate batch and streaming data pipelines using technologies such as Apache Beam, Kafka, Airflow, or similar frameworks Analyze inference and infrastructure telemetry , including data from OpenTelemetry, Grafana, and other observability tools Define and maintain company-wide metrics across product usage, performance, and customer lifecycle Enable self-service analytics through agents and tools, with well-structured semantic layers and context Ensure data reliability and quality through testing, documentation, and governance PREFERRED QUALIFICATIONS Understanding of inference metrics s

machine learningaigo
View job →

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. At Baseten, we are building the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale, the network is the computer. We are looking for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building block in our infrastructure and unlocking the next generation of distributed inference optimizations. THE OPPORTUNITY Networking and compute are no longer separate disciplines; they are converging. The massive throughput of H100, B200, and NVL72 architectures enables and demands a new approach where communication is co-optimized alongside computation. We are entering an era where the network is an active accelerator, leveraging smart hardware offloads and direct interconnects to ensure that data movement operates at wire-speed. In this role, you will go beyond network configuration to architect the software fabric that unifies thousands of GPUs into a cohesive operating system. While you will leverage the best of the open-source ecosystem, you won't be limited by it. Where off-the-shelf solutions stop, you will build from scratch, engineering the primitives required to co-optimize communication and compute for Disaggregated Serving, Wide Expert Parallelism (WideEP), and lightening cold starts. WHAT YOU'LL DO Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack,

pythonkubernetesmachine learning
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE The largest, most demanding enterprises are starting to run on Baseten, and they arrive with a range of security, compliance, and procurement requirements. As a Senior Engineer on Baseten's enterprise engineering team, you'll build the capabilities that enable large organizations like Writer, HubSpot, and Notion to succeed on Baseten. Enterprise engineering authors the core building blocks, APIs, and user experiences powering the Baseten platform: identity and access management, billing, regional isolation, and self-hosted and single-tenant deployment options. This is deep product and systems work across the full stack, from designing authentication and authorization systems using standards like OAuth and OIDC to shipping the admin experiences enterprise IT teams use to manage their organization. EXAMPLE INITIATIVES Recent and upcoming work on the team: Fine-grained authorization for users, service accounts, and agentic workloads SSO and SCIM support, allowing customers to centralize and automate access to Baseten Expanding the billing platform to support evolving pricing models, advanced data exports, and controls to manage spend In-product management and enforcement of customer compliance requirements like data residency and HIPAA Securing network paths in and out of a customer's models with private connectivity and ingress and egress restrictions Allowing customers to run Baseten inside their own VPC, on-pr

kubernetesrestmachine learning
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: As a Software Engineer at Baseten, you will own one of the most critical surfaces of our business: pricing, billing, and revenue infrastructure. As we launch more and more products— billing is no longer just operational plumbing. It is a strategic lever for growth. This role will establish clear ownership of billing as a function and create leverage for Finance, Sales, and GTM teams while maintaining a seamless customer experience. RESPONSIBILITIES: Own Baseten’s end-to-end billing and revenue infrastructure, including pricing, invoicing, metering, and reporting foundations. Build and evolve our billing platform and integrations (including Orb), ensuring correctness, auditability, and a high-trust experience for customers and internal teams. Partner closely with Finance, Sales, GTM, and Forward Deployed Engineering to turn real-world workflows into reliable internal tooling and automation (quoting, approvals, renewals, usage reconciliation, revenue reporting). Design systems that scale with new products, packaging, and go-to-market motions, making billing a strategic lever for growth. Drive reliability and operational excellence for revenue-critical workflows: monitoring, alerting, incident response, backfills, and clear runbooks. Lead from the front on high-impact projects: clarify requirements, propose crisp technical approaches, ship iteratively, and raise the bar on quality and velocity. Debug and resolve

machine learningaigo
View job →
D
Drata
📍 San Francisco• Full-time• $174.5K – $236.1K/yr
1mo ago

Drata is building the trust layer between great companies - automating compliance, managing risk, and helping organizations prove trust continuously as they scale. We're Dratanauts: a global crew of 600+ professionals united by a culture that rewards integrity, ownership, and raising the bar, no matter where in the world we're working from. Why Join the Drata Team? At Drata, you're not maintaining legacy compliance software - you're building the agentic AI platform defining what trust looks like for the next generation of companies. Here's what makes the work itself worth showing up for: Problems without a playbook: You'll work at the edge of AI and security, building agentic governance, continuous compliance, and real-time trust verification to solve problems that don't have an established answer yet. You're writing it as you go. Real ownership, not just process: Our values center on owning outcomes and raising the bar, not checking boxes. You're expected to have opinions and back them. A seat at the table: Your perspective is unique and valued. Open debate and diverse viewpoints are built into how decisions actually get made here, at every level. Growth at rocketship speed: Drata is scaling fast, which means scope grows fast too. High performers get more ownership, visibility, and experience. A crew, not just coworkers: Dratanauts consistently describe a "come as you are" culture with sharp, curious people—the kind of team that makes hard problems genuinely fun to solve. See what they say here and follow us on LinkedIn for company news, employee stories, and career updates. Job Summary: Drata's Identity & Access Management team owns the identity, authentication, and access control infrastructure that every customer uses to access the platform — and that every internal platform service relies on for trust boundaries. Authentication — SSO (SAML 2.0, OIDC), session/token management, MFA. We're focused on authentication for enterprise customers — large user popula

typescriptnodejsaws
View job →
D
Drata
📍 San Francisco• Full-time• $145K – $196.4K/yr
1mo ago

Drata is building the trust layer between great companies - automating compliance, managing risk, and helping organizations prove trust continuously as they scale. We're Dratanauts: a global crew of 600+ professionals united by a culture that rewards integrity, ownership, and raising the bar, no matter where in the world we're working from. Why Join the Drata Team? At Drata, you're not maintaining legacy compliance software - you're building the agentic AI platform defining what trust looks like for the next generation of companies. Here's what makes the work itself worth showing up for: Problems without a playbook: You'll work at the edge of AI and security, building agentic governance, continuous compliance, and real-time trust verification to solve problems that don't have an established answer yet. You're writing it as you go. Real ownership, not just process: Our values center on owning outcomes and raising the bar, not checking boxes. You're expected to have opinions and back them. A seat at the table: Your perspective is unique and valued. Open debate and diverse viewpoints are built into how decisions actually get made here, at every level. Growth at rocketship speed: Drata is scaling fast, which means scope grows fast too. High performers get more ownership, visibility, and experience. A crew, not just coworkers: Dratanauts consistently describe a "come as you are" culture with sharp, curious people—the kind of team that makes hard problems genuinely fun to solve. See what they say here and follow us on LinkedIn for company news, employee stories, and career updates. Job Summary: Drata is reimagining compliance as an intelligent, always-on experience — and AI is at the center of that vision. We are seeking a Senior AI Product Engineer to own the full-stack development of customer-facing AI features, embedded directly within our product teams. This is not a platform or infrastructure role. You'll translate the capabilities of LLMs, agents, and RAG pipelines

typescriptpythonreact
View job →
V
1mo ago

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta’s Developer Experience team builds the tools engineers use every day to bring ideas to production rapidly and reliably. You’ll empower other Vanta engineers to leverage cutting-edge technologies and best practices to make Vanta more performant and scalable on a platform level. Example projects include modernizing our CI/CD pipelines, introducing new test frameworks, launching AI-powered dev tools, and scaling developer environments to support a growing engineering team. This team has a wide breadth of impact across all of product engineering. The work we do compounds in value by making it easier for engineers to diagnose and solve bugs, streamline workflows, and ship value to our customers quickly and safely. Vanta engineers design and develop new product functionality and infrastructure leveraging modern frameworks and tooling, including TypeScript, React, Node.js, MongoDB, Github Actions, and various AWS services such as Fargate and ECS. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. We’d love for you to join us! You will: Set direction for critical dev infrastructure, enabling us to stay ahead of continued rapid growth Design and build CI and build systems that ensure Vanta engineers can develop and ship robust products quickly and confidently Improve the efficiency and reliability of our deployment workflows, including tools for hotfixes, rollbacks, and incident mitigation Lead development of tools that accelerate feedback loops — from typechecking and linting to running tests and deploying changes Build and maintain scalable developmen

typescriptreactnode.js
View job →

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We are looking for a Software Engineer: IaC Platform Experience to join our Interfaces team and own the Terraform provider as a core part of Supabase's developer platform. This is a hands-on engineering role focused on the Go codebase behind the Supabase Terraform provider. You will partner with product and engineering leadership on roadmap priorities, drive technical execution, and ensure we ship a reliable, predictable, and well-documented Terraform experience for developers at scale. You will focus on resource behavior, lifecycle correctness, schema evolution, upgrade safety, and practical migration paths for existing users. This role is ideal for someone who thrives in async, fast-paced environments and enjoys building practical platform primitives that millions of developers can rely on. What You’ll Own Own the Go Terraform provider codebase, including architecture, implementation quality, test strategy, and release readiness. Improve Terraform provider reliability and ergonomics, including resource behavior, data sources, lifecycle edge cases, and upgrade safety. Drive technical strategy for IaC workflows through design docs, RFCs, and iterative delivery. Build practical migration and interoperability paths for existing Terraform users. Partner with product and engineering leadership in a shared roadmap model to define priorities, scope, and outcomes. Monitor customer feedback, OSS issues, and usage signals to continuously improve the Terraform experience. Create clear documentation and examples that make IaC workflows easier to understand and adopt. What You Bring 5+ years of software engineering experience in developer platforms, infrastructure tooling, or distributed sys

typescriptci/cdgit
View job →
S
Supabase
📍 Remote• Full-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the role Auth , written in Go (server) and with client libraries for TypeScript , SSR and for other frameworks and technologies, is one of the most popular products in the Supabase stack. We are seeking someone to help us build new and maintain existing Auth features. What you’ll be responsible for Designing and implementing secure, scalable authentication features in Go and TypeScript. Working across the stack: from server-side protocols to client-side libraries for frameworks like Next.js. Owning the performance, reliability, and scalability of the Auth server across Supabase's infrastructure. Contributing to the evolution of our Auth architecture, including support for OAuth, OIDC, SAML, and other protocols. Planning and executing safe database migrations across a large fleet of Postgres instances. Building and improving observability: metrics, tracing, alerting, and dashboards to keep the system healthy at scale. Writing and reviewing RFCs as part of our product development process. Collaborating with engineers across Supabase to ensure a seamless experience for developers using our tools. Supporting the community and responding to developer feedback on GitHub, Discord, and other channels. You might be a good fit if you (Required) Have 4+ years of professional experience writing and shipping Go in production. (Required) Have 2+ years of professional experience working on an authentication system (implementing protocol support, maintenance at scale). (Required) Have strong relational database experience (Postgres or MySQL); Postgres experience is a bonus. Have strong knowledge of TypeScript in addition to Go (languages used daily). Have strong knowledge of web technology fundamentals (

javascripttypescriptjava
View job →
S
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About this role We’re hiring experienced performance engineers to make performance measurement at Supabase rigorous, repeatable, and actionable by product teams. This role owns the benchmarking and profiling systems we use to detect regressions, explain performance changes, and publish credible numbers for customers and the market. You’ll build the tooling and the methodology and make it easy for product teams to use it correctly. You’d join a new performance team as one of its first members, working alongside a deeply seasoned performance engineer, with the team growing to roughly six this year. There’s no legacy process to inherit — you’ll help define how we measure and communicate performance metrics at Supabase. What you’ll do Build and evolve benchmarking, profiling, and load-testing tooling (cloud, database, and end-to-end). Define benchmark suites and reference workloads that reflect real customer behavior (and evolve them as the products change). Establish performance baselines and regression gates across products (variance tracking, environment control, and detection of changes). Build mechanisms that turn benchmark results into action: triage workflows, owner routing, and clear, reproducible reports. Partner closely with database teams (e.g. Multigres, OrioleDB) to identify bottlenecks and land concrete performance improvements. Enable the company to publish credible, repeatable performance numbers for customers and the market. Help teams self-serve performance testing and make performance a first-class engineering concern. Problems you might work on Build our continuous benchmarking infrastructure so product teams can spot performance regressions early. Ideally, teams can run benchma

aigorust
View job →
S
1mo ago

Supabase is the open-source Postgres development platform that 7M+ developers and thousands of enterprises depend on every day. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the role We’re hiring a Product Manager to own the platform primitives every Supabase product runs on, including internal features like compute , disks , networking, and the API gateway, and external features like read replicas , custom domains , PrivateLink , and Bring Your Own Cloud. What you'll be responsible for: Talk to customers across the full spectrum. Indie developers running a single nano project, fast-growing startups whose costs are dominated by compute and disk, enterprises walking through a network-architecture review, and partners building on top of Supabase. Find the real blockers and bring them back to the roadmap. Own the problem statement and requirements behind every platform bet. Capture the customer evidence behind each decision, name the cost, capacity, and reliability constraints, and give the team a target it can hit. Decide what gets built, what gets deferred, and what gets cut. Every quarter you're choosing between enterprise unlocks blocking deals, reliability and cost wins for the long tail of projects, and net-new capabilities that change what Supabase can run. Set the priorities and defend them. Define how each launch is measured before it ships. Set the metric, agree on the threshold, and track it after launch. Know whether a feature moved enterprise deal velocity, project economics, or platform reliability. Use that to sharpen the next call. Keep engineering, design, and leadership aligned. The platform touches every other Supabase product, every region, and every customer tier. Write the roadmap, surface dependencies before they become blockers, and keep decisions moving. You might be a good fit if you: Have 7+ years of produ

aigorust
View job →
C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! As a Senior Security Operations Engineer you will: Serve as trusted advisor to team’s leadership and partner teams by clearly articulating business risks associated with security issues Harden our cloud-native environments (AWS, OCI, GCP) by introducing secure by default designs and features into network, tooling, and processes Own and drive resolutions for enabling engineers to design, build, and use infrastructure securely at scale by deploying secure architectures using infrastructure-as-code and reusable code libraries Manage IAM / RBAC for cloud infrastructure, and partner with IT on streamling authentication/authorization to ensure unified access control across the board Deploy and operationalize some of the security services and tools (eg: SIEM, SOAR, domain monitoring, endpoint tooling, cloud security tooling) Respond to security incidents and harden environments post-incidents. Support control monitoring and remediation for compliance initiatives Gather and analyze security metrics to address security issues with cross-team dependencies Be a problem solver who is empathetic to developer concerns and will employ construc

pythonawsazure
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! The integration team is responsible for developing and scaling machine learning algorithms and infrastructure for LLM post-training, with a focus on large-scale, distributed RL methods. We strive for excellence in both engineering and science by meticulously designing experiments and design docs. While tasks are assigned according to everyone’s expertise, there is a global team effort to write production code and support the team research efforts, depending on individual interests and organizational needs. In particular, this role aims to enhance the global quality of the post-training codebase by implementing new tools to ease and support research, optimizing post-training algorithms, and scaling distributed RL to unprecedented levels. Please Note: We have offices in London, Paris, Toronto, San Francisco, New York but we are also remote-friendly! Applicants for this role may work anywhere between UTC−06:00 and UTC+01:00. As a Member of Technical Staff, you will: Design and write high-performing and scalable software for training models. Develop new tools to support and accelerate research and LLM training. Coordinate with other

pythonkubernetesgit
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! We’re looking for a senior engineer to help build, maintain and evolve the training framework that powers our frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure. You will design and maintain the core components that enable fast, reliable, and scalable model training — and build the tooling that connects research ideas to thousands of GPUs. If you enjoy working across the full stack of ML systems, this role gives you the opportunity and autonomy to have massive impact. What You’ll Work On Build and own the training framework responsible for large-scale LLM training. Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing). Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100). Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics. Collaborate closely with infra teams to ensure our cluster, container environments, and hardware configurations support high-performance training. Investigate and res

dockerkubernetesgit
View job →
🔔

Get new edge infrastructure engineer jobs by email

Daily job updates · Unsubscribe anytime