ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Baseten’s Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost‑efficient. As part of this team, you’ll focus on Model APIs — the infrastructure powering our hosted API endpoints for the latest open‑source models. This work spans distributed systems, model serving, and developer experience. You’ll join a small, high‑impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale. RESPONSIBILITIES: Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations Productionize performa
Jobiba hiring network
Ai Platform Engineer Jobs
10,000 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current ai platform engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. At Baseten, we are building the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale, the network is the computer. We are looking for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building block in our infrastructure and unlocking the next generation of distributed inference optimizations. THE OPPORTUNITY Networking and compute are no longer separate disciplines; they are converging. The massive throughput of H100, B200, and NVL72 architectures enables and demands a new approach where communication is co-optimized alongside computation. We are entering an era where the network is an active accelerator, leveraging smart hardware offloads and direct interconnects to ensure that data movement operates at wire-speed. In this role, you will go beyond network configuration to architect the software fabric that unifies thousands of GPUs into a cohesive operating system. While you will leverage the best of the open-source ecosystem, you won't be limited by it. Where off-the-shelf solutions stop, you will build from scratch, engineering the primitives required to co-optimize communication and compute for Disaggregated Serving, Wide Expert Parallelism (WideEP), and lightening cold starts. WHAT YOU'LL DO Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack,
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg
Drata is building the trust layer between great companies - automating compliance, managing risk, and helping organizations prove trust continuously as they scale. We're Dratanauts: a global crew of 600+ professionals united by a culture that rewards integrity, ownership, and raising the bar, no matter where in the world we're working from. Why Join the Drata Team? At Drata, you're not maintaining legacy compliance software - you're building the agentic AI platform defining what trust looks like for the next generation of companies. Here's what makes the work itself worth showing up for: Problems without a playbook: You'll work at the edge of AI and security, building agentic governance, continuous compliance, and real-time trust verification to solve problems that don't have an established answer yet. You're writing it as you go. Real ownership, not just process: Our values center on owning outcomes and raising the bar, not checking boxes. You're expected to have opinions and back them. A seat at the table: Your perspective is unique and valued. Open debate and diverse viewpoints are built into how decisions actually get made here, at every level. Growth at rocketship speed: Drata is scaling fast, which means scope grows fast too. High performers get more ownership, visibility, and experience. A crew, not just coworkers: Dratanauts consistently describe a "come as you are" culture with sharp, curious people—the kind of team that makes hard problems genuinely fun to solve. See what they say here and follow us on LinkedIn for company news, employee stories, and career updates. Job Summary: Drata, at the vanguard of compliance software innovation and renowned for its commitment to trust and security across the internet, is on an ambitious path to redefine how AI and General AI technologies bolster compliance automation. Drata is seeking an Applied AI Engineer to drive the quality and effectiveness of our AI systems through rigorous research, experimentation, and evalu
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know Okta Okta is The World’s Identity Company. We free everyone to safely use any technology—anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth. At Okta, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box, we’re looking for lifelong learners and people who can make us better with their unique experiences. Join our team! We’re building a world where Identity belongs to you. The Engineering Opportunity We are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a technical leader within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud s
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know Okta Okta is The World’s Identity Company. We free everyone to safely use any technology—anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth. At Okta, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box, we’re looking for lifelong learners and people who can make us better with their unique experiences. Join our team! We’re building a world where Identity belongs to you. The Engineering Opportunity We are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a technical leader within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud s
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Staff Software Engineer within the Security Platform Engineering group, you'll serve as the primary execution anchor for our most critical regulatory and cryptographic initiatives. This team owns Coinbase's Tier-0 Multi-Party Computation (MPC) engine, which secures 99% of assets under custody. You'll lead the development of high-availability services that directly protect customer assets and ensure global regulatory compliance from our Israel office. What you'll do: Lead the architecture and execution of critical cryptographic infrastructure, including services required for global regulatory compliance Define technical strategies with a multi-quarter horizon, bringing clarity to highly ambiguous requirements for novel cryptographic solutions Own risky and ambitious initiatives, implementing secure and resilient cryptographic protocols in a high-stakes, zero-trust environment Shape engineering quality practices as a technical anchor and mentor within a team of senior engineers, driving rigorous code reviews and eliminating single points of failure Build alignment and secure commitment from cross-functional partners across Security, Product, and Policy to drive flawless execution Required Skills and Experience: 8+ years of experience in software engineering with demonstrated success building highly available, distributed systems Deep domain expertise in applied cr
About the Team IT Systems Operations serves as the operational control layer connecting Security, Engineering, and Employee Technology platforms. The team ensures employee-facing systems, identity workflows, and enterprise applications operate in a predictable, governed, and continuously reliable manner as the organization scales. Beyond implementation, the team establishes structured operating patterns that ensure platform changes, access models, integrations, and lifecycle workflows evolve safely and consistently across the enterprise environment. About the Role This role acts as an operational owner for identity-connected enterprise SaaS platforms and system change controls supporting OpenAI’s compliance requirements. The engineer will be responsible for ensuring that production configuration, access models, and platform changes remain compliant, auditable, and consistently operated through defined controls. You will own and operationalize controlled system and workflow changes across identity platforms, SaaS applications, collaboration tooling, and enterprise infrastructure. You will partner closely with Security, Platform Engineering, and IT Support Operations to: Ensure identity and access workflows behave consistently across systems Implement structured rollout and configuration practices for enterprise applications Improve visibility and traceability of system changes impacting employee workflows Translate operational requirements into durable automation and policy-aligned implementations Success in this role requires not only strong engineering capability, but also sound operational judgment, disciplined documentation practices, and effective cross-functional collaboration. In this role, you will: Enterprise SaaS & Identity Platform Ownership Own administration and operational stewardship of enterprise SaaS and identity-connected platforms, ensuring configuration integrity, access governance, and compliance with defined control requirements. Own onboard
About the Job: As a Senior Sales Engineer , you will be the technical partner to Account Executives, helping LaunchDarkly land and expand enterprise customers across the Indian market. You will guide technical discovery, deliver compelling demonstrations, lead proof-of-value engagements, and help customers understand how LaunchDarkly enables safer, faster software delivery. This role combines deep technical expertise with strong business acumen and the ability to translate complex technology into real business outcomes. You’ll work closely with engineering leaders, architects, and DevOps teams at large enterprises to help them modernize their software delivery practices. Responsibilities: Partner with Sales to Win Enterprise Deals Work closely with Account Executives to qualify opportunities and develop technical account strategies. Help customers understand how feature management and progressive delivery can transform software delivery. Participate in complex enterprise sales cycles and help secure technical wins. Lead Technical Discovery and Solution Design Engage with engineering leaders, platform teams, and architects to understand their software delivery challenges. Design solutions that demonstrate how LaunchDarkly integrates into modern development environments. Align technical solutions with customer business objectives. Deliver High-Impact Demos and Proof of Value Conduct compelling product demonstrations that highlight real customer value. Lead technical evaluations including proof-of-concepts, workshops, and hands-on trials. Guide customers through the evaluation process and ensure successful outcomes. Build Champions and Trusted Relationships Develop strong relationships with both technical and business stakeholders. Enable customer champions who can advocate for LaunchDarkly internally. Support customers throughout the evaluation process and beyond. Be a Technical Thought Leader Stay up to date on trends in DevOps, CI/CD, feature management, AI-driven d
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our diverse team of technologists have developed a high performance RISC-V-based CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a talented engineer to build and run the performance infrastructure shared by our RISC-V Software and RISC-V Performance teams. This is the plumbing that both teams' performance work stands on — benchmarking automation, data collection, and workload capture across silicon, FPGA/emulation platforms and performance models. You'll work across both teams, enabling performance and software engineers to spend their time on analysis and optimization instead of running experiments by hand. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Great at identifying problems and developing solutions, with a bias toward owning the systems you build. Enjoys building tools and optimizing workflows so other engineers can move faster. Strong Linux systems engineer, c
The Forward Deployed Engineer, Staff (FDE) is a high impact technical leader responsible for translating the immense power of Talkdesk's Agentic AI platform into transformative, production grade solutions for our strategic enterprise customers. You are a unique blend of a highly experienced software engineer, a technical project lead, and a hands-on builder. You own the technical execution for deployments, mitigate technical risks, and drive successful customer outcomes across assigned projects. This role is for the experienced builder who excels in high stakes, customer facing environments and is passionate about defining the future of Customer Experience Automation. Responsibilities Technical Project Leadership & Architecture: Act as the technical lead for complex deployments. Design and implement the architectural blueprint for AI agent solutions, manage cross-system dependencies, and ensure designs meet stringent enterprise standards for security and scale. Hands-On Engineering & Delivery: Write production-grade code and leverage Talkdesk and 3rd-party APIs/SDKs to design, build, test, and deploy AI agents. Drive the execution from prototype through to production deployment, ensuring technical quality. Technical Consultation & Alignment: Serve as a trusted technical expert for our AI solutions. Confidently address deep technical inquiries, mitigate technical risk, and build trust with customer Engineering Directors, and technical architects. Influence Product & Engineering Roadmap: Synthesize and codify deployment learnings into reusable solution patterns and tooling. Provide actionable feedback to core Product and Engineering teams to help inform future product direction. Technical Guidance: Mentor junior FDEs and technical specialists on best practices for complex AI architecture, production quality, and client-facing technical delivery. Who You Are We are looking for an autonomous, results-driven technical leader who thrives at the intersectio
About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. Role Overview We are looking for domain-first, systems-oriented engineers who understand how real-world systems (in renewable energy , automotive, manufacturing, IoT, etc.) generate and use data — and can translate that into meaningful AI-driven workflows using DevRev. As a Solutions Engineer, you will act as a trusted advisor , combining domain expertise, engineering depth, and problem-solving to help customers unlock value from their data through DevRev’s AI platform. You will work closely with Sales, Product, and Engineering to design solutions that go beyond demos — enabling real-world automation, intelligence, and agentic workflows . Key Responsibilities Partner with customers to deeply understand domain workflows, systems, and data flows (e.g., connected vehicles, factory systems, IoT environments) Translate real-world signals (sensor data, logs, events) into insights, workflows, and automated actions Build and showcase AI-driven workflows, automations, and agentic use cases Design and articulate end-to-end solutions (data ingestion → processing → decision → action) Act as a technical advisor , guiding customers on how DevRev c
Senior Container Security Engineer – CVE Remediation & Image Hardening About the Role We are looking for a hands-on Senior Container Security Engineer to lead vulnerability remediation and image hardening across Linux-based container environments. This role focuses on deep operating system and container security engineering rather than simple vulnerability scanning. You will analyze, remediate, rebuild, harden, and continuously optimize container images used in modern cloud-native platforms. You will work closely with platform engineering, DevOps, infrastructure, and security teams to build automated remediation pipelines, reduce the attack surface, and deliver production-ready hardened images. What You’ll Do - Own end-to-end CVE remediation across Linux-based container images. - Analyze vulnerabilities across OS packages, libraries, runtimes, and dependencies. - Patch, rebuild, validate, and maintain hardened container images at scale. - Reduce attack surface by removing unnecessary packages, binaries, services, and dependencies. - Build and scale automated remediation pipelines for continuous image patching. - Improve image security posture while minimizing operational disruption. - Generate, validate, and maintain SBOMs to support supply chain visibility and compliance. - Integrate remediation workflows into CI/CD and GitOps pipelines. - Optimize image size, startup performance, and operational efficiency. - Research emerging Linux, container, Kubernetes, and software supply chain threats. - Troubleshoot complex dependency, package compatibility, and runtime security issues. - Help define internal standards for hardened images and secure software delivery. What You Bring - 5+ years of experience in Linux systems engineering, platform engineering, DevSecOps, security engineering, or SRE. - Deep understanding of Linux distributions (Debian, Ubuntu, Alpine, RHEL). - Strong hands-on experience with Docker, Kubernetes, and
Get new ai platform engineer jobs by email
Daily job updates · Unsubscribe anytime