At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Cortex team is building the future of AI for enterprise data. This role focuses on the Search infrastructure that powers our flagship products like CoWork, Cortex Code & Cortex Agents fast, reliable, scalable and secure at the enterprise level. You will be building high-performance retrieval engines (leveraging vector search, hybrid search, and semantic indexing) that power Snowflake Cortex. This involves optimizing how billions of rows of data are indexed and retrieved in milliseconds. What you will do in this role: Architect Agentic Runtimes: Build and scale the orchestration engines that execute complex agentic workflows, ensuring low-latency tool execution and robust state management. Scale Context Engineering Infra: Design high-performance systems for RAG (Retrieval-Augmented Generation), including vector database integration, scalable and efficient search indexing, query processing, and result ranking, semantic caching, and automated metadata extraction. Build the "Evals Engine": Develop the automated infrastructure required to run massive-scale golden set simulations, error analysis pipelines, and "hillclimbing" experiments. Productionize AI Workflows: Collaborate with the modeling team to take raw LLM capabilities and turn them into hardened, multi-tenant mi
Jobiba hiring network
Software Engineer Infrastructure Jobs
6,326 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current software engineer infrastructure jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Privacy Infrastructure team, you’ll design and build large-scale systems that operate across Roblox’s platform and data ecosystem. You’ll tackle complex distributed systems challenges, develop reliable and scalable infrastructure, and work across engineering teams to deliver foundational capabilities that serve millions of users. This is an opportunity to take on high-impact technical problems at Roblox scale while helping shape the next generation of our infrastructure. You Will Design, build, and scale reliable infrastructure and platform solutions that protect user data and support the needs of a global platform. Develop foundational capabilities for emerging technologies, including agentic AI workflows, with a focus on safe and responsible data access. Partner with engineering teams across Roblox to integrate scalable data protection capabilities into their systems and development workflows Drive technical strategy and architecture for complex, cross-cutting challenges spanning data, infrastructure, privacy, and security. You Have 7+ years of proven experience as a software engineer. Expertise in Python or Golang; Strong understanding of system
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale is seeking a Staff Software Engineer to lead the technical vision for our Infrastructure team. As a Staff Engineer, you will be responsible for the architectural evolution of our control plane and data plane, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. You will act as a force multiplier, setting the standards for Kubernetes-based cloud-native infrastructure while mentoring engineers and driving cross-functional alignment across the Ray open-source community and our proprietary product teams. Key Responsibilities Architectural Leadership: Define and drive the multi-year technical roadmap for services that orchestrate Ray clusters across diverse cloud and on-premises environments. Systemic Optimization: Lead the design and optimization of high-performance control plane components specifically tailored for large-scale, heterogeneous AI/ML workloads. Platform Reliability: Establish the organization-wide standards for the reliability, scalability, and observability of Anyscale-managed infrastructure. Strategic Integration: Direct the long-term strategy for accelerator integration (GPUs, TPUs) and container management to ens
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the role. We’re building the next generation of agentic AI infrastructure at Cohere. This team sits at the intersection of ML systems, distributed infrastructure, and developer experience, creating the platform that powers autonomous AI agents at scale. You’ll work on hard, forward-looking problems with few established patterns, including secure code execution, agent state management, model routing, identity and authentication, and resource management for long-running agent workflows. This role is a strong fit for someone who combines systems depth with ML intuition. You should be comfortable building reliable infrastructure, thinking through distributed systems tradeoffs, and understanding how emerging agentic capabilities shape platform design. What you’ll work on. Secure execution environments for agent-generated code Identity, authentication, and trust boundaries for agents Model routing and orchestration across different model types and environments Rate limiting, quotas, and resource management for agent workflows State management, memory, and filesystem abstractions for agents. In this role you will: Turn emerging M
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides tools that help developers find and fix issues in their applications. The Developer Infrastructure team owns the systems that make every engineer at Sentry more effective at shipping high quality software: developer environments, CI/CD for our open source and closed source codebases, the golden path for building new services, our SDK and library publishing tooling, and the metrics and dashboards that keep it all healthy. Our goal is straightforward: developers should spend their time thinking about the software they build for customers, not the software they use to build it. That means everything from local and cloud development environments, to the CI/CD pipeline that gets code out safely and quickly, to the tooling that catches flaky tests before they cost someone a day. As AI coding agents become part of how engineers work, we're also investing in making sure our environments, CI, and deployment systems support that shift. As a Software Engineer on Dev Infra, you'll help build and scale this infrastructure end to end. In this role, you will Build and maintain the tooling that powers local and cloud-based developer environments Improve CI for our codebases and CD for our deployment pipeline, keeping both fast and reliable as the org and its infrastructure grow Define and evolve the golden path for building new services and libraries at Sentry Contribute to our SDK and library publishing tooling and release processes Build the metrics, dashboards, and flaky test detection that give the org visibility into deployment health and CI reliability Build tooling that gives engineers, and the AI codin
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role: This is a Principal Product Engineering role focused on Money Infrastructure at Replit. You’ll work on the financial backbone that powers how Replit earns money, how builders earn money, and how Agents transact. This role sits at the intersection of engineering, product, and the business. The systems you build directly impact revenue, trust, and some of the most critical user journeys on the platform. Getting them right enables growth, experimentation, and global scale. Getting them wrong creates broken payments, confusing pricing, and lost trust. We’re looking for engineers who can design and scale reliable financial systems while translating complex monetary logic into intuitive, user-friendly experiences for both Replit customers and builders on the platform. We love folks who have a passion for monetizing innovation and being a part of the greater pricing story. You will: Lead the design, architecture, and implementation of Replit’s core money infrastructure, spanning pricing, billing, payments, and monetization. Own and scale the global order-to-cash foundation supporting credit-based subscriptions, usage-based billing, marketplaces, in-app payments, and commerce for Agents. Enable rapid pricing and packaging experimentation across the company by building flexible abstractions and APIs for new SKUs, plans, and monetization models. Build high-converting, localized payment experiences across geographies — thinking globally while enabling users to pay locally. Power builder monetization by creating payment rails for apps, Agents, subscriptions, and new monetization primitives. Partner closely on data specifications with finance, accounting, and data teams to produce accurate, auditable, and reliable f
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role The Autonomy ML Infrastructure team is responsible for building & improving the core infrastructure for autonomy teams at Nuro. In this role, you will work closely with teams across Nuro, to design, build and deploy core infrastructure components in machine learning model life cycle, to push the autonomous future forward. You will have an opportunity to work across the full stack of machine learning solutions - from designing robust and scalable model pipelines to building to deploying the optimized models on Nuro’s fleet of self-driving robots! About the Work Optimize Nuro’s autonomy stack with cutting-edge optimization techniques like quantization, distillation, and model compression. Work with autonomy engineers to optimize, validate, and deploy large language models. Develop and maintain a world-class model compiler framework, FTL . Write robust, high quality software to increase our confidence in our vehicle
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations | Austin, US OR Seattle, US About the Role Emerging Technologies & Incubation (ETI) is where new and bold products are built and released within Cloudflare. Rather than being constrained by the structures which make Cloudflare a massively successful business, we are able to leverage them to deliver entirely new tools and products to our customers. Cloudflare’s edge and network make it possible to solve problems at
We are dedicated to ensuring proactive elimination of entire classes of security risk by engineering the core libraries, platforms and frameworks that provide secure guardrails for all Asanas. We are looking for a Senior Software Engineer to join our new Security Development team. This team is focused on building durable, secure-by-default solutions to protect Asana's infrastructure and product. Instead of acting as gatekeepers, we build guardrails that empower engineering teams to move quickly and safely. You will be responsible for engineering preventative controls at scale, focusing on building platforms and frameworks that eradicate systemic risks. This role is based in our San Francisco office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve: Design, build, and maintain secure-by-default frameworks, libraries, and platforms to eliminate entire classes of vulnerabilities. Engineer and improve core security services, including our access control frameworks, secrets management infrastructure, and AWS permissions systems. Develop and own the platform for vulnerability remediation, creating tooling that empowers engineering teams to address risks efficiently. Partner with product and infrastructure teams to architect and implement foundational security controls for Asana including cloud networking, and compute infrastructure. Partner with product teams to effectively offer recommendations for how to secure projects at all phases of implementation (design, development, launch, and/or incidents) Influence engineering initiatives through design reviews, communicating security principles, and helping teams make sound security trad
Senior Software Engineer, Developer Infrastructure Reykjavík The Developer Infrastructure team builds the tooling and systems our Asana engineers use every day to bring their ideas to production quickly and reliably. We build and operate the software that drives Asana’s roadmap. Each day, we combine industry best practices and innovation to support this product-focused company. We’re looking for an experienced Software Engineer with a passion for developer infrastructure. You will work with a world-class team of engineers on deploying and operating existing developer tooling, and building new tools to support our global, growing development team. You will have a unique opportunity to design and develop the systems and applications that drive the Asana development experience, lead complex technical projects, and work on cross-functional initiatives to help define the future of software engineering at Asana. This role is based in our Reykjavík office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Innovate the architecture of our development sandboxes to accelerate development iteration. Modernize our build system, tighten iteration loops, and polish frequent cycles in the developer experience. Lead complex developer infrastructure projects from technical design through implementation, rollout, and operation. Partner with engineering teams to ide
Join the Atlas Search team to design and develop the next generation of Semantic and Vector Search infrastructure. Atlas Search is a growing cloud service that allows users to execute complex search and vector search queries using the MongoDB Query Language. Our users can focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team is building a cloud-based distributed system responsible for the core components of search including data ingestion, performance, query language, query execution, for both relevance-based search and vector search. Our product is being adopted quickly and there are many interesting projects. This is a technical role where you will be responsible for the infrastructure and features enabling our at-scale cloud service powering vector and semantic search. We are looking to speak to candidates who are based in the San Francisco Bay Area for our hybrid working model. What You’ll Do Lead complex projects across the MongoDB ecosystem, for instance, development of a new Search deployment framework within the MongoDB managed cloud Set project level strategy, architect features, and lead projects to successful execution Identify, design, and implement features enhancing our reliability, performance, security and efficiency Perform code reviews with peers and make recommendations on how to improve our software development processes Influence and grow team members through active mentoring and leading by example What We Look For 5+ years experience in data management/search systems, ideally with a strong distributed systems and infrastructure background Experienced in the development and maintenance of concurrent, stateful services Eager to shape the technological direction of a complex system and have the ability to lead initiatives through collaboration with others Experienced in writing features, debugging and optimizing multithreaded applications written in Java Familiarity with LLM
We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls and services within MongoDB Atlas multi-cloud infrastructure. The team sits within the Site Reliability Engineering organization and works with other engineering teams to ensure that our infrastructure adheres to the highest security standards. This role can be based out of our New York City, Austin, Seattle or San Francisco offices, or work fully remotely on standard East Coast business hours. Responsibilities: Design and build core security primitives and services that protect MongoDB Atlas compute, networking, and identity across AWS, Azure, and GCP Build secure-by-default infrastructure using Linux security mechanisms (AppArmor, SELinux, seccomp, cgroups), Kubernetes, and eBPF to enforce runtime policies and gain deep visibility into systems behaviour Develop APIs, automation, and tooling that manage security posture at scale (CSPM, vulnerability management, workload identity) and provide monitoring, logging, and alerting pipelines that integrate with our tooling (Grafana, Splunk, Victoria Metrics.) Integrate security into our CI/CD and infrastructure-as-code workflows (Terraform) so that security controls are versioned, reviewed, and deployed just like any other code Lead complex projects end‑to‑end, from problem discovery and design docs to implementation, rollout, and long‑term ownership Collaborate with SRE, platform and product engineering teams to define secure architectures for new infrastructure and services Qualifications: You might be a great fit if you match some of the following: 5+ years of experience in Software Engineering, Site Reliability Engineering, or similar roles, preferably with relevant security work Proficiency with at least one programming language (Java, Golang, Rust, Python, or C/C++) and experience with infrastructure-as-code tools (Terraform) to automate security configurations
Get new software engineer infrastructure jobs by email
Daily job updates · Unsubscribe anytime