GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As an Assigned Support Engineer, you’ll be a trusted technical advisor to GitLab’s largest Self-managed, GitLab Dedicated, and GitLab.com customers, helping them avoid operational disruption and get the most from GitLab. You’ll combine deep Linux systems expertise, GitLab and CI/CD knowledge, and a proactive support mindset to understand each customer’s environment, anticipate and prevent issues, and solve complex technical and business challenges. In a typical week, you might be prioritizing strategic blockers with customer stakeholders, partnering with Product, Development, Infrastructure, Customer Succ
Jobiba hiring network
Senior Infrastructure Engineer Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Who We Are At Justworks, you’ll enjoy a welcoming and casual environment, great benefits, wellness program offerings, company retreats, and the ability to interact with and learn from leaders in the startup community. We work hard and care about our most prized asset - our people. We’re helping businesses get off the ground by enabling them to focus on running their business. We solve HR issues. We’re data-driven and never stop iterating. If you’d like to work in a supportive, entrepreneurial environment, are interested in building something meaningful and having fun while doing it, we’d love to hear from you. We're united by shared goals and shared motivations at Justworks. These are best summed up in our company values, which are reflected in our product and in our team. Our Values If this sounds like you, you’ll fit right in. Department Platform Engineering Who You Are You are a tooling-focused engineer who obsesses over developer productivity. You’ve built and maintained CI/CD pipelines, developer tooling, and automation that makes engineering teams faster. You understand that the best developer tools disappear into the background - they just work. You’re equally comfortable debugging a flaky GitHub Actions workflow and designing a new internal CLI. You care deeply about reducing friction and cognitive load for your fellow engineers. This is a foundational role. You’ll be one of the first dedicated engineers on a platform organization supporting 300+ engineers. You’ll shape not just the technical foundations but the culture and practices of the team. Your Success Profile What You Will Work On Own and evolve CI/CD infrastructure (GitHub Actions, deployment pipelines) to improve build times and deployment frequency Build and maintain our developer portal (Backstage) - service catalog, golden path templates, and TechDocs integration Create golden path templates for Go, Rails, and Vue.js services that bake in best practices automatically Build and maintain internal
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! About the Team The NR Lens team builds New Relic's Federated Data Platform — a distributed SQL query engine that lets customers connect external data sources (Snowflake, PostgreSQL, Google Sheets, AWS CloudWatch) and query them directly from New Relic. You'll work on the distributed query execution layer, connector architecture, and API gateway that powers cross-source JOINs and unified analytics across customer data stores. What You'll Do Design, build, and maintain cloud-native Java microservices in the NR Lens query path: SQL Gateway, Query Gateway, and data source connector plugins Develop and harden JDBC connector integrations — including connection lifecycle management, credential handling, query pushdown optimization, and security validation Improve query reliability and performance across a multi-tenant distributed SQL deployment serving 50+ customers Build and operate services on AWS (EKS, IAM/STS, S3) with infrastructure-as-code practices Implement
About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About Mixpanel Engineering Mixpanel Engineering is a small, fast-moving team focused on delivering real value to customers. We build powerful AI-powered product analytics while obsessing over clarity, simplicity, and delight. Product innovation drives our business, and product engineering teams own that responsibility. The Product Engineering org is responsible for building our core product that delivers insights to customers. These teams are defining the next phase of our AI-led product experiences and are uniquely positioned to build the next generation of AI-first products. We build the tools product managers rely on to understand users, protect revenue, and scale companies. As we move toward IPO, our work evolves from showing what happened to explaining why. This shift unlocks deeper insight, better decisions, and the next generation of analytics. About the Role The Director of Engineering, reporting directly to the CTO, will be responsible for all of Product Engineering. You'll own the core analysis experience that product managers and engineers rely on daily, the growth infrastructure that brings new builders into the product and keeps them there, and the platform capabilities that let teams test and ship with confidence. Success is measured by revenue and retention. What You'll Do Lead multiple engineering teams across different product areas. This role is for someone who develops people seriously and holds a high bar for execution. Build an AI first product analytics for the next generation of software Hire well, coach deeply, and build a culture where engineers feel ownership and pride in what t
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role The ML Infrastructure team is responsible for building & improving the core infrastructure for autonomy teams at Nuro. In this role, you will work closely with teams across Nuro, to design, build and deploy core infrastructure components in machine learning model life cycle, to push the autonomous future forward. You will have an opportunity to work across the full stack of machine learning solutions - from designing robust and scalable model & data pipelines to building to deploying the optimized models on Nuro’s fleet of self-driving robots! About the Work Design and develop ML workflow pipelines to train, optimize, validate, and deploy Nuro autonomy models. Develop and maintain continuous testing and monitoring systems for core ML infrastructure components. Develop observability to track ML model lifecycles from data generation to on-road validation. Maintain an in-house ML inference platform to serv
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet is hiring a Senior Machine Learning Operations Engineer to architect our machine learning production lifecycle. Your mission is to maintain and deploy ML models to a scalable, reliable, and secure production environment. You will design and maintain the infrastructure, automation, and monitoring systems that ensure our AI products are high-performing and cost-effective. You will report to our Director, Analytics Engineering & Data Governance and work from our Bangalore, India office. You Will: Model and Pipeline Automation Automate the deployment and retraining of ML models, from training through to production inference, by building and managing complete CI/CD/CT (Continuous Training) pipelines, adhering to MLOps best practices. Build, fine-tune, or use pre-trained LLMs, deep learning models or traditional machine learning models. Evaluate and recommend AI or ML solutions for the product using any combination of vendor solutions and/or custom-built models. Governance & Compliance Implement model versioning, lineage tracking, and auditing to ensure compliance with security and ethical standards. Performance Monitoring Continuously monitor the health and performance of production machine learning models, proactively identifying and correcting model drift, staleness, and performance degradation. Incorporate user feedback for iterative improvements and manage necessary model retraining cycles. Cross-Functional Collaboration Act as the "glue" between Data Scientists (who build models
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Job Description/ Responsibilities: Designing, developing and maintaining stable and reliable AI/ML Ops platforms / pipelines Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production. Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow. Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data Technology Eva
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX - Hybrid About the team The Data Intelligence & Analytics organization builds the core data platform and internal products that power decision-making across the company. We design and operate large-scale data systems, own the company’s data lake, ingestion infrastructure, and platform tooling, and develop end-to-end applications that transform complex datasets into fast, reliable, business-critical products used
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Security Engineer on the Detection and Response (D&R) team at Roblox, you’ll protect our user community alongside the underlying platform infrastructure. You’ll design high-fidelity detections, engineer security data platforms, and respond alongside the team during incidents. This is a hybrid in-office role in San Mateo. You Will: Deliver robust D&R capabilities: Engineer high-fidelity detections end-to-end. Lead partners through threat modeling and logging, to deploying actionable alerts, while keeping false positives low. Build security data pipelines: Develop security data pipelines and actively contribute to internal software and data platforms, collaborating across engineering teams. Ensure service reliability: Participate in an on-call rotation to keep detection and response services healthy. Embody security culture: Serve as a trusted security partner across Roblox, helping protect our community and enterprise while fostering a culture grounded in trust, ownership, and shared responsibility. You Have: 3+ years of experience in Security Data Engineering: You have built services that are efficient, reliable, and scalable using programming languages like Golang or Py
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Enterprise Security Engineer, you will play a critical role in executing Roblox’s Enterprise Security Strategy. You will design, deploy, and manage security solutions to protect Roblox’s corporate infrastructure and ensure secure, compliant operations across the organization. Working closely with Corporate Engineering and Trust & Safety teams, you will translate business requirements into robust security implementations that enable secure productivity while mitigating risk. You will be reporting directly to the Senior Manager of Enterprise Security Engineering. You'll partner with security professionals across the Information Security organization, and work cross-functionally with teams throughout Roblox to drive security initiatives that scale with our business. You will: Evaluate and implement security technologies and vendor solutions to ensure alignment with enterprise security requirements, compliance standards, and overall risk management strategy Lead and drive initiatives across core security domains, including Endpoint Security, SaaS Security, Identity & Access Management (IAM), Agentic AI Governance, and Supply Chain Security. Collaborate closely with IT, engin
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Engineering Manager, Communications, you'll lead the team responsible for in-game text chat — one of the most-used surfaces on Roblox and a cornerstone of how our community connects. Every day this system carries billions of messages across hundreds of millions of users, in real time, across 2D and 3D spaces, on every device we support. You'll own the roadmap and the engineering org behind it: building rich, immersive, and engaging communication experiences that let people and creators express themselves safely and seamlessly — better than in real life. This is a role for a leader who thinks like a product builder as much as an engineer. You'll balance the demands of massive scale and rock-solid reliability with a relentless focus on the user experience, shipping features that make conversation on Roblox feel effortless, expressive, and safe. You'll grow and develop a team of engineers, set technical direction, and partner deeply across product, design, trust & safety, and infrastructure to define what communication on Roblox becomes next. You Will Lead, grow, and develop a team of software engineers building the in-game text chat platform, setting a high bar for engineering
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Senior Engineering Manager, Safety Platform You will lead engineering pods within the Safety Platform organization, driving the technical vision and execution for the foundational systems that power every safety workflow at Roblox. You will own the core safety platform end-to-end: from the shared infrastructure and APIs that enable trust & safety capabilities across the company, to the tooling ecosystem that empowers internal operators to investigate, intervene, and resolve issues at scale. The scope also includes the Safety agentic platform — AI-powered systems that automate and augment safety workflows across detection, enforcement, and review pipelines. As the platform layer beneath every safety surface, your work will define how quickly and reliably Roblox can respond to emerging threats. This role requires a leader with a strong platform mindset who can balance technical rigor with broad organizational impact — working across User Safety, Trust & Safety Policy, Data Science, and Machine Learning to deliver scalable, extensible, and high-availability systems for one of the world's largest platforms. You Will: Lead and Develop: Recruit, hire, mentor, and inspire a diverse team of
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's engineering team is expanding rapidly, and we're looking for a seasoned engineering leader to help us scale our AI-powered creation platform and team. As a Senior Engineering Manager within the Creator Group, you will lead the Studio Assistant team — driving the strategy and execution of our agentic AI system that powers creation for millions of Roblox creators. Your team owns everything end-end from the user experience in Studio, to the infrastructure underneath backend by a single cloud-native architecture: one harness, one tool set, one eval framework. Your work will empower millions of creators to plan, generate, and ship content faster than ever before. At Roblox we move fast and ship code to production daily; your leadership will increase productivity by removing obstacles and keeping processes lean. You will identify and mitigate risk and make sure the technology outpaces our growth. Roblox teams are inclusive, helpful, and have a strong sense of ownership over the things they build. If you have a desire to grow and learn, you will fit right in with our highly-skilled and ever-expanding engineering team. You Are: Battle ready: You have experience defining architecture for AI
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Foundation AI organization, you will sit at the epicenter of our foundation model efforts. While the research world is focused on architecture, you will be the architect of the data flywheel that makes VideoGen and 3DGen possible. You aren't just building pipelines; you are building the infrastructure that defines how our models perceive and generate virtual worlds in three dimensions and across time. In this role, you will partner directly with our AI researchers to advance beyond experimental datasets and into the realm of dynamic, high-fidelity data synthesis and evaluation. You will bridge the gap between research prototypes working locally to scaling for millions of users. You will design, implement, and scale robust, high-performance infrastructure to crawl, create, curate, store, and serve the massive datasets required for these models. We are seeking accomplished software engineers with a passion for data, experience building large distributed systems, and a commitment to writing high-quality, well-tested code to solve complex data challenges at scale. Your contributions will ensure that our foundation models receive the highest quality dat
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Why Reliability? Roblox serves over 100 million people every day across a platform that is constantly evolving — and behind every experience is infrastructure that has to work, every time, at massive scale. The Reliability team at Roblox operates at the depth and breadth of the Roblox stack. Availability of the platform is a key company goal. We are hiring our first Senior Machine Learning engineer within our team. As a Senior Machine Learning Engineer within Reliability, you will help set the direction for how machine learning systems/practices can be leveraged to improve the reliability of the overall Roblox platform. You will own the architectural and execution roadmap of leveraging massive data across - logs, traces, metrics, production changes, to proactively detect issues before they become real problems (MTTD) and/or reduce time to resolve incidents (MTTR). You will have the opportunity to cross functionally collaborate with other similar teams at Roblox to define best practices and software. You will: Help define the roadmap for leveraging Machine Learning Engineering to improve Production Systems Reliability at Roblox. Improve r
Get new senior infrastructure engineer jobs by email
Daily job updates · Unsubscribe anytime