Senior Software Engineer, DGX Cloud Production Engineering — 2 Locations. Apply via Workday.
Jobiba hiring network
Senior Software Engineer Production Engineering Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior software engineer production engineering jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: We are looking for a Senior Software Engineer to join our Site Reliability Engineering team. As a Senior Software Engineer in Production SRE, you will be responsible for developing and maintaining the tools and systems that enable our engineering teams to operate our services reliably and at scale. You will work closely with our SREs and other engineering teams to ensure our services are properly instrumented and able to scale with our growing business. The Difference You Will Make: In this role, your expertise in developing and maintaining tools and systems will be instrumental in bolstering our services' reliability and improving how the company manages incidents broadly. By collaborating closely with other engineering teams you will help establish a culture of reliability throughout the organization by providing a comprehensive incident management platform that is being used for instrumentation, operability, and around incidents. Your ability to identify opportunities for improvement and drive their implementation will contribute significantly to our overall operational efficiency and growth, ensuring that our services remain resilient as our business continues to expand. Additionally, as an essential part of this role, you will serve as an active member of the Production SRE team, responding to and managing high severity incidents. Your vast technical experience and leadership skills will be invaluable as you step into the role of Incident Commander during these critical events. You will guide cross-functional teams during crisis situations and ensure timely resolution, minimizi
At Rockstar Games, we create world-class entertainment experiences. Become part of a team working on some of the most rewarding, large-scale creative projects to be found in any entertainment medium - all within an inclusive, highly-motivated environment where you can learn and collaborate with some of the most talented people in the industry. Rockstar is on the lookout for a talented Software Engineer who possesses a strong interest in all the low-level technology that makes a modern video game tick to support the Cfx.re creator platforms, including FiveM and RedM. As a member of our team, you will need a critical and creative eye capable of putting forth innovative solutions to complex problems. If you like to understand how things really work “under the hood” of your favorite games, we’d love to hear from you. This is a full-time, permanent and in-office position based in Rockstar’s unique game development studio in the heart of London. WHAT WE DO The Rockstar Creator Platform Team deliver a technology platform that enables players to experience community created content on fully customized dedicated servers where creators can develop their own game modes and other modifications in a variety of scripting languages. We create technology, tools, and solutions to enhance the creator experience and empower our community to create and share any experience imaginable. RESPONSIBILITIES Maintain and improve existing and new codebases, ensuring high standards of quality, stability, and efficiency in collaboration with cross-functional teams. Support software release processes across multiple branches, coordinating with production and engineering teams to ensure readiness and a high standard of quality. Help identify, prioritize, and resolve critical issues, coordinating timely fixes while maintaining overall system stability. Drive release planning, deployment, and rollback procedures. Maintain and enhance build, test, and release automation
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft is looking for software engineers from a scope of disciplines. We are growing our team with people who want to build, improve and incorporate technologies that make the lives of our community more enriched. As an engineer at Lyft, you'll collaborate with teams like product, data science, analytics, and operations on code that empower us to iterate quickly, while focusing on delighting our passengers and drivers. The Applied AI team is looking for a Backend Engineer to join the AI Entries team. You will build the foundational infrastructure that connects Lyft to emerging AI ecosystems and devices and architect the APIs and orchestration layers that allow third-party agents and multimodal interfaces to interact with Lyft. By building these robust integration, you will help make Lyft available where our riders are. Responsibilities: Establish engineering best practices and patterns; help uplift the team's craft and drive a culture of engineering excellence Drive high-impact projects and innovate new solutions to deliver the best user experience Produce and drive scalable system design for large, complex features — from idea through execution and launch Mentor engineers on the team, providing technical guidance and supporting their growth Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and share knowledge across the team Participate in the team's on-call rotation; identify, triage, debug, and resolve issues across our applications and platforms Have the ability to explain the various trade offs made in decisions Manage project priorities, deadlines, and deliverables. Experience: BS/MS or equivalent in Computer Engineering, Computer Science, or related field or equivalent practical experience. 5+ years of software engineering/production
Must be based in Vancouver The role We're hiring a dedicated data engineer to own the production data platform that our delivery, product, and engineering teams run on; designing integrated, governed data pipelines and delivering automated reporting, AI-assisted workflows, and predictive signals on top of them. You'll write production code, design systems, own CI/CD, and be accountable for the correctness of data that leaders make decisions on. What you'll do Design and operate our cloud data platform: ingestion, transformation, orchestration and serving. Integrate data from across the business (delivery tooling, CRM, product telemetry, finance, support and customer feedback systems) with shared identifiers, data contracts and lineage. Build automated and continuously refreshed reporting so teams manage by exception rather than chasing status. Connect approved AI agents to governed data with structured outputs, provenance, guardrails and human approval in the loop. Build feature pipelines and the MLOps controls behind predictive use cases: tests, versioning, promotion gates and drift monitoring. Own the engineering standards for data: testing, observability, environment promotion, PII classification and access control. What you'll bring Strong software engineering fundamentals: production-quality code, API and interface design, testing discipline, systems design. Real experience building and operating production data platforms on a cloud warehouse or lakehouse (Snowflake and AWS preferred) with dbt and a modern orchestrator. Practical AI tooling experience: something shipped, not prototyped. LLM-backed classification, extraction or structured-output pipelines; agent and tool-calling workflows; retrieval; evals. You can reaso
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Security Software Engineer on the IAM team at Roblox, you'll build the next generation of identity and access management — defining how both humans and AI agents get identity, authenticate, and receive access to Roblox's production infrastructure. As AI agents become first-class actors in our systems, you'll design the tooling and policy that governs what they can do, how they prove who they are, and how we keep that access safe at scale. You'll also continue to evolve our workload authentication, privileged access management, and secure "golden path" for developers. Your work will directly shape the security posture of our entire production environment and set the standard for agentic IAM across the industry. You will: Design Identity and Access for AI Agents: You will define how AI agents get credentials, receive scoped permissions, and have their sessions managed throughout their lifecycle — pioneering the patterns for agentic identity in production. Engineer Hybrid Production IAM at Scale: You will design and implement scalable IAM solutions for Roblox's hybrid production environment, spanning on-premises and cloud infrastructure, ensuring secure and efficient access for hum
As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute. Experience building integrations across multiple model providers. Experience with AW
As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. To provide substantive overlap with the team, this position must be in Eastern Time. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute.
The Senior Enterprise Engineer opportunity Statsig is building an Enterprise Engineering team for engineers who want unusually direct ownership of high-stakes customer outcomes. Like a Forward Deployed Engineering team, Enterprise Engineering works extremely close to customers and tackles the technical problems that matter most to them. The difference is the scope of ownership. Forward Deployed Engineering is often concentrated around implementation or onboarding; Enterprise Engineering is responsible for a customer’s technical success throughout the entire customer journey. You’ll stay close as needs evolve, answering questions, investigating across the stack, identifying root causes, shipping fixes and features, and remaining accountable through resolution. This is a software engineering role with a tight customer feedback loop. It is ideal for a pragmatic generalist, especially someone with backend or infrastructure depth who enjoys ambiguous problems, production systems, fast decisions, and immediate impact. What you’ll do Serve on the engineering front line in Unthread, our Slack aggregation and workflow layer, for key customer accounts. Answer technical questions by reading the code, tracing behavior, inspecting production signals, and building a clear explanation, not by forwarding the thread. Diagnose and fix bugs, then validate the outcome with the customer. Design and implement product or platform features when doing so is the fastest, highest-quality way to solve the customer’s problem. Act as the engineering point of contact for urgent issues in key customer accounts’ Slack channels, coordinating the response while retaining technical ownership. Represent engineering in customer office hours and, for selected accounts, recurring weekly meetings. Partner closely with customer-facing and product teams while maintaining crisp ownership, communication, and escalation paths. Identify recurring patterns and improve tooling, documentation, observability, APIs,
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Customer Experience Engineering Team builds the internal and external technologies that scale Snowflake’s global support and sales organizations. We empower our technical experts by providing the advanced tools they need to resolve complex issues and drive customer success. Our team specializes in software engineering, data-driven decisions, ML, and LLM-based solutions . We build production-grade systems to automate manual processes and augment the capabilities of our technical staff. Our current focus includes: LLMs : Developing and deploying LLM and agent-based architectures for streamlining troubleshooting Scalable Evaluations : Implementing large-scale evaluations to ensure the quality and reliability of our internal and external tools Process Automation : Designing intelligent workflows that eliminate bottlenecks and allow our experts to focus on the most technical aspects of the Snowflake platform Incident discovery: using embeddings, LLMs, clustering, and agents to detect potential widespread issues more quickly Now, the team is growing, and we are looking for a Software Engineer to join us. In this role, you will work closely with the state of the art LLM models, fine-tune them, develop agents, apply various clusterings, summarizations, embeddings, and so on. Ev
About the Role Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work across two areas: Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve. We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure. This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation , solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure. responsible for delivering the software but also for operating and supporting it in production. Why this Role You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. You’ll have the opportunity to build foundational systems from the ground up, work on infrastructur
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. About the team The Apps & Experiences Platform team powers the systems and services behind all of our user-facing applications, including Snowsight, Snowflake Intelligence, and new mobile experiences. Our mission is to build innovative backend services, developer tooling, platform infrastructure, and AI-powered capabilities that enable exceptional product experiences at scale. As part of our team, you’ll work across feature development, platform engineering, infrastructure, and internal tooling to support both end users and developers. We care deeply about building systems that are reliable, scalable, maintainable, and performant. Snowflake is a high-growth AI Data Cloud company, and we’re looking for exceptional engineers to help us scale the next generation of our platform. A key part of this is our work on our internal AI developer agent, which is designed to fundamentally democratize end-to-end web app development across all engineering teams by translating product specs and designs into a fully functional, production-ready features. AS A SENIOR SOFTWARE ENGINEER FOR THE APPS & EXPERIENCES PLATFORM TEAM, YOU WILL: Design, build, and operate scalable backend services and platform infrastructure that power Snowflake’s user-facing applications. Contribute across th
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. About the team The Apps & Experiences Platform team powers the systems and services behind all of our user-facing applications, including Snowsight, Snowflake Intelligence, and new mobile experiences. Our mission is to build innovative backend services, developer tooling, platform infrastructure, and AI-powered capabilities that enable exceptional product experiences at scale. As part of our team, you’ll work across feature development, platform engineering, infrastructure, and internal tooling to support both end users and developers. We care deeply about building systems that are reliable, scalable, maintainable, and performant. Snowflake is a high-growth AI Data Cloud company, and we’re looking for exceptional engineers to help us scale the next generation of our platform. A key part of this is our work on our internal AI developer agent, which is designed to democratize end-to-end app development across all engineering teams by translating product specs and designs into fully functional, production-ready features. AS A SENIOR SOFTWARE ENGINEER FOR THE APPS & EXPERIENCES PLATFORM TEAM, YOU WILL: Design, build, and operate scalable backend services and platform infrastructure that power Snowflake’s user-facing applications. Contribute across the full development l
About the Team: Tubi's Internal Tools team is at the forefront of AI integration, developing everything from developer resources to production-grade AI for business operations. We are the group responsible for turning AI from an experiment into an operating capability: training, infrastructure, developer agents, and AI-powered business systems. Engineers operate with high ownership and autonomy, collaborating on shared architectural decisions and AI infrastructure. What You'll Do: Own systems end to end — design them, build them, and support them in production. Lead the projects you own: sequence the work, decide what lands first, and set technical direction for the engineers working with you. Sit with the people who use what you build, and turn what you learn there into a system. Design the service boundaries, contracts and schema evolution that let our platforms grow without breaking the teams depending on them. Make our AI systems dependable in production: evaluation harnesses, human approval steps before an agent acts, retries that handle a model returning something unexpected, and cost tracking that tells you what a task costs before you run it. Build what other engineers build on — agent skills, tool and MCP integrations, shared libraries — and raise the bar through code review, design discussion and mentoring. Spot the platform work nobody has asked for yet, make the case for it, and build it. Your Background: 5+ years of professional experience building and operating production systems, from design through production ownership. A system you designed and can walk us through end to end — where its boundaries sit, what constrained it, and what you chose against. Strong programming proficiency in a statically typed language such as Rust, Go, C++, Java, Kotlin, C#, or TypeScript. Production Rust is a plus rather than a requirement. You have owned a service in production: you wrote the runbooks, you knew what it cost, and you were the one paged when it broke. Expe
The Development Infrastructure team builds the tooling and systems our Asana engineers use every day to bring their ideas to production quickly and reliably. We build and operate the software that drives Asana’s roadmap. Each day, we combine industry best practices and innovation to support this product-focused company. We’re looking for an experienced Software Engineer with a passion for developer infrastructure. You will work with a world-class team of engineers on deploying and operating existing developer tooling, and building new tools to support our global, growing development team. You will have a unique opportunity to design and develop the systems and applications that drive the Asana development experience, lead complex technical projects, and work on cross-functional initiatives to help define the future of software engineering at Asana. This role is based in our Reykjavík office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Innovate the architecture of our development sandboxes to accelerate development iteration. Modernize our build system, tighten iteration loops, and polish frequent cycles in the developer experience. Lead complex developer infrastructure projects from technical design through implementation, rollout, and operation. Partner with engineering teams to identify opportunities to enable teams to develop faster at Asana and help the company achieve our goals faster. Analyze complex developer infrastructure systems to uncover issues, root causes, and areas for improvement. Keep Asana up to date on open-source trends and identify new opportunities for improved development infrastructure. Champion co
Get new senior software engineer production engineering jobs by email
Daily job updates · Unsubscribe anytime