DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the job DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides an in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. In this role, you will Build core capabilities for our SaaS Platform across multiple clouds Drive development of functional enhancements for Data Discovery, Observability & Governance for both OSS and SaaS offering Lead efforts around non functional aspects like performance, scalability, reliability Lead and mentor junior engineers Work closely with PM, Customers and OSS community Requirements Over 8+ years of experience building and scaling backend systems, preferably in cloud-first or SaaS environments. Solve complex tech
Jobiba hiring network
Lead Lead Junior Software Engineer Reliability Jobs
15 active opportunities · Updated for September 2026
Fresh results
15 shown
Explore current lead lead junior software engineer reliability jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Opportunity Overview: This is a unique opportunity to join a software engineering team that is growing quickly. You will build impactful healthcare technology on a modern stack utilizing your full stack software engineering background. Last but not least: People who succeed here are empathetic teammates who are candid, kind, caring, and embody our core values and principles . We believe that diverse, inclusive teams make the most impactful work. Cohere is deeply invested in ensuring that we have a supportive, growth-oriented environment that works for everyone. What you’ll do: Lead technical design and architecture for complex features across our business intelligence platform Take full ownership of end-to-end feature releases, platform enhancements, and system optimization initiatives Drive technical decision-making by conducting thorough analysis, evaluating trade-offs, and making data-driven architectural choices Architect and implement highly scalable cloud-based data services, APIs, ETL pipelines, and user interfaces Build mission-critical backend systems supporting analytics, reporting, and complex data query capabilities at scale Design and develop intuitive UI platforms and interactive dashboards for data visualization and business intelligence Champion engineering excellence by establishing and enforcing best practices, design patterns, and coding standards Optimize system performance, identify bottlenecks, and implement solutions for scalability and reliability Define and maintain comprehensive test strategies, ensuring high code quality and test coverage Actively participate in ensuring Cohere maintains a disciplined approach to healthcare security, compliance, and data governance Mentor and coach junior and mid-level engineers, fostering their technical growth and career development Collaborate cross-functionally with product, data science, analytics, and business stakeholders to deliver robust technical solutions ISMS roles and respon
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with 5–8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role—leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet. What You'll Do Architectural Leadership: Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management. Platform Orchestration: Lead complex, large-scale infrastructure shifts. Operational Excellence: Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies. This Role Requires Experience: 5–8 years in a DevOps, Site Reliability, or Infrastructure Engineering role. Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators. Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation. Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through auto
Healthcare is complex. We’re here to change that. RVO Health is a health technology company on a mission to make health easier to navigate, more accessible, and more affordable for everyone. Here, you'll help over 40 million people every month, with a team that genuinely cares about the work and each other. AT A GLANCE The Senior Software Engineer is a crucial role within our organization, requiring work in various capacities and adaptation to different work arrangements based on the needs set by the business. The successful candidate will be responsible for fulfilling their job duties in the following work situations: Where You'll Be Location: Denver, CO | Hybrid We believe great collaboration happens when we're together, solving problems, learning from each other, and connecting as a team. That's why we’re in our offices Tuesday through Thursday each week. You are welcome to work remotely Mondays and Fridays if you wish. Address: 1801 California St. Denver, CO 80202 What You’ll Do Lead the end-to-end design, development, and implementation of sophisticated software applications and systems aligned with business goals. Collaborate closely with stakeholders including product managers, designers, and other engineers to gather requirements and translate them into robust technical designs and solutions. Write high-quality, efficient, maintainable, and scalable code adhering to best practices and company standards. Debug, analyze, and resolve complex software defects and performance bottlenecks to ensure optimal system reliability and user experience. Conduct comprehensive testing and validation including unit, integration, and performance testing to guarantee software quality. Mentor and provide technical guidance to junior and mid-level engineers, fostering professional growth and knowledge sharing. Perform thorough code reviews to maintain high code quality, enforce coding standards, and promote best p
Healthcare is complex. We're here to change that. RVO Health is a health technology company on a mission to make health easier to navigate, more accessible, and more affordable for everyone. Here, you'll help over 40 million people every month, with a team that genuinely cares about the work and each other. AT A GLANCE The Senior Software Engineer is a crucial role within our organization, requiring work in various capacities and adaptation to different work arrangements based on the needs set by the business. Where You'll Be Location: Denver, CO | Hybrid We believe great collaboration happens when we're together, solving problems, learning from each other, and connecting as a team. That's why we're in our offices Tuesday through Thursday each week. You are welcome to work remotely Mondays and Fridays if you wish. Office Address: 1801 California St. Denver, CO 80202 What You’ll Do Lead the end-to-end design, development, and implementation of sophisticated software applications and systems aligned with business goals. Collaborate closely with stakeholders including product managers, designers, and other engineers to gather requirements and translate them into robust technical designs and solutions. Write high-quality, efficient, maintainable, and scalable code adhering to best practices and company standards. Debug, analyze, and resolve complex software defects and performance bottlenecks to ensure optimal system reliability and user experience. Conduct comprehensive testing and validation including unit, integration, and performance testing to guarantee software quality. Mentor and provide technical guidance to junior and mid-level engineers, fostering professional growth and knowledge sharing. Perform thorough code reviews to maintain high code quality, enforce coding standards, and promote best practices across the team. Continuously improve software development processes, tools, and methodol
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and automations that power Smartsheet’s core business functions across Finance, Sales/GTM, and People & Culture. Our team owns mission-critical systems and workflows that enable how the company hires, sells, bills, pays, reports, and scales. We operate at the intersection of software engineering, enterprise platforms, and business-critical data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Automation team builds human-to-system and system-to-system automations that reduce manual effort and friction across the business. We combine cloud-native services, agentic AI, and workflow orchestration to enable employees to interact with enterprise systems through intelligent, secure, and auditable automation. As a Senior Software Engineer I (Automation), you will lead the design, build, and operation of systems and workflows that directly support business execution at scale. You will own complex technical initiatives, partner with Product Managers and stakeholders on technical roadmaps, and mentor junior engineers. This full-time position reports to the Sr. Director, Development and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Architect AI Agents: Take a leading role in designing Agentic Workflows using AWS Step Functions and Bedrock Agents that reason
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and automations that power Smartsheet’s core business functions across Finance, Sales/GTM, and People & Culture. Our team owns mission-critical systems and workflows that enable how the company hires, sells, bills, pays, reports, and scales. We operate at the intersection of software engineering, enterprise platforms, and business-critical data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Finance Systems team engineers and operates the platforms that support financial operations, including ERP, procurement, billing, and compliance. We work across configuration, extensibility, and integration to ensure systems are scalable, auditable, and resilient, treating code, configurations, and controls with the same rigor as software. As a Senior Software Engineer I (Finance Systems), you will lead the design, build, and operation of systems and workflows that directly support business execution at scale. You will own complex technical initiatives, partner with Product Managers and stakeholders on technical roadmaps, and mentor junior engineers. You will report into a Manager, Enterprise Systems, and can be based in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Systems Architecture & Optimization: Engineer and lead the end-to-end lifecycle—analysis, prioritization, and tec
Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Design, build, and maintain APIs, services, and systems across Stripe’s engineering teams using Java, Ruby, Scala, and Go Build software infrastructure, including developing, testing, and deploying it Design and develop software systems, using scientific analysis and mathematical models to predict and measure outcome and consequences of design Engineer payments integration with various financial partners software systems Design APIs and underlying data models to support complex financial abstractions and multi-party integrations, enabling flexible billing and settlement configurations across international markets Develop and direct software system testing and validation procedures, programming, and documentation Debug production issues across services and multiple levels of the stack Analyze user needs and software requirements to determine the feasibility of design within time and cost constraints Work with engineers across the company to build new features at large-scale Build new systems to securely store sensitive data; Improve engineering standards, tooling, and processes Integrate observability tools and alerting mechanisms (Datadog, Prometheus) into high-traffic production systems, defining SLAs and real-time diagnostics to meet the reliability standards required by financial partners Mentor junior engineers, lead technical design reviews, and establish best practices in code quality, system design, re
About the Role: We are looking for a Senior DevOps Engineer to join our DevOps team at K Health. You will own and evolve the infrastructure underpinning a healthcare AI platform serving patients and enterprise health system partners. This is a high-ownership role: you will architect and operate cloud environments across K Health and its enterprise partners, lead complex infrastructure migrations, drive disaster recovery programs, and help build the next generation of AI-powered operations tooling. You will also mentor junior engineers and collaborate closely with product and engineering teams across the company. This is a hybrid role based in New York City (4 days/week in office) and includes participation in a daytime on-call rotation. What you will do: Own the design, implementation, and evolution of our GKE-based Kubernetes infrastructure across K Health and enterprise partner environments. Build and maintain our Terraform modular infrastructure library, including reusable modules with automated testing, across GCP, Cloudflare, and AWS. Architect, build, and maintain GitLab CI/CD shared pipeline templates used by all engineering teams (build, test, security scanning, deployment). Own and maintain self-hosted infrastructure software running in-cluster, including GitLab, ArgoCD, Langfuse, DependencyTrack, NGINX Ingress, and others. Implement and support security and compliance controls across infrastructure and the software supply chain - secrets management, pipeline secret detection, container scanning, SOC2 and HIPAA. Drive disaster recovery readiness: design failover scenarios, author runbooks, and lead periodic DR tests. Lead development of AI-powered operations tooling and agentic infrastructure. Monitor, troubleshoot, and improve production system reliability; respond to incidents during on-call shifts. Mentor junior DevOps engineers and establish team-wide engineering standards. What we are looking for: 5+ years of experience in DevOps, platform engineering,
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Team Lead for Initiator & Protocol Engineering, you will spearhead the critical bridge between our industry-leading FlashArray and the Linux/VMWare ecosystems. You will drive the performance and reliability of our storage protocol stacks—spanning NVMe over Fabrics and Fibre Channel—ensuring Pure Storage remains the gold standard for enterprise connectivity. Collaborating closely with cross-functional hardware and software teams, you’ll mentor a high-caliber engineering squad to solve complex kernel-level challenges and influence the global Linux upstream community. WHAT YOU'LL DO Own the Protocol Lifecycle: Lead the development, maintenance, and optimization of Linux and VMWare initiator stacks (NVMeoF, FC-SCSI, iSCSI) and target drivers to ensure seamless, high-performance integration with Pure FlashArray. Drive System Resilience: Architect enhancements for Fibre Channel and NIC driver stacks that improve RAS (Reliability, Availability, and Serviceability), specifically focusing on multipathing logic and link health monitoring. Technical Leadership & Mentorship: Guide a team of senior and junior engineers through complex project deliveries, conducting deep-dive code reviews and setting the technical bar for C/C++ and Python development within the kernel space. Solve the Impossible: Act as the final escalation point for the most challenging system-level bugs found in the field or internal testing, u
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We’re hiring a talented Software Engineering Manager to lead the Snowtrail infrastructure team at Snowflake. Snowtrail is the infrastructure that enables Snowflake to deliver dedicated coverage for customer-specific workloads. Its innovative approach allows Snowflake to precisely test and measure the impact of changes on individual customers, making it essential for ensuring the platform’s reliability, correctness, and performance. Through query replay, Snowtrail helps us catch regressions early. By leveraging machine learning models to intelligently sample queries and workloads, we continuously optimize for both cost and performance. Evolving Snowtrail to incorporate new engine features while improving scalability, efficiency, and reliability is central to our continued success OUR IDEAL MANAGER WILL HAVE : Strong passion and proven track record for shipping quality software in high code velocity environments 10+ years industry experience designing and building distributed data systems. Excellent problem solving skills, and strong CS fundamentals including data structures, algorithms, and distributed systems. Fluency in SQL, Java, C++, Python or Go. Ability to collaborate well across teams, build high-performing teams and mentor junior engineers. Excellent interpersonal co
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Luxury team builds rich and creative features that set the standard for the ride sharing industry. We are looking for a motivated software engineer to join us building features for Black, Black SUV, and other upmarket modes where quality is uncompromising, where we work on balancing and growing these unique marketplaces within the wider Lyft ecosystem and influencing driver behavior to facilitate highest-quality rides. Ownership is a key quality for this team; this person should be driven to track a project to successful completion and beyond, taking initiative to work with other teams and functions to ensure the code they write reaches users and drives impact. You'll collaborate with engineering, product, data science, analytics, and operations on programs that empower us to iterate quickly, delighting our passengers and drivers. Preferred applicants intend to work regularly from our CDMX office, where team members collaborate together in person. Responsibilities: Help establish roadmap and architecture based on technology and understanding of customer needs Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Share your knowledge by giving brown bags, tech talks, and promoting appropriate tech and engineering best practices Can help lead large projects from idea to positive execution Unblock, support and communicate with internal partners to achieve results Experience: Engineering industry experience Experience with object-oriented programming Experience in distributed systems Experience working with databases, relational or NoSQL Write clear, scalable and clear design documentation Design, build and improve a set of team owned components Please submit your resume in English.
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on our Ads Platform team, you will architect and build foundational ads systems that are capable of reliably serving Roblox’s massive scale. You will drive performance improvements to our ads serving and delivery infrastructure, while expanding our systems to support a variety of new and immersive ads experiences. Your scope will span across multiple teams and require close collaboration with engineers across ML, backend, and analytics to shape the future of ads infrastructure at Roblox. This role is a great fit for you if you are proficient in managing large scale systems and data-driven applications and have a zeal for developing inspiring, easily maintainable, and reusable code. Join our highly skilled and rapidly growing team and make a significant impact at Roblox. This position will report to our Senior Engineering Manager and require you to work onsite at our HQ in San Mateo, CA Tuesdays to Thursdays. You Will: Build and ship best-in-class Ads products used by hundreds of millions of users. Develop backend services that scale to billions of products and transactions. Be a tech lead and mentor junior engineers Self-organize and take ownership of projects
ABOUT THE ROLE The Senior Software Engineer acts as the technical lead for an agile team, focusing on the design and development of scalable microservices for Peloton's core features across all platforms (Bike, Tread, Strength, and Digital). The Content AI Team is responsible for training and hosting various models in the domains of Natural Language Processing (NLP) and Automatic Speech Recognition (ASR). This role is highly collaborative, requiring close partnership with Product, Design, and QA throughout the development lifecycle. The candidate will be responsible for designing, developing, testing, deploying, and monitoring microservices that specifically power features like search, voice, and captions, which are crucial for platform expansion and international growth. In addition to technical delivery, excellent communication and engineering skills are essential for providing guidance to junior team members and effectively coordinating with other stakeholders, such as Technical Program Managers (TPM). YOUR DAILY IMPACT AT PELOTON Design, develop and operate business-critical APIs and services with a focus on high availability, low latency, security and scalability Mentor junior engineers in the team and communicate effectively with TPM and Product Write understandable, well tested code with an eye towards maintainability and scalability Architect and build reusable code, libraries, and patterns for use across teams Propose, experiment, and implement solutions to scale services while meeting business and product requirements. Leverage production monitoring/profiling/tracing and load testing tools to discover bottlenecks and using techniques such as data modeling, query optimization, and caching to address the bottlenecks Lead technical discussions during architecture meetings, design reviews, and task breakdown Identify common patterns as well as develop and foster development of reusable components and standards across teams Champion and implement industry
Get new lead lead junior software engineer reliability jobs by email
Daily job updates · Unsubscribe anytime