The Team + The Role Our Emerging Team is focused on building AI Products for our product experience (PX) platform. We build from the ground up to explore, prototype, and ship AI-native experiences that change how software teams understand and serve their users. This is not an AI layer added to existing product; it is a deliberate bet on what product intelligence looks like next. The team operates with high autonomy, moves quickly, and builds products without clear precedents. As a Staff Software Engineer (AI), you will sit at the intersection of deep technical capability and strong product judgment. You will design and build production-grade AI systems, including RAG pipelines, agentic workflows, and LLM-powered features, while making clear tradeoffs across prompting, fine-tuning, architecture, evaluation, and deployment. You will also partner closely with product, design, and engineering stakeholders to frame the right problems and communicate technical decisions clearly. This role is based in our New York office. What this looks like day-to-day Applied AI systems: Design and build AI-native systems, including RAG pipelines, agentic workflows, and LLM-powered product features. You will take ideas from prototype through production and ensure they can support real users. Model strategy: Make principled decisions about when to prompt, when to fine-tune, and when to use a different technical approach entirely. You will explain those tradeoffs clearly to engineers and non-engineers. Evaluation and guardrails: Instrument and evaluate model outputs rigorously by defining evaluation frameworks and identifying hallucinations early. You will implement guardrails that hold up under real-world usage and load. Productionize AI ownership: Own model deployment, monitoring, latency optimization, cost management, and reliability at scale. You will ensure AI systems are observable, performant, and production-ready. Full-stack delivery: Contribute across the stack when needed to get
Jobiba hiring network
Reliability Engineer Jobs
2,027 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The Team + The Role Our Emerging Team is focused on building AI Products for our product experience (PX) platform. We build from the ground up to explore, prototype, and ship AI-native experiences that change how software teams understand and serve their users. This is not an AI layer added to existing product; it is a deliberate bet on what product intelligence looks like next. The team operates with high autonomy, moves quickly, and builds products without clear precedents. As a Sr. Software Engineer (AI), you will sit at the intersection of deep technical capability and strong product judgment. You will design and ship applied AI systems, including RAG pipelines, agentic workflows, and LLM-powered features, from prototype through production. You will make principled technical decisions, evaluate model behavior rigorously, and communicate tradeoffs clearly to engineers and non-engineers alike. This role is based in our New York office. What this looks like day-to-day Applied AI systems: Design and build AI systems including RAG pipelines, agentic workflows, and LLM-powered features. You will take work from prototype through production and ensure it can hold up in real customer environments. Technical decision-making: Make principled decisions on when to prompt, when to fine-tune, and when to use a different tool entirely. You will explain these tradeoffs clearly so the team can move quickly without sacrificing quality. Model evaluation: Instrument and evaluate model outputs rigorously by defining evaluation frameworks and catching hallucinations early. You will implement guardrails that can withstand real-world load and production use. Productionize AI ownership: Own model deployment, monitoring, latency optimization, cost management, and reliability at scale. You will help ensure AI systems are observable, efficient, and dependable in production. Full-stack product shipping: Contribute across the stack when needed because this team ships products, not just models
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. When was the last occasion you had the opportunity to contribute to a company that is shaping an industry and empowering individuals to translate ideas into tangible impact with speed? Smartsheet's core mission is to empower everyone to enhance their work processes. Our business model is founded on identifying exceptional talent and providing them with the autonomy to develop our acclaimed Software as a Service (SaaS) offering. With a user base exceeding 10 million, our platform is utilized across various industries, including construction, retail, and software development, presenting us with unique technical challenges. Smartsheet is seeking a Senior Business DevOps Engineer to join our Corporate Systems Development team in Bangalore. This role will focus on building and scaling our CI/CD pipelines, infrastructure automation, monitoring frameworks, and deployment processes supporting mission-critical integrations across Finance, People, Sales, Legal, IT, and Engineering systems. You’ll work across a variety of systems and platforms (AWS, GitLab, DataDog, Terraform, Boomi, UiPath) to streamline deployment of backend integrations and automation solutions. If you thrive on optimizing developer velocity, ensuring system reliability, and automating everything from build to deploy, this role is for you. The position reports to the Senior Manager, Systems Development and collaborates closely with global developers, architects, and application administrators to ensure our platform foundations are secure, efficient, and scalable
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. We are looking for a skilled and motivated Principal Software Engineer who is passionate about continuous learning and eager to grow along with us in a fast-paced, innovative environment. You will work remotely from Bulgaria and will be reporting to an engineering leader located in Bulgaria. You Will: Lead the design and implementation of Smartsheet's next-generation architecture, ensuring scalability, security, and performance for millions of global users. Define and drive architecture strategy, making key technical decisions that shape the future of the platform. Review and guide technical project designs, providing feedback during design review presentations to ensure system resilience and scalability. Take ownership of cross-functional technical initiatives, aligning teams around common architectural goals while driving large-scale projects to completion. Foster strong technical leadership, mentoring senior engineers and influencing best practices across multiple engineering teams. Lead deployment reviews for high-impact projects, ensuring they meet scalability, performance, and security requirements. Collaborate closely with product management and other business stakeholders to balance market needs with technical constraints, driving innovation while maintaining technical rigor. Advocate for quality and operational excellence, ensuring systems are monitored, tested, and maintained to meet the highest reliability standards. Perform other duties as assigned. You Have: Proven experience in system architecture and the d
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and automations that power Smartsheet’s core business functions across Finance, Sales/GTM, and People & Culture. Our team owns mission-critical systems and workflows that enable how the company hires, sells, bills, pays, reports, and scales. We operate at the intersection of software engineering, enterprise platforms, and business-critical data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Automation team builds human-to-system and system-to-system automations that reduce manual effort and friction across the business. We combine cloud-native services, agentic AI, and workflow orchestration to enable employees to interact with enterprise systems through intelligent, secure, and auditable automation. As a Senior Software Engineer I (Automation), you will lead the design, build, and operation of systems and workflows that directly support business execution at scale. You will own complex technical initiatives, partner with Product Managers and stakeholders on technical roadmaps, and mentor junior engineers. This full-time position reports to the Sr. Director, Development and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Architect AI Agents: Take a leading role in designing Agentic Workflows using AWS Step Functions and Bedrock Agents that reason
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and automations that power Smartsheet’s core business functions across Finance, Sales/GTM, and People & Culture. Our team owns mission-critical systems and workflows that enable how the company hires, sells, bills, pays, reports, and scales. We operate at the intersection of software engineering, enterprise platforms, and business-critical data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Finance Systems team engineers and operates the platforms that support financial operations, including ERP, procurement, billing, and compliance. We work across configuration, extensibility, and integration to ensure systems are scalable, auditable, and resilient, treating code, configurations, and controls with the same rigor as software. As a Senior Software Engineer I (Finance Systems), you will lead the design, build, and operation of systems and workflows that directly support business execution at scale. You will own complex technical initiatives, partner with Product Managers and stakeholders on technical roadmaps, and mentor junior engineers. You will report into a Manager, Enterprise Systems, and can be based in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Systems Architecture & Optimization: Engineer and lead the end-to-end lifecycle—analysis, prioritization, and tec
Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. The role We are looking for a Staff Backend Engineer to help build the next generation of SoFi's digital asset custody platform. This engineer will design and develop the core services that power secure custody, blockchain transaction processing, wallet infrastructure, and institutional digital asset products. This role sits at the center of crypto infrastructure, working closely with Security, Platform Engineering, Product, Treasury, and Operations to build highly available systems responsible for protecting customer assets. The ideal candidate enjoys solving distributed systems problems, designing resilient APIs, and building financial infrastructure where security and reliability are first-class requirements. What you’ll do: Design and build backend services supporting SoFi's digital asset custody platform Develop secure transaction orchestration services for deposits, withdrawals, transfers, staking, and treasury operations Build integrations with custody providers, blockchain infrastructure, and institutional settlement networks Design APIs supporting retail and institutional crypto products Develop services supporting wallet lifecycle management and blockchain interactions Improve platform resiliency through redundancy, disaster recovery, observability, and automated recovery mechanisms Optimize trans
Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. The Role: We are seeking a highly skilled and experienced Staff Software Engineer to join our Test Platform team. In this role, you will have the opportunity to directly impact the design and architecture of our Developer Platform and tooling that enables SoFi engineers to create and deliver high-quality solutions. You will collaborate and partner with a curious team of engineers to design and deliver solutions that raise the testing and reliability standards for our backend and web applications. This new team will be focused on building a green field project to enable autonomous testing for our AI driven SDLC, leveraging cutting edge technologies to deliver a highly resilient and thorough platform. If you are a seasoned Staff Software Engineer with a passion for building products that just work and enabling developers to build reliable services, and a strong background in distributed systems, we invite you to apply for this exciting and new opportunity. What You’ll Do: Provide technical leadership for initiatives in Testing and Reliability, with a focus on integrating AI-driven automation and autonomous testing practices. Collaborate with product engineering teams to understand requirements and design platform capabilities that are efficient, robust, and developer-friendly. Architect and implement solution
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX or San Francisco, CA About the Role AI inference is becoming core infrastructure. Every serious application will need access to many models, across many providers, with reliability, observability, security, cost control, and routing built in from the start. AI Gateway is Cloudflare’s bet that this layer should exist at the network edge: close to users, close to compute, and simple enough that a developer can adopt i
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Senior Software Engineer. About the job This position is needed to design, build, and optimize the core signalling infrastructure that powers real-time video communications for our customers. You will play a key role in ensuring high performance, reliability, and scalability of our video platform, enabling seamless and secure video experiences. Responsibilities In this role, you’ll: Design, implement, and maintain video signalling protocols and server components for real-time video calls (e.g., WebRTC, SIP, RTCP/RTP) in a highly scalable distributed system. Collaborate with cross-functional distributed teams and various stakeholders to deliver high-performance, low-latency media experiences. Ensure secure transmission and compliance with industry best practices (e.g., end-to-end encryption, privacy standards). Contribute to architectural decisions and code reviews, mentoring junior engineers as needed. Stay current with
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as our next Staff engineer (L4), Twilio’s Segment team. About the job As a Staff Engineer on the Twilio Segment Data platform/ pipelines team, you’ll build and scale systems that process several hundred thousands of data points per second. You will lead the development of high-scale ingestion and data processing systems You'll be designing, operating and maintaining complex distributed systems, ensuring reliability, performance, and cost-efficiency while querying petabytes of data for our customer data platform (CDP). Responsibilities In this role, you’ll: Design and deliver robust, high-scale routing experiences for the Data platform/ pipelines team for Twilio Segment. Ship features that opt for high availability and throughput with eventual consistency Collaborate with engineering and product leads, as well as teams across Twilio Segment Support the reliability and security of the platform Build and optimize globally available and high
As the Head of Engineering Operations, you will be a key leadership partner to the CTO and the engineering leadership team, owning the operational backbone of our global R&D organization. In this role, you will serve as a critical bridge to other cross-functional groups, driving operational excellence, organizational effectiveness, and strategic initiatives. We are looking for a strategic leader who can translate company strategy into actionable engineering plans while building lightweight operating models that accelerate velocity and maintain high quality. You will empower our engineering organization to thrive and deliver exceptional impact at scale. This role is based in our San Francisco or Vancouver office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Translate overarching company strategy into actionable engineering and operating plans in close collaboration with cross-functional partners. Design and build lightweight operating models, frameworks, and processes to optimize engineering execution and delivery, integrating evolving AI agentic engineering workflows. Hire, mentor, and lead a high-performing, global team of engineering operations professionals and Technical Program Managers to steer complex horizontal programs. Oversee financial and resource governance, including budget management, program spend, headcount strategy, and vendor/tooling allocations across public cloud and LLMs. Define and own the engineering metrics program (KPIs covering delivery, quality, reliability, efficiency, and capacity) to deliver data-driven insights to leadership. Serve as a trusted proxy and connective tissue for the CTO and R&D
Role Summary: Datadog is seeking a Staff Software Engineer to help shape the future of our Bring Your Own Cloud (BYOC) Logs offering by unifying observability pipelines with log management software that customers deploy and manage in their own infrastructure. This role will focus on building and scaling systems that process, route, and store high-volume observability data within customer-managed infrastructure. You will operate as a hands-on technical leader, driving architecture, cross-team delivery, and product direction across a complex and evolving space. This is a high-impact opportunity to influence product strategy, mentor engineers, and solve deeply technical challenges at scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Make customer-controlled deployments feel like a managed Datadog product: deployment, upgrades, configuration, observability, diagnostics, reliability, and secure operation across diverse customer cloud environments Build and scale high-throughput systems for log processing, routing, and transformation across distributed environments Lead cross-team initiatives, aligning engineers, product managers, and stakeholders to deliver complex, multi-team projects Design and implement software that runs reliably that customers deploy and operate within their own cloud infrastructure. Improve system performance, scalability, and cost efficiency through thoughtful trade-off analysis and capacity planning Contribute hands-on to critical code paths, debugging, and deployment challenges in customer environments Who You Are: You have significant experience building software that is installed, deployed, and operated in customer environments rather than only as a fully managed SaaS service. You have strong expertise in distributed systems,
This role is part of Datadog’s Security Agent team, which powers critical security capabilities across Workload Protection, Vulnerability Management, Cloud Security products, and other emerging security offerings. As a Staff Software Engineer, you will lead the design and development of low-level Linux instrumentation and runtime security technologies that help customers detect threats, monitor system activity, and protect cloud-native workloads at scale. You will work on complex technical challenges involving eBPF, Linux kernel internals, performance-sensitive systems, and large-scale data collection while influencing technical direction across multiple product teams. This role offers significant ownership, broad organizational impact, and the opportunity to shape the future of Datadog’s security platform. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the architecture and development of security agent capabilities that power runtime threat detection and workload protection across Datadog Security products. Design and build reusable eBPF-based monitoring functionality for process, file, and network visibility within Linux environments. Drive end-to-end delivery of new features, from technical strategy and design through implementation, testing, and rollout. Establish and evolve testing methodologies that improve platform coverage, detection quality, reliability, and performance. Partner with product, security, infrastructure, and engineering teams to deliver shared platform capabilities used across multiple Datadog products. Provide technical leadership by influencing engineering direction, mentoring peers, and helping resolve complex cross-functional challenges. Who You Are: You have significant experience building software in Linux environments,
We are looking for a Senior Software Engineer to help us take REDAPL, our Referential Data Platform, to the next level. REDAPL is Datadog's main platform for tracking our customers' infrastructure resources and relationships. The platform enables products where customers can understand, keep track of, and gain insights into their infrastructure related to performance, cost, security, and more. Many Datadog products use REDAPL today such Cloud Security Posture Management, Resource Catalog, Cloud Cost Management, and Service Catalog and others - REDAPL ingests more than 4.5mil updates/second. As a Senior Engineer, you will drive, lead and collaborate on projects both inside and outside the platform. You can expect to contribute to key technical decisions relating to our data ingestion, processing, and query pipelines. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Build a query engine that supports efficient relationship traversals for our most demanding workloads. Contribute to design and drive high-priority, high-visibility projects to increase the platform's value, resilience, and scalability across multiple teams. Lead and guide other engineers through architectural platform decisions Identify potential system risks and trends in reliability and design solutions to address them Provide input on prioritizing engineering-led initiatives in short- and long-term planning and roadmaps Collaborate with internal product teams to understand their requirements and how we plan for their product growth as they integrate and depend on REDAPL Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or related scientific field or equivalent experience You have worked extensively with multiple types of data stores You have contributed to in
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime