About the Team The Safety Systems org is responsible for various safety work to ensure our best models can be safely deployed to the real world to benefit the society and is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Safety Engineering team builds the platforms and tools that make OpenAI’s models safe to use in the real world. We partner closely with researchers, product teams, and policy to turn safety ideas into reliable, scalable systems: measuring risk, enforcing safeguards, and continuously improving how models behave in production. Our work sits at the intersection of product engineering, data, and AI, and directly shapes how millions of people experience OpenAI’s technology. About the Role We’re looking for a self-starter engineer who loves building products in an iterative, fast-moving environment—especially internal tools that unlock real-world impact. In this role, you’ll build full-stack tooling for our Safety Systems teams that directly improves the safety and reliability of OpenAI’s models, including in sensitive areas like mental health and other vulnerable-user protections. Your work will increase the team’s velocity in identifying and fixing safety issues and help tighten the feedback loop between policy, data, and the model training cycle. In this role, you will: Own the end-to-end development of internal tools that help improve the safety of OpenAI’s models (with a focus on areas like mental health and other vulnerable-user protections) Partner closely with Safety Systems researchers, engineers, and model policy creators to understand workflows, pain points, and requirements—and translate them into durable product solutions Build full-stack experiences to support core model policy workflows, such as labeling and inspecting data, analyzing and reviewing failure cases, and surfacing insights for iteration Optimize internal applications f
Jobiba hiring network
Reliability Engineer Jobs
2,028 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. As the Director of Engineering at Smartsheet India, you will build capabilities to empower the world's largest companies to transform their approach to work. You will guide teams that own the grid ecosystem - defining how data linking, synchronization, and grid infrastructure evolve as a cohesive platform. You will ensure architectural decisions are coherent and avoid fragmentation. You will be willing to challenge technical choices. Platform Reliability & Operational Excellence: You will be accountable for the availability and performance of foundational services that other teams depend on. Drive a high bar for on-call health, incident response, and SLA/SLO definition across all the services. You will manage cross-pillar/cross-domain dependencies, negotiate API contracts, and prevent the grid ecosystem from becoming a delivery bottleneck. You will balance the needs of user-facing product features with infrastructural stability and operational health. You will ensure career growth paths are clear for engineers across that spectrum, and develop a strong sense of customer centricity and pillar identity for your teams. You will be comfortable accepting responsibility for impact, service availability, and the effectiveness of your teams. You will be comfortable being at the forefront of AI adoption for delivery and operations, leaning in and helping the team leverage AI for optimum delivery in their ways of working. You are passionate about continuous improvement and have built learning organizations that keep up w
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We’re hiring talented Software Engineers for the Snowflake Dynamic Tables team in Berlin, Germany. Join us to build the next generation data platform that enables customers to transform data with declarative SQL while maintaining control over cost, latency, and throughput. We are looking for strong engineers who are enthusiastic about building new cutting-edge technologies, who look forward to tackling complex database problems, and pick up and understand deep technical areas quickly. You will work alongside seasoned engineers and grow in your scope and influence. AS A SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Work with a talented and collaborative team of engineers and Product Managers in our globally distributed team to design and build Dynamic Tables capabilities Design, implement, support, and evolve new features and performance improvements Help shape technical and product direction with senior teammates Break ambiguous problems down, weigh the tradeoffs, and make technical and product decisions Analyze and solve performance, correctness, and fault-tolerance challenges at scale Dig into unfamiliar parts of a large system to root cause and solve problems Ensure operational readiness of what you build and help meet the commitments to our customers regarding reliability, a
About the Team The Astral team builds high-performance developer tools to power the future of programming, at OpenAI and beyond, including Ruff, uv, and ty. The Astral toolchain sees hundreds of millions of installs per month and powers hundreds of millions of package downloads per day for the Python ecosystem. As a team, we are building on those foundations to continue solving impactful tooling problems as programming evolves. About the Role We are looking for an experienced software engineer to build next-generation programming language tooling. If you like writing high-performance Rust, it could be a good fit; if you like thinking about the future of programming, it could also be a good fit. Strong candidates tend to have deep experience with Rust, Python, open source, compilers, or developer tools — but few candidates are deep in all of these areas, and we've hired candidates without prior Rust or Python experience. In this role, you will: Design and implement features in Astral’s existing open source projects (Ruff, uv, ty, and python-build-standalone, and more). Support Astral’s open source projects as a maintainer, triaging user issues, reviewing pull requests, and participating in community discussions. Evolve the Astral toolchain to accelerate development velocity at OpenAI. Build entirely new tools, in entirely different programming ecosystems, to power the future of agentic software development. Your background might look something like: 5+ years of professional engineering experience, excluding internships, in relevant engineering roles. High agency and comfort operating in a fast-moving environment, with strong ownership of security, reliability, and operational excellence. Strong developer empathy and communication skills, including experience maintaining open source projects. Exceptional systems engineering fundamentals and a track record of leading complex projects from ambiguous problem statements through to user impact. Proficiency in one or more s
About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. In this role you will: As a Hardware Test Engineer, you will work on Machine Learning/AI hardware system projects to craft the solutions for current and future data center deployments. You will bring a strong understanding of hardware system testing, excellent project management skills, and the ability to collaborate across multiple teams to ensure efficient lab operations. You will be responsible for designing, implementing, and executing comprehensive test plans that ensure the reliability, performance, and scalability of our supercomputing hardware systems. You will develop detailed test plans and methodologies tailored to hardware components, including processors, memory modules, custom accelerators and interconnects. You will collaborate with hardware design, manufacturing, firmware teams and vendors to identify, analyze, and resolve issues affecting hardware, power, thermal and high-speed interconnects. You will perform in-depth debugging on the hardware system Excellent analytical skills to diagnose hardware issues, troubleshoot problems, and propose solutions. Ability to interpret complex test data, identify trends, and draw meaningful conclusions. High-speed links, with a focus on SerDes (Serializer/Deserializer) technology to assess signal integrity, error rates, and overall link performance. You will collaborate with the lab manager to maintain the equipment and hardware systems, including oscilloscopes, thermal test chambers, liquid cooling systems, and other mea
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Team The Devices Platform team's mandate is to lay the foundation of Nuro's onboard software for our sensor and compute platform, including device drivers, inter-device protocols and pipelines, and device runtime APIs. Sensors and compute hardware are the eyes, ears, and brains of our self-driving robots. We are creating the hardware-agnostic platform to be used by the perception and autonomy SW stack, and to realize the full potential of our sensor and compute HW in reliability, quality, and performance. The projects we work on are high impact and high visibility within Nuro. This team is also responsible for working with internal stakeholders and external suppliers to define, evaluate, integrate the next generation HW platform for Nuro's products and to build the necessary tooling to assist continuous testing and validation. About the Work Design and develop sensor and compute systems for robotics Architect and/or deploy Nuro
Job Details: Job Description: Join Intel-and build a better tomorrow. Intel creates an environment where employees can prosper while creating the innovative technologies that make amazing possible. As a Facilities Operations ,our scope is vast and includes operating and maintaining all Intel sites, offices, labs, data centers and factories globally as well as onsite services and experiences that help employees stay safe and productive. Our Mechanical Engineers make a big impact by supporting our daily tactical efforts in safety, reliability, and environmental objectives. The Role and Impact As a Facilities Mechanical Engineer, you will play a critical role in ensuring the availability, reliability, and maintenance of mechanical systems (Oil Free Air, HVAC, Exhaust, Process Vacuum, PCW, Wet system, Fire Life Safety System and etc) essential to Intel's manufacturing, clean room, and research and development activities. You'll focus on designing, maintaining, and troubleshooting mechanical systems to support daily operations, system reliability, safety and environmental objectives. Your expertise will directly contribute to maintaining optimal operational performance across Intel's facilities. Key Responsibilities - Own, sustain, and improve system reliability, capability, capacity, operational troubleshooting, system upgrade evaluation and optimization of mechanical systems. - Develop engineering scopes of work for mechanical projects and conduct design reviews and evaluations. - Partner with internal customers and external service providers to ensure systems are maintained and operated within specified limits. - Support environmental compliance for facilities mechanical systems, ensuring adherence to safety and reliability standards. Compliance to EHS / regulatory / insurance audit requirements (Fire protection / Life safety system) - Conduct feasibility studies
About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. About the role We are looking for an Operating Systems Engineer to build and harden the OS foundations for OpenAI products. We are especially interested in experienced, passionate, and innovative operating systems developers who thrive on building foundational platform software and solving hard problems in security, privacy, performance, power, and reliability. You will work across the OS kernel, core OS services, security and privacy primitives, performance and power, and the frameworks that connect applications and UI to the system. This role emphasizes deep debugging and systems ownership from development through production. You will collaborate closely with embedded, firmware, hardware, application, and product engineering teams. Experience with hardware bring-up is a plus, but not required. What you will do Work on end-to-end OS capabilities spanning the OS kernel, userspace services, application frameworks, UI toolkits, and application-facing APIs. Develop, integrate, and maintain OS components, both kernel-bound and in userspace, including scheduling, memory management, filesystems, drivers, IPC/RPC mechanisms, and security-relevant subsystems. Build and maintain core OS services and daemons (init, service management, device discovery, networking primitives, time, logging, update hooks, crash handling, and so on). Design and implement security and privacy mechanisms: Secure boot and measured boot integration points (where applicable). Mandatory access control and sandboxing. Secrets management, secure storage, key handling, and least-privilege service design. Privacy-preserving telemetry, data minimization, and user-consent oriented system behaviors. Establish a perfo
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Senior Software Engineers lead and mentor engineers, delivering high-value products for our customers and infrastructure that enables our business to scale. As a Senior Software Engineer, you'll be responsible for setting technical direction to enable our product and infrastructure to scale with our business, driving complex projects across our technical stack, and mentoring our talented engineering team. Your past experience will be leveraged to enable and accelerate Vanta's growth. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We're growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and to contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! The Authoring Experience team owns how Vanta tests are authored and deployed to customers, from the APIs and services underneath to the experiences built on top. As a Senior Backend Engineer on this team, you will own the services and public-facing customer APIs that turn Vanta's automation platform into products customers touch directly, stitching together systems across the platform to serve them well. What you’ll do as a Senior Fullstack Software Engineer at Vanta: Own the architecture and rollout strategy for critical platform services, from initial design through large-scale rollout and migration. Set the technical bar for how we design, publish, and maintain customer-facing APIs while owning things like security, reliability, and developer experience. Stitch together
About the Team OpenAI's People team helps hire, develop, and support the people building safe and beneficial AGI. Within that team, People Systems builds the technical foundation that enables our HR, recruiting, payroll, benefits, and performance operations to scale with quality, speed, and rigor. We work at the intersection of HR systems, software engineering, and internal tooling. Our goal is not just to keep core systems running, but to build durable technical leverage for the company. About the Role We're hiring a Workday Engineer to help design, build, and operate the systems that power critical people workflows at OpenAI. This is a highly technical role for someone who combines strong Workday expertise with real engineering fluency. You'll build reliable integrations, improve system architecture, automate complex workflows, and help connect Workday to internal tools, external platforms, and emerging AI-driven systems. You should be comfortable going beyond configuration work. We're looking for someone who can reason through ambiguous systems problems, write and debug technical solutions, work effectively in Git-based environments, and use modern developer workflows, including CLI-driven tooling, to build and operate with speed and discipline. You'll partner closely with cross-functional teams across People, Finance, Security, and Engineering, including our People Innovations team, to build systems that are secure, scalable, and practical. Some work will involve improving mature production infrastructure; some will involve building entirely new workflows and capabilities from scratch. In this role, you will Design, build, and maintain Workday integrations, applications, and workflow automations across domains such as payroll, benefits, recruiting, performance, and case management Improve the reliability, quality, and scalability of People systems through strong engineering, testing, and operational practices Build technical solutions that connect Workday with i
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Product Auth0 is a developer-friendly identity platform that simplifies authentication and authorization for applications. Designed by developers for developers, we make access to applications safe, secure, and seamless for the more than 100 million daily logins around the world. Our modern approach to identity enables this Tier-Ø global service to deliver convenience, privacy, and security so customers can focus on innovation. Know more about our product at https://auth0.com/ . The Role We are seeking a founding Engineering Manager to lead and bootstrap our newest team: Core Frontier . This team sits at the vital intersection of deep product innovation and the actual customer experience. Your mission is to ensure that the sophisticated features developed across the Core Identity organization (such as Native to Web, Cross-App Access and Custom Token Exchange) are translated into a seamless and intuitive journey for our users. You will act as the champion for a complete and coherent product experience, bridging the gap between complex internal logic and the polished final result our customers interact with every day. As the inaugural manager for Core Frontier, you will lead the hiring and onboarding of a high-caliber team in Bengaluru while establishing the operational rhythms that drive success. You will partner with global engineering leaders to maintain our high standards for security and reliability, ensuring that every release meets the rigor
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Software Engineer II Overview Who is Mastercard? Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. With connections across more than 210 countries and territories, we are building a sustainable world that unlocks priceless possibilities for all. The Fraud Products team (part of O&T) is developing new capabilities for MasterCard's Decision Management Platform, which serves as the core for multiple business solutions to combat fraud and validate cardholder identity. Our patented Java-based platform processes billions of transactions per month in tens of milliseconds using a multi-tiered, message-oriented approach for high performance and availability. MasterCard software engineering teams leverage Agile development principles, advanced development, design and test automation practices, and an obsession over security, reliability, and perfo
At Playlist, life's richest moments happen when people step away from screens to move, connect, explore, and play. We're building the definitive platform for intentional living, connecting people with inspiring experiences in fitness, wellness, and beyond. With popular brands like Mindbody and ClassPass, Playlist empowers businesses and individuals, making it effortless for aspirations to become actions. Join us in reshaping technology's role to foster meaningful, real-world connections. Mindbody equips wellness entrepreneurs with technology to support thriving businesses and create exceptional experiences. Innovation and curiosity drive our culture, connecting businesses and individuals through cutting-edge solutions. Join us if you're passionate about enhancing wellness through technology. The Role You’ll Play You’ll join Mindbody’s External APIs team, which owns the developer-facing platform that partners, integrators, and enterprise customers build on. That includes our Public API, Webhooks, Affiliate API, and the developer portal experience. You’ll help build, ship, and operate the APIs that thousands of wellness businesses use to connect Mindbody to payroll, marketing, booking, and other tools — with a high bar for reliability, backward compatibility, and partner experience. Build and maintain performant backend systems and applications that drive real-world experiences — especially the public and partner APIs studios, aggregators, and strategic accounts depend on Partner with Product, Partnerships and API Support to bring features to life from ideation through deployment, always iterating with the end-user in mind — here, that user is often an external developer or enterprise integrator Contribute to API platform work such as authentication and authorization, versioning, rate limiting, and changes that protect the experience our partners depend on Help keep existing integrations stable while we ship new capabilities — including compatibility across surfaces t
Our mission and customers: We are creating the freedom for SMEs to succeed by delivering Europe's leading finance workspace with banking at its core, augmented by financial tools. We are proud to be rated 4.8 on Trustpilot, based on 55,000+ reviews. Our culture puts customer satisfaction at the core of what we do, as proven by our Net Promoter Score of 75 (more about our culture here). Our journey: Founded in 2017 by Alexandre and Steve, Qonto has grown to 1,600+ Qontoers serving over 600,000+ customers across 8 European countries. We have been profitable since 2023, and we are just getting started. Our beliefs: We hire for skills and potential. With 80+ nationalities, 45% women, of which 56% of women in our leadership team, diversity isn't a program; It's who we are. We've built a discrimination-free hiring process because the best teams are built on merit. AI at Qonto: AI is deeply embedded in how we work (here) - Every Qontoer gets unlimited access to the best AI tools. We want people who experiment without waiting for permission, push AI beyond the obvious, know when to trust it, and when to question it. ------------------------------------------------------------------------------------------------------ ➡️ Mission: Join us as Analytics Engineer x Business Analytics and become the person our Business Analytics teams can build on without a second thought. You will own end-to-end the data models feeding Product, Growth, and Ops Finance analytics — designing scalable dbt models, pushing back on requests that would trade reliability for speed, and helping the team migrate to Omni and a real semantic layer. You will work closely with Jules Jeanroy, our Analytics Engineering Manager, and partner daily with Business Analysts across Product, Growth, and Ops Finance. The team is at a pivotal moment — investing in scalability, cutting technical debt, and building the standards that will define how Analytics Engineering works at Qonto for years to come. ➡️ As an Analyti
Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability. Key Responsibilities MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime