Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

RS
Redwood Software
📍 Ontario• Full-time• C$132K – C$165K/yr
18 days ago

OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT As a Senior Full Stack Software Developer, you will be responsible for leading the design, development, and delivery of scalable full-stack applications, shaping system architecture, and driving engineering excellence across Redwood’s automation and SaaS platforms. Design, develop, and implement scalable, secure, and high-performance full-stack applications using Java, JavaScript, and related technologies Architect and build backend services, APIs, and microservices with a focus on scalability, reliability, and maintainability Develop responsive, accessible, and high-quality front-end user experiences Partner with product managers and stakeholders to define technical strategy and translate business requirements into system designs Own and contribute across the full software development lifecycle, from architecture and design to deployment and optimization Establish and promote best practices in coding, testing, observability, performance optimization, and AI usage Lead architectural discussions a

javascripttypescriptjava
View job →
S
SonicWall
📍 Bengaluru• Full-time
18 days ago

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . Job Summary We are seeking a experienced Oracle DBA to manage and optimize our Oracle E-Business Suite 12.1.3 environment, with a roadmap to upgrade to 12.2 or migrate to Oracle Cloud. The ideal candidate will have deep expertise in Oracle database administration, EBS application support, and performance tuning, along with familiarity in tools like SharePlex and Hyperion. This role is critical to supporting financial closure cycles and ensuring database reliability and performance. Key Responsibilities: Administer and maintain Oracle EBS R12.1.3 environments, including patching, cloning, backups, and upgrades. Lead and support the upgrade path to Oracle EBS 12.2 or migration to Oracle Cloud Infrastructure (OCI). Support month-end, quarter-end, and year-end financial closure activities, ensuring system stability and performance. Perform database cloning for development, testing, and troubleshooting purposes. Troubleshoot and optimize SQL queries and PL/SQL code for performance and reliability. Monitor and optimize database performance, availability, and scalability. Manage concurrent managers, workflow serv

sqlawsazure
View job →
S
SonicWall
📍 Pune• Full-time
18 days ago

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . Job Summary We are seeking a experienced Oracle DBA to manage and optimize our Oracle E-Business Suite 12.1.3 environment, with a roadmap to upgrade to 12.2 or migrate to Oracle Cloud. The ideal candidate will have deep expertise in Oracle database administration, EBS application support, and performance tuning, along with familiarity in tools like SharePlex and Hyperion. This role is critical to supporting financial closure cycles and ensuring database reliability and performance. Key Responsibilities: Administer and maintain Oracle EBS R12.1.3 environments, including patching, cloning, backups, and upgrades. Lead and support the upgrade path to Oracle EBS 12.2 or migration to Oracle Cloud Infrastructure (OCI). Support month-end, quarter-end, and year-end financial closure activities, ensuring system stability and performance. Perform database cloning for development, testing, and troubleshooting purposes. Troubleshoot and optimize SQL queries and PL/SQL code for performance and reliability. Monitor and optimize database performance, availability, and scalability. Manage concurrent managers, workflow serv

sqlawsazure
View job →
G
Godaddy
📍 United States• Full-time• From $128K/yr
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This position may be a hybrid or fully remote position, as decided by your manager. If designated as hybrid, you’ll divide your time between working remotely from your home and an office location, so you should live within commuting distance. If designated as remote, you’ll be working remotely from your home and may occasionally visit a GoDaddy office to meet with your team for events or meetings. Your hiring manager can share more about this role’s hybrid or remote designation. This position is not eligible to be performed in Alaska, Mississippi, North Dakota, or the Virgin Islands. GoDaddy is not currently considering candidates for this role in California, Seattle, or NYC. Join Our Team Join a team powering secure, scalable email services for millions of customers worldwide! As part of GoDaddy's Professional Email team, you'll solve complex challenges in distributed systems, cloud infrastructure, security, and AI while modernizing critical platforms that businesses rely on every day. If you enjoy owning impactful systems, working across a diverse technology stack, and building innovative solutions at scale, you'll feel right at home here. What you'll get to do... Design, build, and maintain highly available, scalable APIs and services used by millions of customers Deploy, manage, and optimize cloud infrastructure in AWS Architect and implement modern solutions that improve performance, reliability, and security Leverage AI technologies to enhance development workflows and create innovative customer experiences Monitor, troubleshoot, and resolve complex production issues using modern observability and monitoring tools Drive continuous improvement through automation, modernization, and operational excellence C

pythonawsci/cd
View job →
M
Modal
📍 New York• Full-time
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for PhD research interns with strong research experience in reinforcement learning, machine learning, and foundation models, including large language and multimodal models, to join our research team. This internship is well suited to candidates interested in improving existing methods and developing new techniques for large-scale model training, optimization, and inference, extending models to long-context and long-horizon tasks, and improving inference-time efficiency, reliability, and robustness in high-stakes real-world deployments. Preferred Qualifications: Currently pursuing a PhD in computer science, machine learning, or a related field. A demonstrated record of research in reinforcement learning, machine learning, foundation models, or related areas. Experience developing and evaluating large-scale models or machine learning systems. Familiari

restmachine learningai
View job →

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're building a platform that covers the whole life of an LLM: training it, deploying it, and observing it in production. We already run multi-node training, elastic inference, sandboxes, and distributed volumes, and we control the infrastructure underneath. We’re looking for research depth in post-training to sit alongside our systems and product work. What you'll do: We are looking for research scientists with a strong track record in reinforcement learning, machine learning, and foundation models, including large language and multimodal models, to join our research team. This role is well suited to candidates interested in improving existing methods and developing new techniques for large-scale model training, optimization, and inference, extending models to long-context and long-horizon tasks, and improving inference-time efficiency, reliability, and robustnes

restmachine learningai
View job →
S
Supabase
📍 Remote• Full-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the role We're looking for an engineer to drive the evolution of OrioleDB and its collaboration with the upstream PostgreSQL community. OrioleDB is a next-generation storage engine for PostgreSQL, and this role sits at the intersection of database internals, open source community work, and Supabase's managed Postgres platform. This role requires strong working overlap with Americas time zones due to production support and on-call responsibilities. What you'll work on New OrioleDB features. Design and implement new capabilities in OrioleDB — for example, native index access methods (GiST, GIN, HNSW for pgvector), disaster recovery tooling, and other storage-engine-level features that expand what OrioleDB can do. Stability. Strengthen OrioleDB's reliability through deeper test coverage, fault injection, crash and recovery testing, and improvements to CI infrastructure. Ensure regressions are caught early and that OrioleDB behaves predictably under stress, replication, and failure scenarios. Upstream collaboration. Contribute changes directly to PostgreSQL core. Part of OrioleDB lives as a patch on top of PostgreSQL — moving the right pieces upstream shrinks what we maintain ourselves and benefits the wider community. This happens through PostgreSQL's open development process: the pgsql-hackers mailing list, public code review, and commitfests. Supabase integration. Work with Supabase's Postgres team to ensure OrioleDB fits naturally into Supabase's managed offering and roadmap. You Will: Design, implement, and test new OrioleDB features and integrate them cleanly with PostgreSQL's planner, executor, and surrounding subsystems. Build out and maintain test infrastructure: regression suites, f

sqlpostgresqlai
View job →

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Role: Senior Technical Support Specialist (Enterprise Agentic AI) Company: Ema Unlimited Inc. Location: London Employment Type: Full-time/Remote 1. About Ema (Why Ema) Ema is building the world’s first Universal AI Employee — a production-grade agentic AI platform that automates real enterprise workflows across HR, IT, Finance, and Operations. Ema’s customers do not run demos. They replace mission-critical, manual business processes with agentic AI systems that operate across multiple SaaS tools, APIs, and human-in-the-loop workflows. In this world, support is not reactive . Support is production reliability, trust preservation, and system learning . At Ema, Senior Technical Support Specialists are operators of live AI systems , not ticket handlers. 2. Role Overview The Senior Support Engineer owns the health, reliability, and trustworthiness of Ema’s deployed agentic AI systems in production. This role sits at the intersection of: AI behavior Workflow orchestration Enterprise integrations Customer trust Engineering feedback loops This is: ❌ Not L1 / call-center support, ❌ Not a “just escalate to engineering” role, ❌ Not reactive firefighting only This is : A senior technical escal

reactawsazure
View job →
N
Nextdoor
📍 San Francisco• Full-time
1mo ago

#TeamNextdoor Nextdoor is where you connect to the neighborhoods that matter to you so you can belong. Our purpose is to cultivate a kinder world where everyone has a neighborhood they can rely on. Neighbors around the world turn to Nextdoor daily to receive trusted information, give and get help, get things done, and build real-world connections with those nearby — neighbors, businesses, and public services. Today, neighbors rely on Nextdoor in more than 300,000 neighborhoods across 11 countries. Meet Your Future Neighbors As a Software Engineer at Nextdoor, you’ll join a focused team of developers, product managers, and designers who are passionate about using technology to cultivate a kinder world where everyone has a neighbor they can rely on. We are a small team of engineers that wear multiple hats and work across different languages and services to deliver value to our members. We care about moving fast and delivering impact, without compromising on quality and reliability. You will have the opportunity to learn from your co-workers and teach them. As a team, we will make each other better and build great software. What You’ll Bring To The Team If you didn't see an opportunity listed that looked right for you, we would still love to hear from you and consider you for other opportunities! At Nextdoor, we empower our employees to build stronger local communities. To create a platform where all feel welcome, we want our workforce to reflect the diversity of the neighbors we seek to serve. We encourage everyone interested in our purpose to apply. We do not discriminate on the basis of race, gender, religion, sexual orientation, age, or any other trait that unfairly targets a group of people. In accordance with the San Francisco Fair Chance Ordinance, we always consider qualified applicants with arrest and conviction records. #LI-DNI

restairust
View job →
F
Figma
📍 Ca New York• Full-time• From $185K/yr
1mo ago

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! Figma seeks an AI-Native Performance TPM to own performance prevention, diagnostics, and safe rollout for both our flagship products and next-gen AI features. This platform-level role spans desktop, browser and mobile; requires deep experience with performance testing (load, stress, endurance, interference), observability, and incident response. You’ll partner across Product, Platform, Performance Program and Product Support. Join us to keep Figma snappy quick! This is a full time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: This is a specialist Platform TPM owning horizontal, high-visibility performance programs that span: Flagship product performance (load times, FPS, memory) and AI features (model inference latency, throughput, cost, and reliability) Desktop, browser, native mobile (React Native / WebView), and WASM contexts Cross-org programs: observability/telemetry, regression prevention (performance CI), xfn performance forum, SEV mitigation, safe rollout of AI capabilities, and CE Planning (customer engineering / enterprise readiness) We’d love to hear from you if you have: 5+ years in performance engineering, performance TPM, platform TPM, or SRE with hands-on experience shipping performance programs for SaaS products Demonstrated experience with load, stress, performance or scalability testing and new-build comparisons Deep familiarity with web performance (FCP, LCP), rendering/FPS, WASM memory, mobile profiling (Xcode I

reactawsci/cd
View job →
O
1mo ago

About the Team OpenAI’s Forward Deployed Engineering (FDE) team turns research breakthroughs into production-grade systems. We embed deeply with customers to solve high-leverage problems and act as the delivery engine for our most complex large-scale engagements. We move quickly from prototype to production and surface reusable patterns that shape our platform. We operate at the intersection of deployment and development – working closely with OpenAI Research, Product and Partnerships. About the Role As a Technical Deployment Lead (TDL), you will define how OpenAI delivers complex systems to Semiconductor customers. You will own how solutions are scoped, built, shipped, and adopted across high-value engineering workflows such as RTL design, verification, and physical implementation. You’ll translate business outcomes into a technical plan, run day-to-day execution across FDEs, Researchers, and Customer Engineers, and partner with customer teams to ensure delivery supports their goals. You will focus on the semiconductor vertical to deploy next-generation AI capabilities. You will own delivery end-to-end: embedding with customers to map workflows and success criteria, ensuring components ship on time, and leading readiness and change management for adoption. You’ll track progress, manage dependencies, make sequencing decisions, and drive 0→1 prototypes through MVP and scale. You will also share field insights with Product and Research to guide roadmap and priorities. Success will be measured first and foremost by impact - deployments that deliver measurable value against customer goals, drive adoption, and become critical to their workflows. Additional measures of success include delivery reliability (milestones hit, low reopen/churn), operating leverage (patterns reused across deployments), judgment under pressure, and product impact (field signal that shifts roadmaps/architectures). This is a high-trust, high-autonomy role. Success requires deep technical project m

awsrestai
View job →
O
1mo ago

About the team OpenAI’s Forward Deployed Engineering team partners with customers to turn research breakthroughs into production systems. We operate at the intersection of customer delivery and core platform development. About the Role This is a founding role. As a Technical Deployment Lead (TDL), you will define how OpenAI delivers complex systems to customers. You will own how they are built, shipped, and adopted. You’ll translate business outcomes into a technical plan, run day-to-day execution across FDEs, Researchers, and Customer Engineers, and partner with customer teams to ensure delivery supports their goals. This is not a management role, however you'll own delivery end-to-end: embedding with customers to map workflows and success criteria, ensuring components ship on time, and leading readiness and change management for adoption. You’ll track progress, manage dependencies, make sequencing decisions, and drive 0→1 prototypes through MVP and scale. You will also share field insights with Product and Research to guide roadmap and priorities. Success will be measured first and foremost by impact - deployments that deliver measurable value against customer goals, drive adoption, and become critical to their workflows. Additional measures of success include delivery reliability (milestones hit, low reopen/churn), operating leverage (patterns reused across deployments), judgment under pressure, and product impact (field signal that shifts roadmaps/architectures). This is a high-trust, high-autonomy role. Success requires deep technical project management expertise, extreme ownership of outcomes, and an ability to immerse in customer workflows and partner with customer teams to solve complex engineering problems at pace. This role is based in Singapore. We use a hybrid model of 3 days in office and offer relocation assistance. Travel up to 25-50% is required. In this role, you will: Own the technical delivery plan for multiple interdependent workstreams. Translat

awsrestai
View job →
O
1mo ago

About the team OpenAI’s Forward Deployed Engineering team partners with customers to turn research breakthroughs into production systems. We operate at the intersection of customer delivery and core platform development. About the Role This is a founding role. As a Technical Deployment Lead (TDL), you will define how OpenAI delivers complex systems to customers. You will own how they are built, shipped, and adopted. You’ll translate business outcomes into a technical plan, run day-to-day execution across FDEs, Researchers, and Customer Engineers, and partner with customer teams to ensure delivery supports their goals. This is not a management role, however you'll own delivery end-to-end: embedding with customers to map workflows and success criteria, ensuring components ship on time, and leading readiness and change management for adoption. You’ll track progress, manage dependencies, make sequencing decisions, and drive 0→1 prototypes through MVP and scale. You will also share field insights with Product and Research to guide roadmap and priorities. Success will be measured first and foremost by impact - deployments that deliver measurable value against customer goals, drive adoption, and become critical to their workflows. Additional measures of success include delivery reliability (milestones hit, low reopen/churn), operating leverage (patterns reused across deployments), judgment under pressure, and product impact (field signal that shifts roadmaps/architectures). This is a high-trust, high-autonomy role. Success requires deep technical project management expertise, extreme ownership of outcomes, and an ability to immerse in customer workflows and partner with customer teams to solve complex engineering problems at pace. This role is based in Tokyo. We use a hybrid model of 3 days in office and offer relocation assistance. Travel up to 25-50% is required. To succeed in this position you must be fully bilingual—fluent in both Japanese and English (spoken and writt

awsrestai
View job →
S
Snowflake
📍 Bellevue• Full-time
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a talented Tech Lead Manager (TLM) to lead Snowflake’s System Under Test (SUT) team, responsible for evolving how Snowflake engineers test the Snowflake product locally and at scale in CI. The SUT empowers Snowflake engineers by delivering a reliable, low-latency and cost-efficient developer experience across a high-growth, high-demand surface area. As the TLM for SUT, you will lead a small and highly technical team at the intersection of CI, developer infrastructure, and product engineering. You will set direction, drive execution, and partner broadly across Engineering Systems and product teams to deliver a more reliable, faster, and more maintainable test platform for Snowflake’s engineers. In this role, you will: Lead, coach, and grow the SUT team while creating a high-energy, cohesive environment with strong planning, ownership, and career development. Own the roadmap and execution for SUT rollout across development environments, CI and AI workflows. Drive measurable improvements in startup reliability, latency, and cost, using clear SLOs, dashboards, and operational metrics to guide decisions and raise the bar on execution. Serve as the technical anchor for the SUT domain, shaping architecture and guiding the evolution from legacy systems to a composable

kubernetesaigo
View job →
S
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. About the team The Apps & Experiences Platform team is powering systems and services for all our user facing apps including Snowsight , Snowflake Intelligence , and new mobile apps . Our mission is to craft innovative backend services, features, tools, infrastructure, and AI tooling that bring such products to life with delightful user experiences. As part of our team, you'll dive into a mix of building & managing platform infrastructure, and building AI self-serve tools to support the platform and its developer’s needs. We're passionate about building a platform that is highly reliable, available, maintainable, and scalable. We are a high growth AI data cloud company and we are looking for exceptional talent like you to help build and grow our infrastructure to scale us to the next level. AS TECHNICAL LEAD & MANAGER FOR APPS & EXPERIENCES STORAGE TEAM, YOU WILL: Own the storage team’s technical strategy and execution , driving projects from initial idea formulation and detailed system design to high-quality implementation and successful deployment. Provide deep technical oversight by actively participating in design reviews, architecture discussions, and drilling into complex system implementations to ensure reliability and scalability. Serve as a key techn

redisawsazure
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Top cities for Reliability Engineer

City links are canonicalized and require at least 20 current jobs.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.