Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

SL
Sumo Logic
📍 Bengaluru• Full-time
18 days ago

Principal Product Manager - Agentic Investigation & Reliability Experiences Sumo Logic is hiring a Principal Product Manager to lead how engineers and operators investigate incidents, understand reliability risk, and act on their operational and security telemetry. The observability category was built around collecting telemetry and giving people tools to navigate it: dashboards, queries, monitors, traces, and alerts. Customer expectations are now shifting. Teams don't just want more dashboards; they want help getting from a signal to a resolution, understanding what's broken, why, what's impacted, and what to do next. As AI agents move into production operations, this role owns how Sumo Logic brings intelligent, agent-assisted investigation and reliability workflows to customers, grounded in evidence, context, and enterprise governance. This is a senior, high-ownership role. It requires genuine observability domain background. You should have lived in this space and understand how monitoring, troubleshooting, and reliability actually work, combined with the ambition to define a new category of experience on top of it. What You Will Own The current data experiences. Log Search, Live Tail, query and query optimization, Metrics Search, Tracing, Dashboards, and the data-experience UI. This is a live, revenue-generating product with real customers, and keeping it strong is part of the job. You own its health, roadmap, and competitiveness today while steering it toward an AI-native future, focusing new investment where it strengthens investigation, speed, and value for both new and power users. The reliability and alerting surface. Monitors, Alerts, SLOs, Scheduled Searches, and the reliability workflows around them. You will own alerting accuracy, noise reduction, and operational health signals both as capabilities customers depend on today and as the foundation for more automated, agent-assisted detection and investigation. The agentic investigation experience. You

reactsqlaws
View job →
SL
18 days ago

Principal Product Manager - Agentic Investigation & Reliability Experiences Sumo Logic is hiring a Principal Product Manager to lead how engineers and operators investigate incidents, understand reliability risk, and act on their operational and security telemetry. The observability category was built around collecting telemetry and giving people tools to navigate it: dashboards, queries, monitors, traces, and alerts. Customer expectations are now shifting. Teams don't just want more dashboards; they want help getting from a signal to a resolution, understanding what's broken, why, what's impacted, and what to do next. As AI agents move into production operations, this role owns how Sumo Logic brings intelligent, agent-assisted investigation and reliability workflows to customers, grounded in evidence, context, and enterprise governance. This is a senior, high-ownership role. It requires genuine observability domain background. You should have lived in this space and understand how monitoring, troubleshooting, and reliability actually work, combined with the ambition to define a new category of experience on top of it. What You Will Own The current data experiences. Log Search, Live Tail, query and query optimization, Metrics Search, Tracing, Dashboards, and the data-experience UI. This is a live, revenue-generating product with real customers, and keeping it strong is part of the job. You own its health, roadmap, and competitiveness today while steering it toward an AI-native future, focusing new investment where it strengthens investigation, speed, and value for both new and power users. The reliability and alerting surface. Monitors, Alerts, SLOs, Scheduled Searches, and the reliability workflows around them. You will own alerting accuracy, noise reduction, and operational health signals both as capabilities customers depend on today and as the foundation for more automated, agent-assisted detection and investigation. The agentic investigation experience. You

reactsqlaws
View job →
T
18 days ago

Toradex is a global company strongly focused on engineering & technology. We’re powered by a diverse & uniquely gifted workforce. We pursue the best people to propel our innovative vision of embedded computing and IoT. If you’re interested in being a driving force at an agile technology company, engineering clever computing solutions & helping other companies bring their products to life, we should talk. Description We are looking for an Integration Platform Specialist to manage company-wide integrations using Workato. The role focuses on building reliable workflows, supporting API-based connections, and working with teams to automate business processes. The position requires strong knowledge of REST APIs, GraphQL, Postman, and the Workato platform. Experience with Workato MCP functionality and coding skills in Python, Ruby, or similar languages are preferred. About You You enjoy solving complex system and process problems with practical, scalable solutions. You can work independently and take ownership of integrations from discovery through deployment and support. You communicate clearly with both technical and non-technical stakeholders. You document your work well and create clear support material for future maintenance. You are curious, hands-on, and willing to investigate issues until you find the root cause. You care about reliability, data quality, security, and a good internal user experience. You are comfortable working across teams and balancing business priorities with technical constraints. Key Responsibilities Own the integration platform roadmap and day-to-day operation, ensuring business-critical automations are reliable, observable, and maintainable. Partner with business and application owners to turn process gaps into pragmatic integration designs and delivery plans. Build Workato recipes, custom connectors, and reusable patterns that reduce manual work and improve data flow between systems. Maintain and modernize existing integrations,

javascriptpythonjava
View job →
S
Sofi
📍 San Francisco• Full-time
1mo ago

Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. Role Description As the Director of Corporate Infrastructure, you will drive efforts to oversee the design, implementation, and operation of our corporate networks. This includes the IT Infrastructure DevOps, Security and SRE teams. You have a deep understanding of infrastructure as code, configuration as code & networking technologies, strong business acumen, lead through data and metrics, and operate with a high level of rigor and accountability. You are responsible for the reliability, scalability, sustainability, and efficiency of SoFi’s network infrastructure. As a member of the Corporate Infrastructure leadership team, you are directly accountable for the teams that are responsible for driving efforts to evolve and build our next-generation corporate network (to include infrastructure and wifi). This includes managing all aspects of our network - engineering and operations to improve user experience and performance, and also supporting the multi-terabit backbone network that interconnects edge PoPs, corporate offices, data centers, and cloud gateways. You will lead a diverse team through highly technical problems to achieve our overall strategic, operational, and financial goals. You will build and lead a high-performance team, displaying technical proficiency to support and scale the

awsrestai
View job →

Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. This position is based at our office in Chennai, India. Appian was built on a culture of in-person collaboration, which we believe is a key driver of our mission to be the best. You will be the product manager working closely with the team whose mission is to strengthen and optimize site infrastructure by delivering essential upgrades, resource efficiency, and scalable solutions. You will be responsible for the direction and roadmap of a component of the Appian Cloud data plane that ensures reliable, high-performance operations for all Appian Cloud customer sites. This role is specifically focused on the cloud-native persistence and messaging layer. You will oversee the backend sub-systems—including technologies like S3 and Redis—that power the platform's internal data plane and core services. This component of the software is not directly user-facing but has strong implications on the scalability and reliability requirements our customers expect. What you will be doing: Prioritize and Define: Work on an agile team to prioritize, define, and ensure the success of infrastructure and managed services features for a high-level strategic roadmap. Stakeholder Collaboration: Prioritize what we should build and when by collaborating with stakeholders on product vision and strategy, while taking customer feedback into account. Technical Discussions: Define how infrastructure features will work through close collaboration with engineers in design sessions an

redisawskubernetes
View job →
J
Jumio
📍 India• Full-time• Remote
18 days ago

Role Purpose: At Jumio, the Software Engineer II (QA) will focus on ensuring the quality and performance of highly scalable web and backend applications. In this role, you will design and implement automated tests for web (Playwright/Selenium) and API-based solutions, leveraging your knowledge of Java or JavaScript. Collaborating closely with development and product teams, you will ensure Jumio's products meet the highest standards of quality and reliability. You’ll have an opportunity to learn and grow in a fast-paced environment while exploring new tools and methodologies. A problem-solving mindset, willingness to innovate, and attention to detail will make you successful in this role. T-Shaped Engineering Expectation: As part of Jumio’s engineering culture, you will adopt a T-shaped engineering approach. In addition to developing expertise in test automation and quality engineering, you will collaborate across the development lifecycle, including understanding software architecture, contributing to design discussions, and ensuring robust and scalable test solutions. Role Value: This role is critical to ensuring the reliability, scalability, and security of Jumio’s products. By building and maintaining automated testing frameworks, you will enable faster releases and higher confidence in the quality of our software. Example Responsibilities: Develop, maintain, and execute automated test scripts for web applications using Playwright, Selenium, or similar automation frameworks. Create and execute API test suites using tools such as Postman, REST Assured, or equivalent testing frameworks. Design and execute functional, regression, integration, and exploratory test cases based on business and technical requirements. Validate application functionality, backend services, APIs, and data flows across different environments. Identify, document, track, and verify defects, working closely with developers to ensure timely resolution. Execute automated test suites as part of C

REMOTEjavascriptjavaci/cd
View job →

Are you ready to do your life’s work at the heart of the autonomous revolution? NVIDIA’s SWQA organization is seeking a world-class Software QA Test and Tool Developer to join our Automotive Platform team, where the code you validate ensures the safety of millions on the road. In this role, you won't just be testing software; you will be architecting the security and reliability of the next generation of intelligent vehicles. We are looking for engineers who are as comfortable navigating low-level product architecture as they are deep-diving into complex product use cases with passion for quality. This is a high-impact, hands-on position focused on our industry-leading automotive products, offering a rare opportunity to influence the core of our tech stack. You will also build the tools and frameworks that define performance standards for systems running on Linux and QNX. What you’ll be doing: Design, execute, and automate comprehensive test cases and test scenarios to validate our automotive platforms using various test methodologies to identify and track actionable defects and track them to closure. Participate in deep-dive reviews of product requirements and technical designs, providing critical feedback to ensure features are built for testability and security from day one. Partner closely with project management, hardware teams, and software developers to provide rigorous technical analysis of bugs and publish data-driven statistical reports for global team members. Architect and maintain a distributed test automation framework capable of managing high-concurrency workloads across an extensive automation farm of hundreds of concurrent systems. Develop sophisticated test libraries and automation solutions to accelerate development cycles and expand automated test coverage for re

pythonlinuxai
View job →
S
Stripe
📍 Seattle• Full-time
1mo ago

Who we are Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet. Stripe Apps connect Stripe with the tools businesses already use, bringing relevant data and workflows into the Stripe Dashboard and third-party systems. The Strategic Apps team owns a portfolio of integrations with enterprise systems that are critical to our users’ workflows—from quote-to-cash and support to reporting. We help users embed Stripe more deeply into the way they run their businesses. What you’ll do We’re looking for a Staff Product Manager to set the strategy and lead execution for Strategic Apps. You will own a portfolio of existing and new integrations, including those with systems such as Salesforce, and determine where Stripe should invest based on user needs, Stripe’s product and industry priorities, and market signals. You will work across Stripe and directly with external partners to deliver reliable, high-quality integrations that drive adoption and business outcomes. Own the vision, product strategy, roadmap, and success metrics for a portfolio of strategic apps, making and communicating clear trade-offs amid significant ambiguity. Identify, prioritize, and validate the highest-value user problems and integration opportunities through direct user research, ecosystem analysis, and quantitative evidence. Lead cross-functional teams across Engineering, Design, Data Science, Finance, Sales, User Support, Marketing, Partner Management, and other product teams to deliver and evolve integrations end to end. Partner with Engineering to define product requirements, technical approach, and quality standards for integrations, including APIs, documentation, edge cases, reliability, security, compliance, and ongoing maintenance. Establish feedback loops

aigoexcel
View job →
H
Hp
📍 Colorado• $83K – $127.8K/yr
10 days ago

Software Developer in Test Description - This role is responsible for ensuring quality, reliability and performance of software applications throughout the development lifecycle primary through software test automation. The role designs, codes, and implements software test automation using appropriate programming languages, frameworks, and tools. The role works closely with cross-functional teams to gather requirements, provide technical insights, and ensure the successful execution of test automation with the main goal of improve quality of the solution. The role also creates and executes comprehensive test plans, test cases, and test scripts based on project specifications. The role sets and provides design guidance to other developers and SQA engineers regarding test automation and the test framework. *Onsite in Ft Collins 4-days a week Responsibilities • Designs quality assurance and test processes for portions of end-user video conferencing/collaboration application, systems software running on android hardware, local, networked, and Internet-based platforms. • Analyzes design and determines test scripts, coding, automation, and integration activities required based on specific objectives and established project guidelines. • Designs and maintains Test automation framework • Executes and writes portions of testing plans, protocols, and documentation for assigned portion of application; identifies and debugs issues with code and suggests changes or improvements. • Identifies opportunities for performance improvements and optimizes code and application performance. • Utilizes latest AI tools and technologies in speeding up test automation • Executes test cases depending on the needs of the project • Provides valuable input into the development of user stories and acceptance criteria, shaping a quality-oriented d

pythonjavadocker
View job →

NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer! What you'll be doing: Primary responsibilities will include building AI/HPC infrastructure for new and existing customers. Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting. Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement. Maintain services once they are live by measuring and monitoring availability, latency, and overall system health. Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements. What we need to see: BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields. At least 5+ years of professional experience in networking fundamentals, Ethernet or InfiniBand World. Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS, etc. Possess solid working knowl

pythonlinuxai
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As an Intern Product Yield Enhancement (PYE) Engineer at Micron Technology Inc., you will assist in taking next generation DRAM devices from design to mass production! PYE Engineers are crucial partners between Research and Development, Manufacturing, Product Engineering and Global Quality teams. Our goal is to offer hands-on and substantial engineering projects that allow interns to gain valuable experience with memory systems and the product life cycle while adding impact to Micron’s business. Our intern programs help prepare engineers for future roles and offer an extraordinary experience that supplements your academic experience! The PYE Engineering role is critical to Micron as it serves the company's charter to be the most cost-effective producer of DRAM memory while delivering the quality and reliability that customers expect from the Micron brand. This position is for a 3-month internship. Responsibilities: Your job duties will include detailed DRAM electrical and physical failure analysis. Identify defects introduced by the semiconductor fabrication process. You will data mine using statistical analysis, AI assisted engineering tools, and automation workflows. Find the root cause of failure mechanisms, communicating results with Management, Product, Process and Test Engineering teams. Complete 3-month internship. Minimum Qualifications: Must be pursuing a minimum of a bachelor’s degree in Electrical Engineering, Computer Engineering,

pythonairecruitment
View job →
D
Datadog
📍 Massachusetts• Full-time• From $152K/yr
1mo ago

The Manager of Networking at Datadog leads the global management of network services across all Datadog offices worldwide. This role is responsible for ensuring seamless, high-performance Wi-Fi and direct internet access in our global offices and conference room technology, supporting a rapidly growing global enterprise. As Datadog continues to grow rapidly, this role plays a critical part in scaling both the team and network infrastructure to meet increasing demand. This is a hybrid role that sits in global headquarters in New York city and requires three days in the office each week with occasional travel to our offices around the world. What You’ll Do: Lead the global network engineering teams, managing both full-time employees and third-party vendors to ensure consistent and high-quality service delivery for Datadog offices worldwide. Oversee the design, deployment and scaling of office network infrastructure, including Wi-Fi (Cisco Meraki) and edge networking devices (Cisco, Palo Alto, Juniper, and others), ensuring these services operate according to Datadog's defined service level objectives. Define and implement standards, policies, and processes for network infrastructure to ensure security, reliability and scalability. Collaborate with IT Security, Enterprise Technology, and Workplace teams to align network services with broader IT and business objectives. Develop and track operational metrics for service availability, network performance, driving continuous improvement and optimization. Who You Are: An experienced people manager with at least 5+ years of leadership experience managing teams of network engineers. Proven expertise in Wi-Fi network engineering, including deep knowledge of Cisco Meraki and edge networking solutions from vendors like Cisco, Palo Alto and Juniper. Experience managing large-scale office technology projects in global enterprises with more than 10 offices and 7,000+ employees, ensuring infrastructure keeps pace with ra

aigorust
View job →

Job Details: Job Description: As a Module Equipment Technician, you will play a vital role in ensuring the smooth operation of advanced manufacturing equipment used in semiconductor production. You will perform hands-on troubleshooting, maintenance, and calibration of electromechanical systems while supporting experiments and equipment modifications that drive cutting-edge advancements. Your work will directly contribute to enhancing equipment reliability and optimizing production efficiency, making an essential impact on Intel's manufacturing operations. Business Group Join Intel's Advanced Packaging Technology Development - Substrate and Wafer Assembly (ATPD: SWA) organization, a leader in delivering innovative and cost-effective substrate packaging solutions. This business group is focused on advancing Intel's capabilities in manufacturing technology to maintain competitive excellence across global markets. As part of this dynamic team, you will support critical equipment and collaborate on continuous improvement initiatives that align with Intel's broader mission of innovation and growth. Key Responsibilities Perform electrical and mechanical troubleshooting to diagnose issues with manufacturing equipment. Execute setup, calibration, corrective and preventative maintenance on production equipment, including wet chemistry, plating, and dry high-vacuum toolsets. Monitor tool performance and analyze data to identify and address equipment-related issues. Collaborate with engineering and support teams to improve equipment reliability and reduce downtime. Document maintenance activities, repair findings, and parts usage using established systems and procedures. Lead or contribute to continuous improvement projects, focusing on equipment optimization and yie

recruitment
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron is seeking a Facilities Water Services UPW and Wastewater Coordinator to support wafer manufacturing by coordinating daily work, maintenance, vendors, contractors, and projects for ultrapure water, process water, reclaim, and wastewater systems. This role partners with Operations, Engineering, Maintenance, Construction, Procurement, EHS, vendors, and leadership to plan, implement, document, and align work with production needs. Success requires strong organization, technical understanding, communication, attention to detail, and the ability to lead priorities across operations, maintenance, projects, and production schedules. Ideal candidates are proficient with SAP or other CMMS tools, Microsoft Office, trackers, dashboards, drawings, and documentation systems. They coordinate schedules, track action items, communicate status, support scope development, identify gaps, and improve safety, reliability, documentation, cost control, and execution quality. This role helps maintain critical facility systems while demonstrating Micron’s core values of People, Innovation, Tenacity, Collaboration, and Customer Focus. Responsibilities: Coo

aiSAPprocurement
View job →

Job Title Senior Software Technologist I - C++ Job Description Software Engineer II Your Role: • Design, develop, test, and maintain software components and applications using modern C++, C# in a Windows-based environment. • Participate in the full software development lifecycle including requirements analysis, design, implementation, testing, debugging, and maintenance. • Develop and maintain software applications using Visual Studio and associated C++, C# development tools. • Work with Windows operating system fundamentals including processes, services, registry, file system, User Account Control (UAC), and application configuration. • Create, enhance, and troubleshoot software modules while adhering to coding standards, design guidelines, and software development best practices. • Utilize GitHub for source control management including branching strategies, commits, pull requests, merges, rebasing, and code reviews. • Support and troubleshoot CI/CD pipeline issues using GitHub Actions and participate in continuous integration activities. • Manage software dependencies and package management using NuGet and Conan. • Configure and maintain build systems using CMake and Visual Studio project configurations. • Develop and execute unit tests using Google Test (GTest) to ensure software quality and reliability. • Perform debugging, root cause analysis, and defect resolution for software issues identified during development, testing, and field support activities. • Participate in peer code reviews and contribute to software quality, maintainability, and technical excellence. • Collaborate effectively with Software Verification, Product Management, DevOps, Architecture, and cross-functional teams to deliver high-quality software solutions. • Create and maintain

🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Top cities for Reliability Engineer

City links are canonicalized and require at least 20 current jobs.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.