Jobiba hiring network

Software Reliability Engineer Jobs

6,428 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

E
10 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. The Role As part of the Core Products Business Unit (CPBU) Systems Software team, you will join a high-leverage engineering initiative dedicated to transforming upgrade workflows and lifecycle operations. Our goal is to shift hardware upgrades from manual, support-intensive workflows toward a safer, guided, automated self-service experience integrated directly into Pure1 and Skyline. In this role as a Systems Software Engineer (MTS3), located on-site in Santa Clara, CA, you will build dedicated systems infrastructure and controller-upgrade automation for our next-generation upgrade framework. You will bridge on-array software execution with cloud-guided operations, eliminating manual choreography and runbooks while maintaining uncompromising standards for safety and availability. What You'll Do Design & Build Upgrade Infrastructure: Architect, implement, and maintain high-reliability systems software components for controller-upgrade automation and the next-generation hardware upgrade framework. Automate Upgrade Workflows: Transition complex hardware replacement and upgrade tasks into repeatable, self-service workflows, significantly reducing support and field dependencies. Engineered Safety & Resilience: Build robust state machines, health checks, workflow orchestration, pre/post-upgrade validations, failure-handling mechanisms, and automated recovery paths. Cross-Functional Collaboration: Partner closely with int

pythonaic++
View job →
H
Hp
📍 Texas• $123.1K – $150.4K/yr
1mo ago

Software Product Security Engineer Description - This role supports the development and maintenance of secure software products under the guidance of senior engineers. The position focuses on learning software engineering and security best practices while contributing to the design, implementation, testing, and maintenance of desktop, web, and cloud-based applications and services. Key Responsibilities Assist in developing, testing, and maintaining software applications and security solutions. Participate in software development activities including coding, debugging, testing, and integration. Support the development and maintenance of Windows desktop applications and services. Assist in developing and maintaining web applications, APIs, and cloud-connected services. Troubleshoot software issues with guidance from senior team members. Write clean, maintainable, and well-documented code. Create and execute unit tests to verify software functionality and reliability. Participate in code reviews and learn software development best practices. Contribute to Agile ceremonies, sprint planning, and team activities. Learn and apply secure coding and software security principles. Support the deployment, monitoring, and maintenance of cloud-based applications and services. Collaborate with cross-functional teams to deliver end-to-end software solutions. Support product release and maintenance activities. Education & Experience Bachelor's or Master's Degree in Computer Science, Software Engineering, or a related discipline. 0-2 years of software development experience. <li

javascripttypescriptpython
View job →
N
1mo ago

We are looking for a highly motivated AI/ML Software Engineer to join the Enterprise Agentic AI Platform team within IT. You will work closely with Business Analysts, and Engineering teams to design, develop, and deploy enterprise AI solutions that improve productivity and automate business workflows across Engineering, Operations, and Manufacturing. What you'll be doing: Design, develop, and deploy Agentic AI applications using Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and AI orchestration frameworks. Build scalable AI services and reusable components integrated with enterprise applications such as PLM, SAP, and other business systems. Collaborate with business and IT teams to translate business requirements into AI-driven solutions. Develop secure, scalable APIs and enterprise integrations to enable intelligent workflows and automation. Improve AI solution quality, performance, and reliability through prompt engineering, evaluation, and continuous optimization. Partner with cross-functional teams throughout the Software Development Lifecycle (SDLC), from solution design through deployment and production support. What we need to see: Bachelor's or Master's degree in Computer Science, Information Technology, AI/ML, or a related field. 6&#43; years of software engineering experience with strong proficiency in Python and backend application development. Hands-on experience with Generative AI, LLMs, RAG, AI agents, REST APIs, and cloud-native application development. Experience integrating enterprise applications and building scalable, production-ready software solutions. Strong analytical, problem-solving, communicatio

pythonazureai
View job →

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Infrastructure Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for Palantir platforms, products, and deployments. You'll get to use your creativity to develop novel solutions to evolving challenges and automate processes wherever possible, using whichever tools are best for the job including industry-leading LLM and AI technology! As a Forward Deployed Infrastructure Engineer, every day is different! You will be developing software and owning reliability and operations for software systems that are critical to solving our government’s greatest challenges. We strongly believe in engineering teams being responsible for the operations of their services in production. As such, you’ll work closely with both forward deployed teams and product teams to participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Infrastructure Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for Palantir platforms, products, and deployments. You'll get to use your creativity to develop novel solutions to evolving challenges and automate processes wherever possible, using whichever tools are best for the job including industry-leading LLM and AI technology! As a Forward Deployed Infrastructure Engineer, every day is different! You will be developing software and owning reliability and operations for software systems that are critical to solving our government’s greatest challenges. We strongly believe in engineering teams being responsible for the operations of their services in production. As such, you’ll work closely with both forward deployed teams and product teams to participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Infrastructure Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for Palantir platforms, products, and deployments. You'll get to use your creativity to develop novel solutions to evolving challenges and automate processes wherever possible, using whichever tools are best for the job including industry-leading LLM and AI technology! As a Forward Deployed Infrastructure Engineer, every day is different! You will be developing software and owning reliability and operations for software systems that are critical to solving our government’s greatest challenges. We strongly believe in engineering teams being responsible for the operations of their services in production. As such, you’ll work closely with both forward deployed teams and product teams to participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.

PE
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Infrastructure Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for Palantir platforms, products, and deployments. You'll get to use your creativity to develop novel solutions to evolving challenges and automate processes wherever possible, using whichever tools are best for the job including industry-leading LLM and AI technology! As a Forward Deployed Infrastructure Engineer, every day is different! You will be developing software and owning reliability and operations for software systems that are critical to solving our government’s greatest challenges. We strongly believe in engineering teams being responsible for the operations of their services in production. As such, you’ll work closely with both forward deployed teams and product teams to participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.

Job Title AI Senior Systems Engineer (AI for RAMS & Systems Engineering) Job Description As an AI Senior Systems Engineer, you will help shape the future of Systems Engineering at Philips by driving the application of Artificial Intelligence within Systems Engineering and Reliability, Availability, Maintainability and Safety Engineering (RAMS) practices. You will identify opportunities where AI can enhance engineering activities, improve engineering productivity, and increase the quality, consistency, and traceability of engineering deliverables throughout the product development lifecycle. Acting as a thought leader and trusted advisor, you will help engineering teams understand, adopt, and effectively apply AI technologies within their daily engineering work. Working closely with systems engineers, RAMS engineers, architects, software teams, and AI specialists, you will bridge the gap between engineering challenges and emerging AI capabilities. You will contribute to the evolution of engineering methodologies, tools, and best practices that enable the next generation of AI-enabled Systems Engineering and RAMS Engineering at Philips. Your role: Drive the adoption of AI within Systems Engineering and RAMS practices across Philips. Identify, develop, and scale AI use cases for requirements engineering, system architecture, modelling, verification, validation, traceability, and engineering knowledge management. Identify, develop, and scale AI use cases for RAMS engineering like FMEA, data analysis, HALT/ALT, Modelling Simulation & Analysis, for Hardware and Software Support engineering teams in evaluating and implementing AI-enabled engineering workflows, methods, and tools. Coach and educate Systems Engineers, RAMS Engineers, Architects, and technical leaders on the opportunities, limitatio

machine learningartificial intelligenceai
View job →

Job Title Senior Software Technologist I - C++ Job Description Software Engineer II Your Role: • Design, develop, test, and maintain software components and applications using modern C&#43;&#43;, C# in a Windows-based environment. • Participate in the full software development lifecycle including requirements analysis, design, implementation, testing, debugging, and maintenance. • Develop and maintain software applications using Visual Studio and associated C&#43;&#43;, C# development tools. • Work with Windows operating system fundamentals including processes, services, registry, file system, User Account Control (UAC), and application configuration. • Create, enhance, and troubleshoot software modules while adhering to coding standards, design guidelines, and software development best practices. • Utilize GitHub for source control management including branching strategies, commits, pull requests, merges, rebasing, and code reviews. • Support and troubleshoot CI/CD pipeline issues using GitHub Actions and participate in continuous integration activities. • Manage software dependencies and package management using NuGet and Conan. • Configure and maintain build systems using CMake and Visual Studio project configurations. • Develop and execute unit tests using Google Test (GTest) to ensure software quality and reliability. • Perform debugging, root cause analysis, and defect resolution for software issues identified during development, testing, and field support activities. • Participate in peer code reviews and contribute to software quality, maintainability, and technical excellence. • Collaborate effectively with Software Verification, Product Management, DevOps, Architecture, and cross-functional teams to deliver high-quality software solutions. • Create and maintain

SC
Sigma Computing
📍 San Francisco• Full-time• $170K – $240K/yr
16 days ago

Senior Software Engineer - Observability and Reliability About the Role We are growing the engineering team and looking for engineers who have the chops to build and deliver world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible. What You Will Be Doing Build observability tools and platforms, including: metrics, logging, distributed tracing, dashboarding, alerting, application performance management Build with modern tools and languages like Go, Open Telemetry and Kubernetes Participate in on-call rotation and ensure uptime of services Create runtime tools/processes that optimize cloud triaging and limit downtime Define best practices around making our systems and services measurable Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies. We expect successful candidates to be coding a majority of their time Qualifications We Need Strong Computer Science fundamentals 5+ years industry experience building and maintaining high-quality software, especially software other engineers use You apply a product mindset to infrastructure systems and feel accomplished enabling others Desire to be a great teammate and have fun at work Strong sense of craftsmanship, and a healthy academic curiosity Qualifications We Want (also, skills you’ll learn!) Experience building systems for data analytics Distributed systems monitoring and profiling skills Knowledge of cloud application security models Administered cloud service infrastructure (GCP, AWS, Azure) Startup experience Additional Job details Additional Job details The base salary range for this position is $170k - $240k annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience. Base pay is one part of the Total Package that is provided to compensate and recognize e

pythonsqlaws
View job →
O
1mo ago

Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role As OpenAI continues to grow, we are looking for experienced, problem-solving engineers to ensure our systems scale. Our success depends on our ability to quickly iterate on products while also ensuring that they are performant and reliable. You will work in a deeply iterative, collaborative, fast-paced environment to bring our technology to millions of users around the world, and ensure it’s delivered with safety and reliability in mind. Successful candidates will play a crucial role in ensuring the reliability, scalability, and performance of our systems as we continue to expand. As a reliability expert, you will be at the forefront of maintaining and enhancing the stability, scalability, and performance of our rapidly evolving infrastructure. You will work closely with cross-functional teams, including software engineers, product managers, and data scientists, to build and maintain resilient systems that can handle our growing user base and workload. In this role, you will: Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands. Build and maintain the load, chaos and synthetic testing software leveraged by development teams to make the systems they design and operate more reliable. Build and maintain automation tools to streamline repetitive tasks and improve system reliability. Build and maintain the platform for CPU/storage, GPU, and network lifecycle management to drive efficiency, accountability and support dynamic optimization of our resources. Implement fault-tolerant and resilient

awskubernetesrest
View job →
C
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations Austin, US Responsibilities The Argo team was formed to own a very important aspect of Cloudflare's systems: enable more reliable network connectivity for Cloudflare’s products than the Internet itself provides. Almost all products in Cloudflare’s portfolio are or will be powered by Argo technology, including CDN, Spectrum, Magic Transit, Stream, Workers, Workers AI, R2, WARP, and more. As a member of the Argo team, you’ll be a

sqlpostgresqlaws
View job →

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX What You’ll Do The Argo team was formed to own a very important aspect of Cloudflare's systems: enable more reliable network connectivity for Cloudflare’s products than the Internet itself provides. Almost all products in Cloudflare’s portfolio are or will be powered by Argo technology, including CDN, Spectrum, Magic Transit, Stream, Workers, Workers AI, R2, WARP, and more. As a member of the Argo team, you’ll

sqlpostgresqlaws
View job →
A
Asana
📍 Warsaw• Full-time• $372K – $432K/yr
1mo ago

Asana’s rapid growth brings new challenges in keeping our systems fast, reliable, and resilient. As our product evolves, we’re making a major investment in reliability – and building a brand new SRE team in Warsaw is a key part of that strategy. This is your chance to help shape it from day one. This isn’t a traditional “ops” role – we’re looking for strong software engineers who are passionate about building reliable, distributed systems. You’ll work closely with a small SRE team in San Francisco, infrastructure engineers in Reykjavik, and an established infrastructure team in Warsaw. Warsaw will be a significant hub for our future infrastructure engineering and operations. As one of the first engineers here, you’ll have a real say in how we build reliable infrastructure, manage incidents, and support the rest of the company. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you’ll achieve Influence the future of Asana’s SRE practice, especially as we grow the Warsaw team. Lead reliability-focused projects across our stack – from infrastructure to tooling to incident response. Define and implement Asana’s incident management process – we’re investing here, and you’ll help shape how it works. Build internal platforms and frameworks that help other teams improve the reliability of their services. Be part of (and help shape) a sustainable on-call rotation – shared across teams in Warsaw, San Francisco, and Reykjavik. On average, we handle ~1 page per day, but it’s not constant, and we care about keeping things sane. Work with our stack: AWS, Kubernetes (EKS), Datadog, MySQL (RDS), ElasticSearch (OpenSearch), Redis

typescriptpythonsql
View job →
H
Hp
📍 Colorado• $59.4K – $89.6K/yr
8 days ago

Software Quality Engineer Description - This role is responsible for maintaining the quality, reliability, and performance of software applications throughout the development lifecycle. The role identifies and rectifies defects, ensures adherence to established quality standards, and contributes to the overall improvement of the software development process. The role involves various activities aimed at preventing and detecting issues, thereby enhancing the end user experience. The role creates and executes comprehensive test plans, test cases, and test scripts based on project specifications. *Onsite in Ft. Collins 5-days a week Responsibilities • Executes established test plans and protocols for assigned portions of code for end-user applications, systems software, and firmware running on hardware, local, networked, and Internet- based platforms; identifies, logs, and debugs assigned issues. • Perform Functional and Solution Testing of Video/Collaboration Software • Additionally, codes and programs test scripts, automation, and integration activities based on specific test requirements. • Conducts functional, integration, regression, and performance testing to validate software functionality. • Automates testing processes using appropriate tools and frameworks to improve efficiency and repeatability. • Monitors and enforces adherence to established coding standards, design guidelines, and best practices. • Monitors software performance and conducts load and stress testing to identify bottlenecks and performance issues. • Prepares and maintains QA-related documentation, including test plans, test matrices, and testing reports. • Develops understanding of and relationship with internal and outsourced development partners on software applications design and development. • Participates as a member of project

pythonaijenkins
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime