Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations organization is currently seeking a Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking an experienced cloud infrastructure professional to help design, build, and sustain the foundational cloud environment supporting critical program needs. In this role, the selected candidate will help establish and operate secure, scalable, and resilient cloud infrastructure environments in Microsoft Azure to enable enterprise applications, software toolchains, and digital engineering workloads. As both an individual contributor and technical leader, this position will work across network, computer, storage, identity, security, and automation domains to deliver repeatable cloud infrastructure patterns and operational excellence. This role is focused on infrastructure operations, sustainment, automation, and reliability, rather than application software development. Position Responsibilities: Design, implement, and maintain Microsoft Azure-based infrastructure solutions including networking, compute, storage, identity integration, and supporting services Develop and maintain Infrastructure as Code (IaC) and configuration automation solutions using Terraform, Ansible, PowerShell, and Bash Implement cloud policies to enforce security, ensure regulatory compliance, and manage user access Build repeatable landing zones and cloud infrastructure patterns that support mul
Jobiba hiring network
Software Reliability Engineer Jobs
6,326 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . The Developer Infrastructure team within Coinbase's Platform product group builds the tools and platforms that every Coinbase engineer relies on to build, test, ship, and operate software. As the Group Product Manager for this team, you'll own the product vision and multi-year strategy for the entire code lifecycle, from code change to production deployment. You'll define what world-class developer infrastructure looks like at Coinbase and drive the execution that makes every engineer faster, safer, and more productive. What you’ll do: Own the end-to-end product roadmap and success metrics for Developer Infrastructure, including CI/CD pipelines, release automation, testing infrastructure, deployment systems, and production readiness standards. Drive a continuous simplification agenda by leading regular systems migrations that reduce complexity, improve reliability, and make the platform easier for engineers to adopt and use effectively. Build quality scorecards and leaderboards for production systems across Coinbase, ensuring all systems meet defined standards with executive-level visibility and accountability. Partner with Engineering, SRE, Security, and infrastructure leadership to ensure developer lifecycle systems are reliable, performant, secure, and scalable. Deepen understanding of internal customer needs by analyzing usage patterns, gathering feedback f
About the Team The People Technology team builds and operates the systems that support how OpenAI hires, develops, and supports its people. The team brings together People Systems and People Innovation Labs, a product engineering group focused on rethinking how we find and retain exceptional talent and help employees do their best work. People Systems owns the company’s core people-technology ecosystem, including platforms such as Workday and Ashby. People Innovation Labs builds new employee and People Team experiences on top of that foundation, including OpenHouse, our internal employee hub, and AI-powered products and automations. Together, we are working toward a model in which our enterprise systems provide reliable data, controls, and core business logic, while employees and managers can complete more of their work through simple, integrated, and AI-native experiences. About the Role We are looking for a People Systems Lead to manage the People Systems team and shape how our core systems evolve. You will be responsible for the reliability and effectiveness of our current environment while helping us move beyond the constraints of traditional enterprise software. This includes stabilizing and improving platforms such as Workday and Ashby, designing the integrations that connect them to the broader technology ecosystem, and partnering with People Innovation Labs to surface workflows through OpenHouse, Slack, and AI-powered experiences. This role requires someone who is comfortable moving between strategy, technical design, and team leadership. You should understand People systems deeply, be able to work through integration and architecture decisions with engineers, and translate complex organizational needs into scalable solutions. You will also manage vendor relationships, develop the People Systems team, and drive alignment across People, Engineering, Finance, Security, Legal, and other partners. This role could be a fit for someone who has grown up in People S
AI/ML – Investment Services A Career with Point72's AI/ML – Investment Services Team The AI/ML – Investment Services team at Point72 spearheads the development of cutting-edge AI solutions that seek to transform our business processes and enhance enterprise intelligence. The team aims to bridge the gap between business challenges and technological innovation, collaborating with stakeholders across the firm and leveraging expertise in generative AI, data engineering, and machine learning. WHAT YOU'LL DO Build and scale core backend services and platforms that power generative AI applications and data infrastructure used across the firm’s investment workflows Design and implement high-throughput, low-latency data pipelines to ingest, normalize, and serve both structured and unstructured data Develop robust APIs and microservices to support model inference, feature serving, and downstream applications Integrate generative AI tools and model-serving workflows into production, including embedding stores, retrieval components, and fine-tuning pipelines Optimize system performance, cost, and reliability through profiling, capacity planning, and architectural improvements Implement automated testing, continuous delivery pipelines, monitoring, and incident response practices to maintain production health Partner with data scientists, AI engineers, product owners, and operations to translate models and prototypes into scalable, production-grade solutions Mentor engineers, lead code reviews, and establish engineering best practices for maintainability, security, and observability Own end-to-end delivery, operational runbooks, and metrics-driven measurement of feature impact and system reliability WHAT'S REQUIRED Bachelor’s degree in computer science, software engineering, or a related technical field Minimum 5+ years of professional experience building backend systems and production services Demonstrated experience designing and operating large-scale data engineering pipelines
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe Terminal helps businesses extend their online presence into the physical world. Our mission is to make it as easy for businesses to accept in-person payments as Stripe has made it to accept payments online. Terminal provides an integrated in-person payments platform spanning software, payment processing, device fleet management, and hardware. Our hardware portfolio includes Stripe-designed readers and accessories as well as devices from third-party hardware partners. Together, these products enable businesses to build reliable, differentiated in-person payment experiences across a broad range of use cases, countries, and operating environments. What you’ll do We’re looking for an experienced Product Manager to build the strategy, development, and expansion of Terminal’s hardware portfolio. This role will work across both first-party and third-party hardware, helping define the multi-year roadmap for the devices, accessories, and partner ecosystem that power Terminal. You will partner closely with hardware engineering, firmware and software engineering, design, operations, logistics, partnerships, sales, finance, legal, and regional teams. This is a highly cross-functional role for someone who can connect customer needs, technical constraints, business strategy, and execution across user experience, device capabilities, cost, quality, reliability, supply, and time to market. Responsibilities Set and execute a multi-year strategy and ro
We are seeking an experienced IT/Lab Manager to lead the planning, deployment, and operations of our physical lab environment and IT systems. This role will focus on building and maintaining scalable, reliable, and secure environments to support engineering teams involved in research, quality assurance, validation, and related activities. It will also support internal collaborators. You will have an outstanding opportunity to drive innovation in a multidimensional, technology-focused company that is crafting the future of data-center and lab technologies. If you bring perfection and creative thinking while solving issues as they arise, and enjoy working with distributed teams – your place is with us! What You’ll Be Doing: Own day-to-day operations, planning, and roadmap for the engineering lab and IT infrastructure (servers, storage, networking, and related services). Lead and mentor an IT/Lab team, driving guidelines, standards, and a culture of ownership, partnership, and continuous improvement. Collaborate closely with R&D, QE, Verification, and other engineering teams to design, provision, and maintain environments that meet their performance, reliability, and security needs. Lead all aspects of running data center and lab operations, including rack layout, cabling, power and cooling, hardware lifecycle, and resource availability. Lead procurement and vendor management for hardware, software, and services, including evaluation, negotiation, and ongoing relationship management. Implement and maintain automation for system provisioning, configuration, and operations using tools such as shell/Perl/Ansible. Design and maintain monitoring, logging, and alerting for servers, network, and storage systems to ensure high availability and rapid incident response. Investigate and resolve sophisticated infrastructure issues across OS, networking, storage, virtualization, and appli
About BlockTech BlockTech is an algorithmic trading firm operating at the frontier of global crypto derivatives and spot markets. We trade 24/7 across some of the fastest-moving, most data-rich venues in finance. Crypto is one of the few markets where the gap between a good model and live PnL comes down to how fast and how reliably you can ship it. Abundant data, novel microstructure, and a short path to production mean the quality of our infrastructure directly moves the edge. We're looking for a Quantitative Developer to build that infrastructure and get our models into production. The role This is a hands-on engineering seat on our trading floor. You'll own the infrastructure that turns research models into production trading systems, working primarily in Python and Rust, the language our models and systems are written in. You'll build the frameworks and tooling that let quants ship their code to production and keep those systems running once they're live. You'll work shoulder-to-shoulder with researchers, traders, and fellow engineers, building the tooling, pipelines, and standards that let good ideas reach production quickly and safely, and helping raise the engineering bar across the floor. You will Build and own the infrastructure that takes models from research to production - data pipelines, backtesting, deployment, and monitoring. Turn research prototypes into robust, performant, production-grade code that runs reliably in live trading Work closely with quantitative researchers and traders to get new models live quickly and safely Own the systems you ship in production: monitor performance, diagnose issues, and keep latency and reliability high Improve the tools, standards, and pipelines the team relies on, raising the engineering bar across the floor What we're looking for 2+ years of software engineering experience, with a track record of shipping production-grade systems A strong academic foundation in a STEM discipline (computer science, mathematics, p
Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. We are an extremely product focussed team. We work in partnership with Product and Design functions of teams we support. Our team's dedicated ML product engineers enable us to move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting-edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure the real customer impact of each model we deploy. What will I be doing? Play an active role in hiring, mentoring and career development of other engineers Raise the bar for technical standards, performance, reliability, and operational excellence Identify areas
Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. We are an extremely product focussed team. We work in partnership with Product and Design functions of teams we support. Our team's dedicated ML product engineers enable us to move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting-edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure the real customer impact of each model we deploy. What will I be doing? Play an active role in hiring, mentoring and career development of other engineers Raise the bar for technical standards, performance, reliability, and operational excellence Identify areas
About Target As a Fortune 50 company with more than 400,000 team members worldwide, Target is an iconic brand and one of America's leading retailers. Joining Target means promoting a culture of mutual care and respect and striving to make the most meaningful and positive impact. At Target, we have a timeless purpose and a proven strategy. Some of the best minds from different backgrounds come together to redefine retail in an inclusive learning environment that values people and delivers world-class outcomes. Target in India operates as a fully integrated part of Target's global team and supports the company's global strategy and operations. About the team The IT Data Platform (ITDP) team enables data-driven management of Target's technology ecosystem by bringing together trusted data and insights across technology assets, software delivery, infrastructure, reliability, security, engineering effectiveness, and technology operations. The Analytics team within ITDP transforms this data into metrics, analytical products, dashboards, predictive insights, and decision-support capabilities that help technology teams understand what is happening, why it is happening, where risk may be emerging, and where action is needed. The team is building toward an analytics capability that progresses from descriptive and diagnostic analytics to predictive and prescriptive insights, using statistical methods, applied data science, and GenAI where each approach is appropriate. As a Senior Data Analyst for Target's IT Data Platform Analytics team you'll: <p style="co
About us: Working at Target means helping all families discover the joy of everyday life. About Target As a Fortune 50 company with more than 400,000 team members worldwide, Target is an iconic brand and one of America's leading retailers. Joining Target means promoting a culture of mutual care and respect and striving to make the most meaningful and positive impact. At Target, we have a timeless purpose and a proven strategy. Some of the best minds from different backgrounds come together to redefine retail in an inclusive learning environment that values people and delivers world-class outcomes. Target in India operates as a fully integrated part of Target's global team and supports the company's global strategy and operations. About the team The IT Data Platform (ITDP) team enables data-driven management of Target's technology ecosystem by bringing together trusted data and insights across technology assets, software delivery, infrastructure, reliability, security, engineering effectiveness, and technology operations. The Analytics team within ITDP transforms this data into metrics, analytical products, dashboards, predictive insights, and decision-support capabilities that help technology teams understand what is happening, why it is happening, where risk may be emerging, and where action is needed. The team is building towards an analytics capability that progresses from descriptive and diagnostic analytics to predictive and prescriptive insights, using statistical methods, applied data science, and GenAI where each approach is appropriate. <p style="color:!im
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Senior Staff Rust Developer to join our Platform Convergence Team. This is a hybrid role based in San Jose, CA reporting to the Sr. Director, Software Engineering. Join us to build a new platform from the ground up that can scale hundreds of millions of users with high reliability and low latency. You will design and implement distributed system and core infrastructure components while collaborating closely with various stakeholders. What you’ll do (Role Expectations) Design and build a low-latency, high-throughput data forwarding plane using Rust, leveraging its async/await model for efficient I/O and service-oriented infrastructure Develop distributed, scalable systems with a focus on concurrency, fault tolerance, and messaging Implement and maintain gRPC-based APIs and services to integrate forwarding plane capabilities with control and orchestration layers Optimize system
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. You Will: Design, build, and maintain backend integrations using AWS services such as Lambda, S3, SQS, SNS, and EventBridge Develop API-based automations and orchestrate data flows between Salesforce, NetSuite, Workday, Snowflake, and other SaaS platforms Refactor and enhance legacy integration logic to align with modern architecture standards Own your code from design through CI/CD (GitLab) and deployment via Terraform/CloudFormation Collaborate with US-based engineers on shared initiatives, sprint planning, code reviews, and technical handoffs Participate in a shared production support rotation, resolving critical integration and data pipeline issues Partner with architects and BSAs to review integration design and testing strategies for reliability and resilience Support both ETL-style batch pipelines and real-time event streaming using CDC and platform events You Have: BS in Computer Science, related field, or equivalent industry experience 5+ years of experience in backend software development or automation using Node.js, Python, VS code or similar Strong experience with AWS services, especially Lambda, S3, SQS/SNS, EventBridge, DynamoDB Expertise in database design, schema modeling, and query optimization (DynamoDB & Snowflake preferred) Strong grasp of integration design patterns, secure API development (REST/SOAP), and system interoperability Develop and maintain GitLab based CI/CD pipeline implementations for tests, linting, deployment, etc. Experience with managing infrastructure as code (Terraform, Cl
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role The AI team is building cutting-edge solutions that bring the power of AI directly to edge devices while seamlessly integrating with cloud infrastructure. We are looking for a Senior Software Engineer to design and develop high-performance, scalable services to support AI workloads across edge and cloud environments. What You Might Do Design, build, and maintain services that power AI-driven applications, ensuring scalability and performance. Develop APIs and microservices that facilitate seamless integration between cloud-based AI models and edge devices. Optimize data pipelines and storage solutions for real-time AI inference and processing. Implement security and privacy best practices for distributed AI systems. Work closely with AI researchers, infrastructure engineers, and frontend developers to deliver end-to-end AI-driven solutions. Build and optimize an agent orchestration runtime that enables tool use, memory management, and multi-step reasoning across LLMs, APIs, and edge-connected systems. Develop robust logging, monitoring, and alerting systems to ensure system reliabilit
Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. The Developer Experience (DX) team at Amplitude builds and maintains the foundations that power how developers integrate, extend, and trust Amplitude across platforms. Our mission is to make Amplitude’s SDKs reliable, easy to adopt, and a joy to build on, so customers can confidently instrument their products and unlock insights at scale. We’re looking for a Staff Software Engineer, iOS to play a key technical leadership role on our DevEx team. In this role, you will lead the design and development of Amplitude’s core iOS SDKs, including Analytics and Session Replay , and serve as the iOS platform expert that other SDK teams, such as Statsig, Guides, and Surveys rely on. As a Staff Engineer, you’ll operate with a wide scope and high impact: setting technical direction for the iOS platform, driving cross-SDK architecture, improving performance and reliabilit
Get new software reliability engineer jobs by email
Daily job updates · Unsubscribe anytime