Jobiba hiring network

Engineering Excellence Engineer 2 Jobs

8,135 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current engineering excellence engineer 2 jobs. Use filters to narrow by work mode, employment type, experience and date posted.

A
Affirm
📍 Spain• Full-time• Remote• From €1M/yr
19 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions

REMOTEpythonsqlmysql
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng

pythonsqlpostgresql
View job →
O
Okta
📍 San Francisco• Full-time• From $165K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in a global on-call rotation supporting highly available customer-facing systems. Participate in incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with en

pythonsqlpostgresql
View job →

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<

pythonawsazure
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron’s DRAM Design Engineering Group (DDEG) is where innovation meets excellence. We are advancing memory and storage technologies through collaborative engineering and creative problem-solving. Our team works at the forefront of semiconductor design, developing solutions that shape the future of memory products in a fast-paced, learning-focused environment. As a Design Verification Engineer, you will help develop next-generation memory technologies by verifying and optimizing digital and analog circuit designs. In this role, you will work closely with global multi-functional teams across the product lifecycle to deliver high-quality, manufacturable memory solutions that meet performance, reliability, cost, and customer requirements. Your work will directly contribute to bringing advanced memory products from concept to production. Responsibilities: Verify circuit functionality, reliability, power, and compliance with product specifications Drive verification planning, coverage closure, circuit debug, and design improvements Perform circuit modeling and simulation using industry-standard tools Support silicon validation, reticle experiments, and tape-out activities Partner with engineering, manufacturing, and product teams to deliver manufacturable designs Minimum Qualifications: Bachelor’s degree in Electrical Engineering or a related field 4&#43; years of semiconductor design, verifi

pythonairecruitment
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. The Staff SRE, Classified Opportunity This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a technical leader within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. Security Clearance: Active U.S. TS/SCI clearance with Full Scope Poly Compliance Expertise: Proven experience navigating Federal and DoD compliance frameworks, specifically FedRAMP and Impact Level 6 (IL6) What You’ll Do Work with various teams to design and implement scalable, and reliable network solutions Maintain a highly available cloud infrastructure edge for the Okta identity platform C

pythonawsdocker
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. The Firmware & Product Test (FPT) team plays a critical role in delivering high-quality enterprise SSD solutions by ensuring firmware functionality, reliability, and compliance. We work across simulation, FPGA, and hardware environments to validate modern storage technologies, build scalable automation, and drive continuous improvement in validation methodologies. Our team values technical excellence, collaboration, and innovation, including the use of AI-enabled tools to enhance engineering productivity and quality. As a Principal Test Development Engineer, you will serve as a technical leader for firmware validation, defining verification strategies, advancing automation frameworks, and driving complex failure analysis efforts. This role offers the opportunity to influence product quality across multiple SSD programs while mentoring engineers and partnering closely with firmware architects to improve testability and validation effectiveness. Responsibilities: Lead verification strategy, test planning, automation, and coverage closure for NVMe front-end firmware features across multiple product lines Architect and enhance scalable Python-based test automation frameworks, CI/CD integration, regression infrastructure, and reporting capabilities Drive root-cause analysis and failure triage using firmware traces, protocol analyzers, system logs, and structured debug methodologies Define validation standards, review test code, mentor engineers, and promote standard methodologies in automation and qua

pythongitlinux
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Site Reliability Engineer (SRE) - Security and Data Systems Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering and systems administration, where you'll be responsible for building and maintaining highly reliable, scalable, and secure infrastructure. You will be a key contributor, applying your expertise to automate manual processes and proactively solve complex problems before they become incidents, handling incidents, and includes on-call shifts. Responsibilities Platform & Reliability: Design, build, and maintain the core infrastructure that underpins our security SaaS offerings, ensuring high availability, performance, and scalability. This includes building and operating the tooling for our Snowflake data systems. Automation: Develop robust automation using code to eliminate toil and ensure consistency across our environments. You'll be a key driver in automating everything from infrastructure provisioning to application deployment and incident response. Security & Compliance: Work closely with our security teams to embed a security-first mindset into all our processes and infrastructure. You will be responsible for ensuring our systems and data platforms are compliant with industry standards. Incident Response: Participate in on-call rotations and be a primary responder for critical inci

kubernetesmachine learningartificial intelligence
View job →
M
Mongodb
📍 New York• Full-time• From $126K/yr
1mo ago

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. We are looking for a talented Senior Software Engineer to join our team as we execute on a multi-year roadmap and prepare to launch and rapidly scale services that handle petabytes of data. Come do some of the most interesting work of your career as we Think Big and Go Far for our customers! This role can be based out of our New York City office (hybrid working model) or remotely in the North America region. What you’ll do Design, build, and operate control plane services powering an elastic and multi-tenant storage layer for thousands of database instances. Solve problems around maintaining high availability and performance during load spikes, hardware failures, cloud provider outages, and other disruptions. Contribute to a culture of operational excellence through dashboards, playbooks, and on-call improvements. Lead complex technical projects from planning through successful deployment with clear stakeholder updates. Partner closely with peers across database, cloud, and infrastructure engineering teams as well as project management to investigate incidents and develop long-term roadmaps. Mentor junior engineers and foster a collaborative team environment. We’re looking for someone with 5+ years of professional software development experience building, deploying, and operating multi-tenant cloud services with a focus on operational excellence. Experience with large backend/compiled codebases, such as Rust or C/C++. Experience with containerization and orchestration platforms (e.g. Kubernetes). Experience with observability tooling (e.g. time series metrics, dashboards). Experience with distributed systems

mongodbawsazure
View job →
O
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Auth0 provides an unparalleled authentication experience for hundreds of millions of users worldwide. Our commitment to reliability is a key foundation of our product and our dedication to exceeding customer availability expectations is a core engineering focus. As a Senior Site Reliability Engineer, you'll join our SRE team based in Europe to ensure our production systems are not only operational but also resilient, scalable, and ready for exponential growth. This isn't just about keeping the lights on; it's about directly contributing to the platform's core resiliency and robustness. You'll be a hands-on builder, crafting solutions that make our system more reliable by design. What you’ll do: Design and build custom software in Go to enhance the platform's reliability, resiliency, and redundancy. Partner with engineering teams to embed reliability principles, improving the availability, performance, and observability of our services. Use your deep understanding of infrastructure and observability principles to identify opportunities for improvement within the product and implement solutions. Contribute to our follow-the-sun on-call rotation, providing rapid, effective response to critical incidents and using your expertise to troubleshoot, mitigate or accurately escalate production issues. Because our team is globally distributed, your on-call shifts will only occur during your standard local working hours. Develop and refine our SRE tooling and proc

sqlawsazure
View job →
O
Okta
📍 Toronto• Full-time• From C$110K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Team The Auth0 Platform Tools team owns the incident management tooling, Slack-based tooling, StatusPage, and local development environments that Auth0 engineers rely on every day. That includes incident.io and the services we have built around it, Statuspage, custom Slack bot applications that automate our incident response and engineering operations workflows, the customer-facing web application behind status.auth0.com, Vivaldi, and Tilt - the tools engineers use to run Auth0 locally. We are seeking an engineer to help build new features across all of these tools. Our stack is primarily TypeScript and Node.js, with a React and Next.js front end, backed by Postgres and Redis, and deployed on Kubernetes on AWS. A significant portion of our incident and engineering operations automation is built on Tines, a no-code automation platform. Prior no-code experience is welcome, but we expect you to learn Tines here and become effective with it. Current initiatives include extending our incident tooling to meet FedRAMP requirements, taking full ownership of the status page, and improving how we communicate incident status to customers. There is real room to improve along the way, from test coverage to resilience to inherited technical debt. We build for two audiences: Auth0 engineers, who depend on our tooling every day, and Auth0's customers, who rely on the status page during incidents. We are looking for an engineer who cares about both and enjoys working wi

typescriptpythonreact
View job →
O
Okta
📍 Washington• Full-time• From $174K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. The Staff SRE, Classified Opportunity This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a technical leader within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. Security Clearance: Active U.S. TS/SCI clearance with Full Scope Poly Compliance Expertise: Proven experience navigating Federal and DoD compliance frameworks, specifically FedRAMP and Impact Level 6 (IL6) What you’ll be doing: Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting hi

pythonsqlpostgresql
View job →
O
Okta
📍 Washington• Full-time• From C$180K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Staff Product Security Engineer Opportunity As a Staff Product Security Engineer, you will play a critical role in safeguarding Okta’s products by conducting comprehensive security reviews, guiding engineering teams in secure development practices, and handling externally reported vulnerabilities. You will engage in code reviews, penetration testing, and architectural security assessments to ensure the security of Okta’s platforms and features. This role is not suited for individuals who rely solely on automated vulnerability scanning. Instead, you must possess a deep technical understanding of web applications, backend services, penetration testing methodologies, and secure design principles. A successful candidate will have expertise in authentication protocols (SAML, OAuth, OIDC), threat modeling, and a strong desire to automate security processes by building tools that proactively identify vulnerabilities. You will also be responsible for communicating risks, impact, and remediation strategies to developers, leadership, and external audiences through documentation, presentations, and external publications. The ideal candidate will also demonstrate a deep technical background in assessing AI-integrated software architectures and securing Large Language Models (LLMs) against emerging threats and modern vulnerability classes. The ideal candidate will have an attacker mindset—the ability to think critically, creatively, and like an adversary when solvin

pythonjavaaws
View job →
R
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Storage Platform team builds and operates the platform that powers database access across Robinhood. We own relational (Postgres/Aurora), key-value (DynamoDB), and caching systems, along with the SDKs, control plane automation, and data plane services that enable safe and reliable access at scale. Our mission is to standardize and strengthen how services connect to storage, improve reliability and performance, and reduce operational overhead through automation. We manage thousands of databases and hundreds of caching clusters supporting millions of users and critical brokerage workloads. Availability is our highest priority — our systems are designed to meet strict uptime targets, including no downtime during market hours. As a Senior Software Engineer , you will build and improve core infrastructure used by many engineering teams, with a focus on reliability, performance, and operational excellence. You’ll deliver key components of data plane and control plane systems (for example: connection pooling, query routing, automation workflows, and observability) and help evolve patterns for safe, consistent database access. You’ll work closely with peers to design pragmatic s

vuesqlpostgresql
View job →
R
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Data Lake Ingestion team builds the foundations that move data from Robinhood’s operational platforms into our data lake so partner teams can make data-driven decisions. Our mission centers on operating secure and reliable ingestion paths, spanning near-real-time CDC pipelines, large-scale snapshot systems, and streaming flows via Kafka and Flink. We sit at the center of Robinhood's data platform, directly enabling analytics, regulatory retention workflows, and secure access patterns across the company. We collaborate with partner teams like Storage, Streaming Platform, and Governance to maintain high data quality and schema validation at scale. Our team is built on a culture of technical excellence and collaborative problem-solving to address deep platform engineering challenges! As a Senior Software Engineer, you will design and implement the next generation of our hybrid ingestion systems, focusing on performance, schema evolution, and validation. You will scale our flagship Rocket CDC platform and migrate legacy snapshot pipelines to newer, backup-based architectures. Your work will involve developing high-throughput streaming applications with Apache Flink and Kafka

pythonjavavue
View job →
🔔

Get new engineering excellence engineer 2 jobs by email

Daily job updates · Unsubscribe anytime