Integration Reliability Engineer, Technical Operations About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team APAC Payins TechOps is a newly formed team based in Singapore. Our charter is to improve the health and resilience of our payment acquiring volume, making it easier for Stripe to build and operate the integrations that support hundreds of billions of dollars in payments annually. Our team partners closely with payments and platform engineering teams to assess systems or processes that create high operational workloads, then develops durable solutions via code, tooling, data, and process improvements. We are responsible for the financial data quality of key systems at Stripe, ensuring that data is quarantined without impacting downstream systems while implementing the right changes upstream to prevent recurring issues. The team's work has a direct impact on Stripe's ability to expand into new markets and offer more sophisticated payment features to merchants. What you’ll do Responsibilities Scope and lead technical initiatives end-to-end: identify problems worth solving, propose the right solution approach, and deliver on that solution — not just execute on a pre-defined plan. Investigate problems in systems by tracing problems through Stripe’s stack. You’ll examine code, write SQL queries, read logs, and inspect data pipelines to understand system behavior, then make changes to address. Examine updates being made by Stripe’s financial partners to understand impact, and make the changes within Strip
Jobiba hiring network
Tech Ops Team Lead Jobs
1,774 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current tech ops team lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Community Support Engineering (CSE) China Foundations team works together with our US partners to co-own the Community Support foundations systems, collectively empowering every step of the customer support journey. We're building the next-generation case management platform that gives a 360-degree view of every customer interaction, helping ambassadors resolve guest and host issues faster. At the same time, we're developing the centralized supervisor portal alongside agent quality systems that turn performance insights into actionable coaching signals. Our routing platform connects Airbnb customers to the right ambassador, through the right channel, at the right time. The team also provides AI-powered solutions to support TechOps around system inconsistencies and incidents to ensure high availability and reliability of critical CS flows. The Difference You Will Make: We are seeking an experienced Engineering Manager to lead the China Routing and Data Service team. In this role, you will: Ramp up on the team’s technology, product, and business goals, including hands-on work to develop a deep technical understanding. Lead, mentor, and grow a high-performing engineering team in a fast-paced environment. Foster a team culture that emphasizes product thinking, technical depth, and customer-centric problem solving. Drive and partner closely with teams across CSE to understand business cases and deliver intuitive, scalable routing and data solutions. Set clear priorities and strategic direction in collaboration with cross functional partners.. Improve team processes to strengthen coll
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Staff Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimising costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimised for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimise AWS cloud infrastructure, including EKS, ECS, S3, VPCS, RDS, IAM, and more. I
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. Stripe was built with simplicity in mind. We strive to deliver frictionless experiences for all of our users, whether they are an Independent Business, Startup, SMB, or Enterprise, and our mission is to provide all Stripe Users with the best support experience possible. Today, Stripe handles over a million support cases per year and processes millions of internal transactions. We're going to achieve excellence by thinking of support in a novel, solution-oriented way, and viewing operations as an integral enabler of all of Stripe's growth. About the team At Stripe, our Mexico City office is a vibrant hub at the forefront of our mission to reshape the financial landscape for businesses worldwide. Our team prioritizes collaboration, innovation, and excellence. As a member of the Mexico City team, you'll be part of a mission-driven community dedicated to enhancing the global economy and increasing the GDP of the internet. We strive for excellence by creating with craft and beauty, while having fun and celebrating our successes together. We cultivate a culture of collaboration, inclusivity, and support where every team member's voice matters. Our commitment shines through as we handle over a million support cases each year, empowering our users not just to solve problems but to achieve their goals. What you'll do Responsibilities Troubleshoot and solve complex problems Analyze our processes and instigate changes to help scale our operations and improve user exp
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Auth0 provides an unparalleled authentication experience for hundreds of millions of users worldwide. Our commitment to reliability is a key foundation of our product and our dedication to exceeding customer availability expectations is a core engineering focus. As a Senior Site Reliability Engineer, you'll join our SRE team based in Europe to ensure our production systems are not only operational but also resilient, scalable, and ready for exponential growth. This isn't just about keeping the lights on; it's about directly contributing to the platform's core resiliency and robustness. You'll be a hands-on builder, crafting solutions that make our system more reliable by design. What you’ll do: Design and build custom software in Go to enhance the platform's reliability, resiliency, and redundancy. Partner with engineering teams to embed reliability principles, improving the availability, performance, and observability of our services. Use your deep understanding of infrastructure and observability principles to identify opportunities for improvement within the product and implement solutions. Contribute to our follow-the-sun on-call rotation, providing rapid, effective response to critical incidents and using your expertise to troubleshoot, mitigate or accurately escalate production issues. Because our team is globally distributed, your on-call shifts will only occur during your standard local working hours. Develop and refine our SRE tooling and proc
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. With the Okta's Auth0 organization’s increased dedication to ensuring customer availability expectations are exceeded in every way, you will play a key role as we evolve our system architecture to meet the demands of enormous growth and support the hundreds of millions of users who rely on us to provide uninterrupted access to business-critical Reporting to the Manager of Engineering, in this role as a SRE Operations Engineer, you will ensure smooth operations of our Customer Identity Cloud at Okta. Working closely with the SRE team, your primary focus will be on ensuring production systems remain operational at all times, while continually setting and achieving long-term operational success for the platform with potential career growth into Site Reliability Engineering. What you’ll be doing Executes operational work including updating/patching and maintaining the Engineering Service Desk queue Responsible for ensuring team requests are triaged and/or actioned in a timely manner Monitors Platform health and take steps to alleviate issues related to deployment and operations Assist with capacity, performance and scalability testing where required Escalation point for Platform issues from customer support teams Execute runbooks and update processes as required Interface with the SRE team to report core issues, required improvements and new feature requests What you’ll bring to the role General platform infrastructure knowledge, including high availability / l
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. As a Senior Site Reliability Engineer you will champion all things pertaining to reliability at Okta for Auth0. Working closely with the Product Engineers, Quality Engineers, Platform Engineers and Architecture teams, your primary focus will be on ensuring production systems remain operational at all times, while continually setting and achieving long-term performance, reliability and scalability goals in a platform with an exponential growth plan for the coming years. With Okta’s increased dedication to ensuring customer availability expectations are exceeded in every way, you will play a key role as we evolve our system architecture to meet the demands of enormous growth and support the hundreds of millions of users who rely on us to provide uninterrupted access to business-critical enterprise and consumer applications. Skills Exceptional communication skills, including technical writing in the English language Systematic problem-solving approach, coupled with a strong sense of ownership and drive Understanding of microservices, cloud infrastructure (AWS, Azure), databases (SQL, No-SQL, Key/Value), containers (docker, kubernetes), web technologies (web sockets, http) and networking (SSL, routing, VPN) Live and breathe SLIs, SLOs, error budgets and SLAs Strong belief in automating everything and reducing toil for yourself and teammates Loves to work as a team, but is able to work effectively in a remote environment where tasks may be self-driven Knowledge
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies from the world's largest enterprises to the most ambitious startups use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team As Stripe's user base and global footprint grows dramatically, we face uniquely complex operational challenges at the intersection of payments infrastructure, financial data integrity, and global scale. The Stripe Delivery Center (SDC) provides operational leverage across Stripe's portfolio supporting external users and internal teams with precision and speed. What you’ll do As a Technical Operations Analyst on the Payments team, you will own the reliability and accuracy of some of the world's most critical payment flows, a platform processing hundreds of billions of dollars annually. You will sit at the intersection of engineering, finance, and operations: triaging live payment failures, driving reconciliation accuracy, and partnering with financial institutions globally to resolve issues fast. This is not a passive monitoring role. You will be expected to independently own investigations end-to-end, identify systemic issues before they escalate, and drive process improvements that scale with Stripe's growth. Responsibilities Own reconciliation and payments-related investigations, ensuring accurate ledgering of funds and resolving discrepancies with banking partners swiftly and precisely. Proactively monitor payment performance across users and payment methods; triage degradation alerts, diagnose root causes, and communicate findings clearly to technical and non-technical stakeholders. Serve as the escalation point for complex support inqui
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team As Stripe’s user base and global footprint grow dramatically, we have distinctly unique support problems resulting from both our type of scale and the type of businesses we partner with. The Stripe Delivery Center (SDC) strategy will provide operational leverage and expand Stripe’s portfolio of operational capabilities to support the scaled needs of external users and internal Stripe teams. At Stripe, our Mexico City office is a vibrant hub at the forefront of our mission to reshape the financial landscape for businesses worldwide. Our team prioritizes collaboration, innovation, and excellence. As a member of the Mexico City team, you'll be part of a mission-driven community dedicated to enhancing the global economy and increasing the GDP of the internet. We strive for excellence by creating with craft and beauty, while having fun and celebrating our successes together. We cultivate a culture of collaboration, inclusivity, and support where every team member’s voice matters. Our commitment shines through as we handle over a million support cases each year, empowering our users not just to solve problems but to achieve their goals. What you’ll do As an analyst on the Payments Health Operations team, you will be charged with monitoring and maintaining payments performance for Stripe’s largest users. This will include triaging, investigating, and responding to detected regressions in authentication and cost rates; providing in-depth analysis of
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Trust and Safety Operations team is focused (ok, maybe obsessed) on scaling Roblox’s Operations organization and transforming our customer experience through our multi-year vision and strategy execution. You will be reporting to the Trust & Safety India Operations Manager. The team provides support to Roblox’s global players, developers, and advertisers. As a Product Support Specialist on our Safety Ops Team, you’ll be focused on Escalations, Product Reliability, Process Improvement and Design along with playing a critical role in managing and resolving escalations and high-impact incidents affecting our Roblox users and platform. You will coordinate across multiple teams to ensure rapid response to minimize the impact on customers and business operations, identifying trends, enhancing support processes, improving agent efficiency, and reducing operational costs while maintaining high reliability and service quality. Additionally, you will serve as a product expert working closely with the PSM team, providing detailed synthesis of new and existing issues our users experience to prioritize and enable engineering teams to improve the platform. You Will: Coordinate a
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. About the Role The Trust and Safety Operations team is focused (ok, maybe obsessed) on scaling Roblox’s Operations organization and transforming our customer experience through our multi-year vision and strategy execution. Creators are the engine of Roblox. When something breaks for them, whether it is a broken publish flow, a payout that didn't land, an asset that won't upload, how quickly and how well we resolve it directly shapes whether they keep building with us. You will be reporting to the Trust & Safety India Operations Manager. The team provides support to Roblox’s global players, developers, and advertisers. As a Senior Product Support Specialist on our Safety Ops Team, you’ll own the hardest escalations end to end, drive root cause analysis with engineering and product partners, and then close the loop by turning what you learn into support content that scales for both our human agents and our automated support experiences. Additionally, you will serve as a product expert working closely with the Product Support Management team, providing detailed synthesis of new and existing issues our users experience to prioritize and enable engineering teams to improve the platform. You'
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Content Systems is the heart of Lyft's language ecosystem. The team is made up of two distinct disciplines, Content Design and Learning Experience Design, who create and own language across Lyft's in-app experiences that help customers succeed on the Lyft platform. We're looking for a Content Systems intern to join our team for the summer. This role will both help design and build our internal content system (think information architecture), and embed in a rotation of design teams as a content design/learning design partner in building products across the Lyft community. Of course, you won’t be alone. You'll work with designers, researchers, marketers, and product managers day in and day out to develop experiences that reach and resonate with riders. This internship will take place in our Toronto office during Summer 2027 . What we're looking for Responsibilities: Support our content system by establishing and translating information architecture to help manage content across Lyft Help teams reconsider existing systems and processes (eg publishing workflows) to align to the new information architecture Support product strategy and vision through content, working on projects as a content designer and learning experience designer Translate complicated concepts into clear in-product copy (UI, notifications, errors, and more) and learning materials Collaborate with designers, visual content creators, researchers, engineers, marketers, data science, product managers, and other stakeholders across tech and ops The ability to translate concepts across the UI, learning center, and knowledge base into connected content objects Manage multiple projects and competing priorities Maintain, evolve and champion the Lyft brand voice and style Skills: Excellent writing skills; ability to
About the Team: Being part of Meesho's Fulfilment and Experience team will zip you to the cockpit of our ever burgeoning rocketship, where you get to directly shape the experience of the country's next billion e-Commerce users. We are an eclectic mix of 100+ professionals with diverse skill sets ranging from running operations / support, supply chain know-how, analytics and the holy grail, first principles problem solving. At Meesho, we are trying to do what's never been done before - taking e-Commerce to the masses. This leaves us with no choice but to completely reimagine logistics from the ground up, to cater to our customers' price and delivery expectations. That means a host of "zero-to-one" projects (takers, anyone?) to build a supply chain which changes how folks think about e-Commerce not just in India, but globally. We focus on personal growth and fun at work just as much as we do on working hard. About the Role As an Assistant Manager - F&E, you'll be responsible for identifying key problems, setting the priorities, coming up with solutions and driving implementation. In order to drive implementation, you'll get complete autonomy in terms of the team and processes that you would want to set up. You'll also be responsible for shaping up the right solutions in coordination with the product team, in case your solution requires tech interventions.
At Trustpilot, we're on an incredible journey. We're a profitable, high-growth FTSE-250 company with a big vision: to become the universal symbol of trust. We run the world's largest open customer review platform, and while we've come a long way, there's still so much exciting work to do. Come join us at the heart of trust! We are growing our engineering team at Trustpilot and are looking to welcome a Software Engineer I into the Trust Tech department! You will join a cross-functional team with full ownership of our products and codebase, where you will take part in every step of the development process, from ideation to maintenance. This is a place where you can grow as an individual, learn from senior mentors, and have a real influence on the direction of our projects. About the team: You will be joining a brand new team being built within the Trust Tech department, focusing specifically on businesses. The team is responsible for building and maintaining automated systems that enforce Trustpilot’s terms of use for businesses. This includes developing systems for automatic detection, enforcement of misuse, and internal investigation tools. This work is vital to maintaining Trust and Transparency on the platform. To achieve this, we collaborate closely with Data Scientists, Data Ops, and ML Ops to build sophisticated automated architectures that integrate detection models and robustly scale our misuse enforcement. What you’ll be doing: Work in a cross-functional “full ownership” team alongside Product, Design, and Data Science. Implement and release new features with a focus on backend stability and scalability. Help build solutions to handle high-volume data processing for fraud detection. Maintain and improve the internal tooling frontends (React) used for investigations. Troubleshoot existing software, squash bugs, and learn how to optimize database performance. Participate in technical discussions and learn best practices for Infrastructure a
At Trustpilot, we're on an incredible journey. We're a profitable, high-growth FTSE-250 company with a big vision: to become the universal symbol of trust. We run the world's largest open customer review platform, and while we've come a long way, there's still so much exciting work to do. Come join us at the heart of trust! We are growing our engineering team at Trustpilot and are looking to welcome a Software Engineer I into the Trust Tech department! You will join a cross-functional team with full ownership of our products and codebase, where you will take part in every step of the development process, from ideation to maintenance. This is a place where you can grow as an individual, learn from senior mentors, and have a real influence on the direction of our projects. About the team: You will be joining a brand new team being built within the Trust Tech department, focusing specifically on businesses. The team is responsible for building and maintaining automated systems that enforce Trustpilot’s terms of use for businesses. This includes developing systems for automatic detection, enforcement of misuse, and internal investigation tools. This work is vital to maintaining Trust and Transparency on the platform. To achieve this, we collaborate closely with Data Scientists, Data Ops, and ML Ops to build sophisticated automated architectures that integrate detection models and robustly scale our misuse enforcement. What you’ll be doing: Work in a cross-functional “full ownership” team alongside Product, Design, and Data Science. Implement and release new features with a focus on backend stability and scalability. Help build solutions to handle high-volume data processing for fraud detection. Maintain and improve the internal tooling frontends (React) used for investigations. Troubleshoot existing software, squash bugs, and learn how to optimize database performance. Participate in technical discussions and learn best practices for Infrastructure a
Get new tech ops team lead jobs by email
Daily job updates · Unsubscribe anytime