At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. The Fraud Engineering organization builds the foundational and tactical software that enables Affirm to respond strategically to fraud - both in real time and after the transaction. Our mission is to support business growth while protecting Affirm, our buyers, and our merchants. We evaluate risk at critical decision points in the user journey, rapidly adapt to evolving fraud patterns, and equip operations teams with the tools needed to investigate and mitigate fraud at scale. We’re looking for a Senior Software Engineer to join our fully remote team based in Europe. You’ll collaborate closely with stakeholders across North America - including teams in Product, Compliance, Machine Learning, Fraud Operations, Fraud Strategy, and other Fraud Engineering teams. The Fraud Engineering organization builds the foundational and tactical software that enables Affirm to respond strategically to fraud - both in real time and after the transaction. Our mission is to support business growth while protecting Affirm, our customers, and our merchants. We evaluate risk at critical decision points in the user journey, rapidly adapt to evolving fraud patterns, and equip operations teams with the tools needed to investigate and mitigate fraud at scale. The Fraud International Team (FIT) leads the engineering effort to extend Affirm's Fraud Decisioning System into markets beyond the US - Canada, the UK, Australia, and growing rapidly. We own real-time fraud decisioning for these markets, integrate region-specific risk vendors, and are building the platform that lets Affirm stand up fraud protection in a new country quickly and safely. As Affirm expands globally, FIT scales with it. We're looking for a Senior Software Engineer to join our fully remote team based in Europe. You'll collaborate closely w
Jobiba hiring network
Senior Engineering Manager Data Platform Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior engineering manager data platform jobs. Use filters to narrow by work mode, employment type, experience and date posted.
At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions
At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions
Sigma is transforming how businesses allow customers to build apps, agents and dashboards on top of governed enterprise data. Hence, we are growing the design team and looking for designers who are excited to solve challenging problems, deliver impactful capabilities throughout our stack to build world-class technology. You will be part of a talented team of designers with a shared mission to make data easily accessible for all users. We're looking for a Senior Product Designer / Design Engineer who sits at the intersection of interaction design and AI engineering: someone who uses AI to ship faster, builds the skills and evals that make AI more effective, and invents new interaction paradigms for how people work alongside intelligent systems. This isn't a traditional design role. Yes you'll be using Figma, but also writing code with AI, training it, evaluating it, and questioning every assumption about what a "UI" can be when the interface itself reasons. Please note this is a 4 day on-site role in our San Francisco office. What You'll Do Start with AI, stay with AI. Use LLMs to clarify scope, draft specs, surface edge cases, and align your team before committing to a direction, use AI coding tools to build and iterate on the solution itself, and merge code to prod when fits. Prototype in code. Build working interfaces with Cursor and Claude Code, guiding structure, behavior, interaction, motion and UX quality while AI handles implementation. Partner directly with engineering to decide what moves into the product and what stays as a validated spike. Bring it to production. Fix small interaction and refinement issues directly on prod code. Design new AI interaction paradigms for conversational interfaces. Invent and validate novel patterns for how users converse with, direct, and trust AI systems - especially in data contexts where precision and confidence matter. Write evals, skills, and help on tools. Build the scaffolding that makes AI reliabl
Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge As GTM complexity grows across sales, Solution Engineering, forecasting, enablement, and execution, leaders need stronger visibility, reliable insights, and operational support to drive performance and execute strategic priorities. Your Mission GTM Business Partnership & Performance Partner with Sales, Solution Engineering, and GTM leaders to drive performance through data-driven insights. Support forecast reviews, pipeline inspections, MBRs, QBRs, and leadership reporting. Deliver KPI analysis, executive-ready insights, and follow-up on key actions and decisions. Solution Engineering Strategy & Operations Analyze SE capacity, productivity, engagement, and alignment to sales priorities. Evaluate SE impact on pipeline, deal progression, and strategic opportunities. Support planning, coverage models, and resource allocation decisions. Strategic Analysis & Decision Support Identify trends, risks, and opportunities across sales, pipeline, productivity, and capacity metrics. Build analyses to support territory planning, forecasting, segmentation, and GTM priorities. Translate complex data into acti
Silicon Verification Engineer Multiple roles across different levels Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Bengaluru which will play a central role in Graphcore's work building the future of AI computing. The verification team sits within the Silicon design team and is responsible for ensuring that the RTL created by the logical design team and used by the physical design team matches the architecture specification for Graphcore silicon. The silicon verification engineer is responsible for verification activities within Graphcore, helping the team meet the company objectives for quality silicon delivery. Responsibilities Verification planning, specification and closure of functional coverage Providing feedback to architects Test generation and failure diagnosis/triage Contributing to shared verification infrastructure Ensuring good communication between sites Essential skills: • Verification experience in relevant industry • Proven leadership and planning skills • Highly motivated, a self starter, and a team player • Ability to work across teams and programming languages to find root causes of deep and complex issues • Experience of the verification process applied in CPU and/or ASIC environments • System Verilog, Python, C++, Linux Desirable skills: • UVM • SVA • Assembly languages • LLVM, GCC • DVCS e.g. Git • SGE or other DRMS • XML and XPath/XSLT • Web programming – HTML/DOM, Javascript, SQL Benefits: In addition to a competitive salary, Graphcore offers a competitive benefit
Silicon Physical Design Engineer Multiple roles across different levels Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Bengaluru which will play a central role in Graphcore's work building the future of AI computing. The physical design team sits within the wider silicon design team which includes RTL, verification and DFT. Our work also involves strong links with architecture, packaging and product engineering. We are responsible for working with those teams to create high-quality RTL and building the final chip layout (e.g. GDSII) ensuring a signoff-quality design is delivered to the Foundry (e.g. TSMC). We are looking to hire high-quality silicon physical design engineers to join our team. The successful candidate will support the team with achieving our goals and creating the right engineering solutions. We are a collaborative team and good communication is essential, as is the ability to adapt and learn. For the successful candidate we offer an open, honest and collaborative environment working on leading-edge designs at the most advanced nodes. Our engineers are not siloed, and they are trusted and encouraged to ta ke ownership of their designs and problem solutions. You will be part of a team that looks for improvements to everything we do: our designs, our flows, our methodologies, our infrastructure. Responsibilities and Duties Applicants will be expected to contribute technically to the development of Graphcore's next generation of AI superchips, fo
Senior: GBP 73,500 - 99,500 Staff: GBP 97,300 - 131,700 Subject to alignment to the responsibilities and duties of the role - we currently have multiple positions available at both Senior and Staff level. About the job Build the Linux distribution foundation that turns upstream software into trusted Graphcore platform releases. You will help create the Linux distribution that powers Graphcore AI systems. The team produces production-ready system images from proven upstream distributions. Your work will shape how releases are built, validated and prepared for deployment. You will strengthen the engineering path from upstream Linux software to dependable platform releases. You will build and improve automated pipelines, run established Linux test suites, and diagnose issues across build and validation flows. As the platform evolves, you will introduce controlled configuration and tuning changes with evidence-led validation. This is hands-on systems engineering with visible impact. You will help define reliable processes for a new team building a critical part of Graphcore’s platform. The team and culture You will join one of Graphcore’s newest engineering teams, helping shape its culture from the start. It is a small, co-located team where ownership matters and progress is visible. Work happens through close technical discussion, practical problem-solving and evidence-led decisions. Ideas are challenged openly, and the best path wins regardless of hierarchy. You will report to a leader who values technical credibility and invests in people’s growth. The team moves with pace, takes responsibility and changes direction when the evidence demands it. What we're looking for Strong practical experience working in Linux environments Experience building or maintaining automated CI/CD pipelines for reliable engineering workflows Proficiency in Python, Bash or similar
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng
Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Lead the technical design and architecture of major platform initiatives, author design documents and build consensus across engineering teams. Define technical roadmaps for complex, multi-quarter projects that span multiple teams. Make critical architectural decisions for company documentation infrastructure, balancing scalability, reliability, and developer experience. Evaluate and set direction for integrating emerging technologies, including AI/LLM capabilities, into company documentation platforms and authoring tools. Establish and evolve engineering standards, best practices and technical guidelines for the team and broader organization. Partner with engineering teams across the company to understand documentation needs and design integrated solutions. Design, build and maintain scalable, reliable and performant services and systems. Contribute high-quality code across the full stack and navigate codebases with different languages and tools. Debug and resolve complex production issues and improve system reliability. Take ownership of system health and incident response. Who you are Minimum requirements Must have a Bachelor's degree or foreign equivalent in Computer Science, Software Engineering, Engineering, or a related field, plus four (4) years of experience in Software Engineering. Must have four (4) years of experience in each of the following: - Working in a full stack environment with a foc
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About The Team The Commerce Site Reliability Engineering team is responsible for the reliability, scalability, and day-to-day operation of the platforms that power GoDaddy's Commerce ecosystem. We build and operate shared infrastructure, support critical production systems, and partner closely with engineering teams to ensure services remain secure, resilient, and highly available. As a Senior Site Reliability Engineer, you'll join a team that values ownership, operational excellence, and continuous improvement. Engineers are empowered to identify problems, drive meaningful change, and influence how reliability is delivered across the broader Commerce organisation. From improving operational maturity and reducing toil to modernising delivery platforms and strengthening incident response practices, this team plays a key role in enabling engineering teams to move quickly and safely. You'll work closely with engineers across infrastructure, cloud, security, networking, and application teams while helping shape the future of reliability engineering at GoDaddy. This role offers significant opportunity to broaden your impact, develop technical leadership skills, and grow toward Staff and Principal engineering positions over time. What you'll get to do... Lead reliability and operational improvement initiatives across GoDaddy's Commerce platform, helping engineering teams build and operate services safely and at scale. Own critical production systems, drive incident response and post-incident improvements, and continuously raise the bar
The Applied AI team designs and builds algorithmically driven features in the Datadog app. We work across a range of applications, primarily focusing on analysis on streaming data such as anomaly detection, error outliers and faulty deployment analysis. As an Applied Scientist you will work on building models and algorithms for machine learning powered features within the Datadog platform. You will work closely with our engineering and product partners to explore, build, scale and deliver these features that we incubate within the Applied AI team. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Design solutions for our different use cases. You will research and benchmark relevant algorithms to find the best fit for our use-cases Leverage machine learning algorithms and statistical techniques to build new scalable product features Develop, deploy and monitor new and existing features to production Participate in our journal club by reading and presenting the latest academic research papers to the team Explore, analyze and tell the story behind high volumes of data flowing through Datadog systems Maintain and monitor the models, services and infrastructure owned by your team Participate in your team’s on-call rotation Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering, Machine Learning or related scientific field or equivalent experience You have experience working with high-scale systems and datasets including building models, applying machine learning to real business problems, and writing production data pipelines You can explain complex ideas and algorithms to non-technical audiences You care about code simplicity and performa
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in a global on-call rotation supporting highly available customer-facing systems. Participate in incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with en
NVIDIA is leading company of AI computing. At NVIDIA, our employees are passionate about AI, HPC , VISUAL, GAMING. Our SA team is more focusing to bring NVIDIA new technology into difference industries. We help to design the architecture of AI computing platform, analysis the AI and HPC applications to deliver our value to customers, focusing on defining and solving computational challenges in LLM inference and training acceleration, as well as network communication and data transfer optimization. What You'll Be Doing: Contribute to the development of open-source inference frameworks such as SGLang and vLLM, including feature and operator development, performance optimization, and model support, in collaboration with the community. Develop and optimize KV cache offloading frameworks for LLM workloads, supporting multi-level cache offloading and reuse across CPU, SSD, and remote storage to improve inference efficiency. (Team project: FlexKV) Drive R&D on compute performance in distributed training, and explore methods and technologies for performance optimization. Study computational challenges in machine learning systems, identify common needs and bottlenecks, and build example code, acceleration libraries, or frameworks accordingly. What We Need to See: Over 5 years working experience in the technology industry, with master’s degree or above in computer science, mathematics, electrical engineering, automation, or related fields. Strong interest in accelerated computing, parallel computing, and heterogeneous computing, with the motivation to explore these areas in depth. Solid programming skills, with a good understanding of data structures and computer systems fundamentals. Strong learning agil
We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<
Get new senior engineering manager data platform jobs by email
Daily job updates · Unsubscribe anytime