Jobiba hiring network

Senior Systems Engineer Product Platform Tools Jobs

7,101 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior systems engineer product platform tools jobs. Use filters to narrow by work mode, employment type, experience and date posted.

A
Affirm
📍 Spain• Full-time• Remote• From €1M/yr
21 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions

REMOTEpythonsqlmysql
View job →
DU
DoorDash USA
📍 San Francisco• Full-time• From $1.6M/yr
21 days ago

About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that will run DoorDash's data, analytics, and ML compute for the next five years and beyond. You will own the deep, cross-cutting problems that span the runtime, the shuffle service, the scheduler, and the overall service reliability — making the architectural calls that compound across the platform's lifetime. You will partner with the Engineering Manager on technical roadmap, hiring, and team shape, and act as the senior technical voice in cross-team partnerships with Data Engineering, ML Platform, and product engineering teams that depend on the platform. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Set the multi-year technical direction for an in-house Spark-on-Kubernetes platform — runtime, shuffle, scheduler, reliability — and make the architectural calls that compound for years. Own the deepest distributed-systems problems on the team: shuffle architecture, multi-tenant scheduling, runtime performance, and the failure modes that only show up at scale. Partner with the Engineering Manager on technical roadmap, hiring, inte

pythonjavasql
View job →
N
Nvidia
📍 Remote, United States• Remote
17 days ago

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data

REMOTEpythonsqlaws
View job →
M
Mongodb
📍 Gurugram• Full-time
1mo ago

The Infrastructure Engineering team is responsible for building and maintaining a self-service internal development platform that enables MongoDB engineering teams to reliably deploy and operate their own production services and products. We work with numerous engineering teams across the company to understand their infrastructure requirements and development workflows, develop broadly applicable self-service platform services and tooling, continuously monitor how platform services are being utilized, and look for ways to improve developer productivity through automation and education. We are big open source enthusiasts and use a number of open source tools in our stack (contributing upstream whenever possible). Some of the tools we use regularly include Go, AWS, Kubernetes, Crossplane, Terraform, Helm, Drone, Prometheus, and Grafana. However, technology is nothing without a stellar team of engineers that are focused on doing high quality work and working as a team to solve complex distributed computing and platform engineering problems. This is where you come in! We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Our ideal candidate Has built and operated large-scale distributed systems in cloud providers (AWS strongly preferred) Has a strong backend programming background. Fluency in Go is strongly preferred; deep experience with another compiled or strongly-typed backend language is acceptable Has experience working with AI coding agents and can demonstrate building high quality context to yield high quality outputs Has experience designing and implementing medium-to-large software projects, including driving design reviews and mentoring less-senior engineers Pragmatic, detail-oriented, self-motivated, and understands the benefits of collaboration Strong experience operating production Kubernetes clusters, not just deployed to it Has practical experience defining and operating against SLI/SLOs for services they owned Str

mongodbawsazure
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $187K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . The Enterprise Applications & Architecture (EAA) team within Platform is the primary engineering partner for the Sales, Trading, and Prime (STP) business, building and operating Salesforce capabilities that power institutional sales, account management, and billing workflows. As a Senior Enterprise Engineer on this team, you'll design and deliver scalable, secure CRM and billing solutions while integrating Salesforce with pricing, contracting, and downstream finance systems. You'll join a small, high-ownership engineering team that partners closely with Sales Operations, Account Management, Billing, and Institutional Operations to unlock revenue, improve forecasting, and create resilient, auditable processes. What you'll do: Own end-to-end Salesforce solution design and delivery using Apex, Lightning Web Components, Flows, and managed packages across Sales Cloud and Service Cloud Build and operate integrations between Salesforce and internal platforms using REST APIs, Platform Events, middleware, ETL tooling, and contracting systems (e.g., Ironclad) Author technical design documents that account for security, data architecture, system limits, and scalability Partner with product managers, architects, and business stakeholders (Sales, Legal, Compliance, Institutional Ops) to translate requirements into technical designs and implementation plans Drive engineering q

REMOTEawsci/cdgit
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $186.1K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Software Engineer on the AI Platform team within the Platform group, you'll build and operate the LLM and agent infrastructure that every team at Coinbase depends on. This team owns the company's single path to large language models and the full agent lifecycle: build, deploy, run, observe, and improve. You'll lead multi-quarter technical initiatives across the platform, from gateway and runtime systems to knowledge bases and applied AI agents, directly shaping how Coinbase scales AI across the organization. What you'll do: Own the architecture and delivery of core platform systems including the LLM Gateway (60+ models, auth, PII redaction, fallbacks, cost optimization), AI Hub, and agent runtime with microVM sandboxes and governed MCP gateway Drive the design and implementation of Knowledge Base infrastructure, connecting data sources to auto-provisioned vector and markdown stores queryable by any agent Lead AI FinOps capabilities including spend attribution, governance, and cost optimization across all AI workloads company-wide Partner across engineering, security, legal, finance, product, and external partners at frontier labs and major cloud providers to ship high-impact platform capabilities Build evaluation and observability tooling including LLM-as-judge harnesses, full tracing, and feedback loops that let subject matter experts refine production a

REMOTEpythonawsmicroservices
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $180.4K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Senior Analytics Engineer As a Senior Analytics Engineer on the Platform team, you'll build the scalable data models and pipelines that power analytics, experimentation, and decision-making across Coinbase. Our Analytics Engineering team transforms raw data into trusted, well-modeled sources that stakeholders across Product, Engineering, and Data Science rely on daily. You'll own end-to-end data solutions for specific business domains, turning complex data flows into clean, reusable frameworks that unlock commercial value at scale. What you'll do: Own end-to-end data modeling for assigned business domains, from understanding source system data flows through designing modular, reusable models (star/snowflake schemas) that serve as the single source of truth for downstream teams. Build and optimize ETL/ELT pipelines using modern tools like dbt and Airflow, ensuring data quality, reliability, and performance at scale across Snowflake or similar warehouse architectures. Partner with Engineering, Product, and Data Science teams to identify data gaps, define requirements, and deliver data products that directly enable experimentation, ad hoc analysis, and business metric optimization. Develop scalable abstractions and frameworks (UDFs, Python packages, internal data apps) that multiply the efficiency of other data teams and reduce time-to-insight across the organization. D

REMOTEpythonsqlaws
View job →

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 The Mission The Foundry is ClickUp's internal AI innovation lab — embedded inside GTM Systems and accountable for turning AI capabilities into production-grade, internally deployed products that make every GTM function faster and smarter. We build the infrastructure that powers AI-first work across Sales, Marketing, Post-Sales, and Revenue Operations. As the Senior Software Engineer on this team you will own the technical delivery of our MCP server platform, agent orchestration layer, and internal tooling — shipping production systems used daily by hundreds of ClickUp employees, and scaling your own throughput by treating AI tools as first-class engineering collaborators. What You'll Own MCP Server Platform Design, build, and operate Model Context Protocol servers that expose CRM, ticketing, analytics, and communication data to AI agents across the GTM stack Implement Okta PKCE authentication flows and RBAC policy enforcement so agents access only the data they're authorized to touch Maintain deployment infrastructure on AWS (Bedrock, Lambda, ECS, API Gateway) and contribute to GCP workloads where applicable Own observability: structured logging, distributed tracing, latency SLOs, and on-call runbooks for every production server Agent Orchestration & AI-Native Products Build and maintain multi-step autonomous agents that execute end-to-end GTM workflows — lead qualification, deal room assembly, onboarding automation, support triage, and more Architect prompt engineering frameworks, tool-call schemas, and agent evaluation harnesses that make AI behavior predictable and auditable Integrate with LLM p

typescriptpythonreact
View job →
V
Vanta
📍 United States• Full-time
1mo ago

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. As a Senior Applied AI Engineer at Vanta, you will play a crucial role in shaping Vanta’s AI offerings, setting technical strategy, and leading projects that leverage AI to deliver smarter, faster outcomes for our customers. You'll be part of a team integrating AI into the Vanta product, working alongside a multidisciplinary group of product engineers, machine learning engineers, product managers, designers, and security and compliance experts to implement, scale, and maintain AI-enabled product experiences. In this role, you’ll build products that enable Vanta’s customers to leverage AI to accelerate their journey towards compliance, managing risk, and earning trust. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll do as an engineer working on Applied AI at Vanta: Work cross-functionally to design and implement AI-powered features to deliver customer value and integrate LLMs with Vanta’s existing products and systems. You’ll work with other product engineers across Vanta to understand how AI systems can accelerate product adoption at Vanta Instrument evaluations, guardrails, and monitoring, and review customer usage to continually improve quality Collaborate with AI Platform engineers shaping foundational AI systems and tooling that accelerate product teams Make pragmatic tradeoffs that consider business priorities, user experience, and a sustainable technical foundation Mentor engineers, champion good technical and product instincts, and model a collaborative, high-ownership engineering culture How to be successful in this role: At least 7 years of industry experience as a software

typescriptreactnode.js
View job →
L
Linear
📍 United States• Full-time
1mo ago

At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. As a small team, we’re all generalists that work across the full stack (built in Typescript end-to-end). We’re looking for experienced engineers that thrive in an environment of autonomy and individual responsibility to help us build the future of product development. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in the US and Europe. You can work from anywhere within those regions. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you'll do Build new user-facing features with everything from database models to GraphQL resolvers and UI components Optimize our data synchronization stack by applying better serialization protocols Add real-time collaborative editing to our content editor Improve performance by profiling and tweaking virtualized list rendering Add analytics, monitoring, and alerts to our service so that we can better respond to operational incidents Open-source any non-trivial innovations that come out of our work on the product Redefine best-in-class software development processes so that we can build a purpose-built product. What we're looking for 5+ years of experience building customer-facing products at a high-quality software company Strong React and Typ

typescriptreactsql
View job →
G
GHX
📍 Hyderabad• Full-time
21 days ago

Role: Senior AI Engineer Location: Hyderabad, India (Hybrid) Department: Product Development About the Role GHX is building a cutting-edge LLM-powered document understanding platform focused on classification, structured data extraction, and intelligent orchestration at scale. This is a high-impact AI engineering role where you will own the full lifecycle—from problem framing to production deployment . Initially, you will focus on prompt engineering and evaluation systems , building the quality foundation for AI performance. Over time, the role expands into agent orchestration, system architecture, and migration of rule-based systems to LLM-driven pipelines . A strong foundation in software engineering (5+ years) is essential. This role demands engineering rigor across both traditional system design and AI system behavior . Core Responsibilities 1. Prompt Engineering Design prompts for diverse document classification and extraction tasks Treat prompts as formal specifications (precise, structured, and edge-case-aware) Develop few-shot, chain-of-thought, and structured output templates Manage prompt lifecycle: versioning, testing, and rollback 2. LLM Output Evaluation Create and maintain ground truth datasets Build automated evaluation pipelines (precision, recall, field-level accuracy) Identify and resolve conceptually incorrect outputs despite surface correctness 3. AI Agent Orchestration Design multi-agent workflows for document processing Implement tool-use patterns and integrate MCP servers Optimize orchestration for scale and efficiency 4. Software Engineering Develop production-grade APIs and backend services Apply Clean Architecture / DDD principles Write maintainable, testable Python code Contribute to CI/CD, deployment, and observability systems 5. Stakeholder Collaboration Act as a bridge between business stakeholders and AI systems Translate product requirements into technical architectures Communicate system behavior, limitations, and quality

pythonawsazure
View job →
C
Cohere
📍 Toronto• Full-time• Remote
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About North: North is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications. In this role, you will: Build and ship features for North, our AI workspace platform Develop autonomous agents that talk to sensitive enterprise data Write and ship minimal code that runs in low-resource environments, and has highly stringent deployment mechanisms As security and privacy are paramount, you will sometimes need to re-invent the wheel, and won’t be able to use the most popular libraries or tooling Collaborate with researchers to productionize state-of-the-art models and techniques You may be a good fit if: Have shipped (lots of) fullstack code (Python and React) in production You excel in fast-paced environments and can execute while priorities and objectives are a moving target You’ve worked in both large enterprises and st

REMOTEpythonreactai
View job →
C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About North: North is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications. In this role, you will: Build and ship features for North, our AI workspace platform Develop autonomous agents that talk to sensitive enterprise data Write and ship minimal code that runs in low-resource environments, and has highly stringent deployment mechanisms As security and privacy are paramount, you will sometimes need to re-invent the wheel, and won’t be able to use the most popular libraries or tooling Collaborate with researchers to productionize state-of-the-art models and techniques You may be a good fit if: Have shipped (lots of) fullstack code (Python and React) in production You excel in fast-paced environments and can execute while priorities and objectives are a moving target You have strong coding abilities and are comfo

pythonreactgit
View job →
D
1mo ago

We are building the best platform in the world for engineers to understand, scale, and protect their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, application tracing, and security insights for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way . At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Solve a scaling bottleneck in a critical service Deploy a new feature to production, progressively rolling it out with feature flags Investigate and fix a production issue from a service your team owns Design a way to scale up a service for more traffic With your team, plan the most important projects to work on next Who You Are: You have significant experience in one or more languages You value code simplicity and performance You can design architecture to solve problems at high scale You have a BS/MS/PhD in a scientific field or equivalent experience You want to work in a fast, high-growth startup environment that respects its engineers and customers You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how You have demonstrated ability to use AI coding tools in day-to-day workflows and build, validate, and refine AI-generated output in products You can design AI Backend systems, with awareness of quality, cost, and latency tradeoffs Bonus: You’re motivated to push the boundaries of how AI can improve software engineering best practices and contribute to building AI-enabled products Datadog values people from all walks of life. We understand not everyone will meet all

aigorust
View job →
R
Replit
📍 Foster City• Full-time• Remote
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About Replit Replit is building the world's most ubiquitous AI coding agent. Replit Agent can be used by anybody to bring their ideas to life. Whether it's an app for yourself, the next great startup idea, or a tool to make you more productive at work, Replit Agent can help build it. Replit is also the leader in secure vibe coding. We protect apps, give users features to manage security risks, and help them vibe code more safely. About the role: Replit is changing how people and companies turn ideas into software. Reaching those customers requires engineering that connects acquisition to the product experience and gives teams trustworthy evidence about what works. This role goes beyond operating conventional marketing technology. You will rethink growth systems for a world where agents can observe performance, diagnose problems, take action, and learn from the result. As the founding engineer for Growth Enablement, you will set the technical direction for this area. You will build agentic systems alongside shared foundations for attribution, audiences, lifecycle engagement, referrals, and promotions. The work spans web, mobile, billing, and data. You will work directly with marketing, sales, and partnerships to find the highest-leverage problems and ship the first solutions. Successful projects will become platforms that help these teams move faster and give Replit a clearer view of what drives durable growth. You will: Design agentic growth systems that monitor performance, diagnose failures, run approved experiments, and improve from the results. Explore AI-native approaches to answer engine optimization, campaign operations, audience discovery, and measurement. Build acquisition measurement across web and mobile, in

REMOTEtypescriptpythonsql
View job →
🔔

Get new senior systems engineer product platform tools jobs by email

Daily job updates · Unsubscribe anytime