Jobiba hiring network

Senior Systems Design And Integration Specialist Jobs

7,101 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior systems design and integration specialist jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine

pythonkubernetesci/cd
View job →
M
20 days ago

MeltPlan | Planning Engine for the Built Environment MeltPlan is building the “planning engine” for the $14 Tn construction industry, an AI system designed specifically to optimize decisions before construction begins. While design software optimizes use and aesthetics and construction software optimizes execution and control, MeltPlan is building the missing layer - software that optimizes decisions and tradeoffs upstream, before scope is locked, procurement begins, and change orders become inevitable. MeltPlan’s long-term goal is to help teams make construction “boring” by making planning more intense: surfacing constraints and tradeoffs early, aligning stakeholders before plans are frozen, and reducing the need for late-stage redlines, rework, and change orders. MeltPlan is founded by operators who have built at scale. Kanav previously co-founded Innovaccer, a $3Bn health-tech company focused on making US healthcare more affordable and accessible. He’s now applying that systems-level thinking to construction.He’s joined by Tanmaya Kala, former Project Executive at DPR Construction, who led large commercial, healthcare, and life sciences projects. We combine deep tech scale with real construction execution. What This Role Really Is Code writing is a commodity. We’re not hiring someone to close tickets. We’re hiring someone who: Is biased toward building Thinks in systems Designs for scale and resilience Reviews architecture, not just PRs Can move fast without creating long-term mess You’ll ship products daily. You’ll also help shape the system so it doesn’t collapse under growth. What You’ll Do Build end-to-end features (DB → API → UI) Design scalable data models and backend architecture Ship clean, production-grade frontend experiences Review and improve system design decisions Work directly with founders on product direction Help define engineering standards as we scale What We’re Looking For Strong backend fundamentals (APIs, databases, performance, system desi

reactgitai
View job →

Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team... Here at GoDaddy, the ML Engineering (MLE) team exists as the backbone of our machine learning infrastructure, enabling ML scientists and product teams across Domains to ship models to production reliably, efficiently, and at scale. This team owns the full lifecycle of ML systems — from CI/CD pipelines and model serving infrastructure to GPU workload orchestration and observability. Through disciplined engineering practices, thoughtful system design, and close collaboration with ML scientists, data engineers, and product teams, we deliver the platform that powers domain search, pricing, recommendations, and emerging AI experiences for millions of customers worldwide. We are currently looking for an experienced, highly motivated Senior Engineering Manager to lead our ML Engineering team based in India. This is an established team with existing engineers — we expect the candidate to ramp up quickly on our ML infrastructure stack, build strong relationships with the team, and partner with both India-based teams and US-based teams to drive execution and grow the team further. This individual will join us on our journey to build and scale ML infrastructure that serves real-time predictions at low latency, automates model deployment and promotion, and provides the observability and reliability guarantees that production ML systems demand. Become part of a team that bridges the gap between ML research and production engineering — shipping systems that directly impact GoDaddy's core revenue. What you'll get to do... Lead a team o

typescriptpythonaws
View job →

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What you will be doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuous

pythondockerkubernetes
View job →
H
1mo ago

Become a part of our caring community Most AI engineering jobs are a thin wrapper around a model API. This role is different. We build the platform that transforms millions of clinical documents into trusted, actionable data. Our systems use large language models (LLMs) to read medical records, extract structured facts, answer complex questions with citations back to the source document, and route ambiguous cases to human experts for review. Our users make decisions that impact real healthcare outcomes, so “good enough” is not good enough. Building AI systems that are accurate, reliable, auditable, and scalable is at the core of this role. As a Senior AI Applied Engineer, you will design, build, deploy, and operate production AI systems used at scale within one of the largest health insurers in the United States. You will own solutions end-to-end, from user experience and APIs to model orchestration, evaluation frameworks, infrastructure, and production operations. Why Join Us Build production AI systems where LLMs are in the critical path, not just demos or proofs of concept. Work on extraction, retrieval, agentic workflows, and human-review systems that process real healthcare data at scale. Own projects end-to-end across frontend, backend, AI orchestration, infrastructure, deployment, and operations. Solve challenging problems around accuracy, explainability, traceability, and reliability in regulated environments. Ship quickly in a small, high-impact team that embraces AI-assisted development and rigorous quality standards. Build systems that continuously improve through expert feedback, evaluations, and human-in-the-loop workflows. Key Responsibilities Design, develop, and deploy full-stack AI-powered application

javascripttypescriptpython
View job →

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What You Will Be Doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuously improve reliability and efficiency. Build and maintain runbooks, standard operating procedures, and comprehensive documentation for storage services and automation.<

pythondockerkubernetes
View job →

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir build software at scale to transform how organizations use data. Our Software Engineers are involved throughout the product lifecycle, from idea generation, design, prototyping, and production delivery. You will collaborate closely with technical and non-technical teammates to understand our customers' problems and build products that solve them. We encourage movement across teams to share context, skills, and experience, so you'll learn about many different technologies and aspects of each product. Engineers work autonomously and make decisions independently, within a community that will support and challenge you as you grow and develop, becoming a strong technical contributor and engineering leader. Our Product Development organization is made up of small teams of Software Engineers. Each team focuses on a specific aspect of a product. Our infrastructure teams are responsible for the lowest layers of our software stack, often focused on database technologies, distributed systems, large scale data systems, security, and application infrastructure. As a Software Engineer on infrastructure working on our Foundry platform, you'll contribute high-quality code to underpin Palantir Foundry and Gotham with performant, secure, and scalable building blocks, enabling products deployed to the most important institutions in the public and private sector. You'll build the foundational capabilities that power our products used by research scientists, aerospace engineers, intelligence analysts, and economic forecasters, in countries around the world. We’re hiring engineers who are passionate about solving real-world problems and empowerin

PE
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir build software at scale to transform how organizations use data. Our Software Engineers are involved throughout the product lifecycle, from idea generation, design, prototyping, and production delivery. You will collaborate closely with technical and non-technical teammates to understand our customers' problems and build products that solve them. We encourage movement across teams to share context, skills, and experience, so you'll learn about many different technologies and aspects of each product. Engineers work autonomously and make decisions independently, within a community that will support and challenge you as you grow and develop, becoming a strong technical contributor and engineering leader. Our Product Development organization is made up of small teams of Software Engineers. Each team focuses on a specific aspect of a product. Our infrastructure teams are responsible for the lowest layers of our software stack, often focused on database technologies, distributed systems, large scale data systems, security, and application infrastructure. As a Software Engineer on infrastructure working on our Foundry platform, you'll contribute high-quality code to underpin Palantir Foundry and Gotham with performant, secure, and scalable building blocks, enabling products deployed to the most important institutions in the public and private sector. You'll build the foundational capabilities that power our products used by research scientists, aerospace engineers, intelligence analysts, and economic forecasters, in countries around the world. We’re hiring engineers who are passionate about solving real-world problems and empowerin

G
Godaddy
📍 Bulgaria• Full-time
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Our Global Sustaining Engineering team sits at the intersection of software engineering and infrastructure, ensuring the services our customers depend on are fast, resilient, and always available. As a Senior Site Reliability Engineer, you'll take direct ownership of production services — from initial design through day-to-day operation — while partnering with product, engineering, and security teams to build and maintain business-critical systems. In this role, you will deepen your technical expertise and grow your leadership presence by mentoring the next generation of SREs. You will also gain hands-on experience with intelligent tooling in real-world workflows. What you'll get to do... Design, implement, and operate scalable, highly available production services while diagnosing and resolving complex infrastructure, network, and application issues Build and maintain alerting pipelines, dashboards, and SLO-driven monitoring strategies using Icinga, Prometheus, and Grafana Lead incident response end-to-end — performing root-cause analysis, authoring blameless post-mortems, and driving corrective actions to closure Develop and extend Infrastructure as Code coverage and build internal tooling that eliminates manual, repetitive operational work Mentor SRE I and SRE II engineers through code reviews, debugging sessions, and knowledge-sharing talks Apply LLM-driven log analysis, anomaly detection, and generative AI tools to accelerate incident response and runbook creation — validating all outputs before use Your experien

pythondockerkubernetes
View job →
D
Datadog
📍 New York• Full-time• From $116K/yr
1mo ago

Our GTM Enablement team plays a critical role in helping Datadog’s Sales and Customer Success teams succeed by equipping them with the right skills, knowledge, and tools at the right time. We partner closely with teams across Sales, Customer Success, Product Marketing, Strategy & Operations and more to ensure enablement programs translate into real performance in the field. The Program Management team within Sales Enablement is a group of interdisciplinary thinkers leveraging diverse business backgrounds in engineering, education, management consulting, business operations (and more!) to define, manage and measure Datadog’s global go to market motion. We solve complex, systemic challenges that impact go-to-market effectiveness by diagnosing root causes, whether behavioral, skills-based, or structural, and shaping the strategic solutions that resolve them. We partner closely with our Sales Enablement stakeholders across Curriculum, Field Enablement, Operations, and more to deliver on measurable results and operate as one unified enablement team. As a Senior Program Manager on the Sales Enablement Program Management team, you will design, plan, and execute complex, global, cross-functional programs that change how our Go-to-Market (GTM) teams operate at a systems level. You will own the end-to-end enablement process for the programs and stakeholder relationships you’re responsible for: diagnosing the gap, gathering field evidence, building a multi-faceted strategy (not just a content solution), partnering across teams to deliver the experience, and reporting business impact to Sales and Customer Success leadership. This is a high-ownership role for someone who moves quickly, thinks strategically and creatively, and is equally comfortable analyzing hard data and reading what’s happening on the ground. You’ll drive adoption at scale with limited guidance, influence senior stakeholders without positional authority, and help shape how the Program Management function o

aigorust
View job →
M
Mongodb
📍 New York• Full-time• From $164K/yr
1mo ago

Our Internal Data team is on a mission to power decisions and applications across MongoDB, oriented around the business areas of GTM, Product & Technology, Finance and HR. With this role, we are focusing on delivery of scalable, insightful, and impactful data products for our Product and Technology (P&T) teams - the org that builds MongoDB products for our external customers. These products drive smarter decisions, automate workflows, enable AI agents and deliver actionable insights that accelerate growth. If shaping the future of data-as-a-product excites you, then this opportunity is for you. As a Senior Data Product Manager on our Data team — which includes Data Pipeline/Platform Engineering, Data Architecture, and Data Governance — you will lead efforts to build, enhance, and scale the data products MongoDB employees use every day across a growing, high-impact company. You’ll deliver data products like data pipelines, datasets, reports/dashboards, APIs, ML models, playbooks, and frameworks while enhancing more traditional data management tools/services to power both AI and BI. These data products form the backbone of our internal data ecosystem and are instrumental in empowering teams to achieve their KPIs and innovate faster. In this role you will also have the opportunity to drive the vision and strategy for internal data as a product, with a focus on telemetry and product adoption/engagement/churn signals. By developing tools and systems that provide deep insights into how customers use or don’t use MongoDB’s products, you will empower teams to understand customer behaviors and improve the end-user experience. The goal is to enable data-driven decisions that enhance product design, drive customer value, and fuel innovation across the MongoDB portfolio. This role can be based out of our Palo Alto, San Francisco, or New York City office or remotely in the East Coast region. Responsibilities Lead Product Strategy and Ownership Define product vision, stra

mongodbawsazure
View job →
L
Lyft
📍 Toronto• Full-time• C$149.6K – C$187K/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. With a billion rides per year and counting, Lyft is solving hard problems in a rapidly growing domain with a lot of data and creative solutions in Rider, Marketplace, Growth, and beyond. While traditional approaches to optimization and problem decomposition are sufficient to disrupt transportation, building a next-generation platform for low-cost, ultra-immersive transportation to improve people's lives warrants modern ML utilizing peta-byte scale data. Our highly motivated Machine Learning Engineers work on these challenging problems and define solutions to directly impact various aspects of our core business. If you are a critical thinker with experience in machine learning workflows and LLMs, passionate about solving business problems using data and working in a dynamic, creative, and collaborative environment, we are searching for you. We are seeking a Senior Machine Learning Engineer to join the Rider Applied AI team and lead the design, development, and deployment of state-of-the-art machine learning and artificial intelligence systems. This role requires a strategic thinker who can balance high-level system architecture with hands-on technical implementation. You will collaborate across teams to shape the future of ride-sharing by leveraging AI, Machine learning and Data science. Responsibilities: Model Development & Research: Design, build, and deploy machine learning models for real-time applications, including translating state-of-the-art research into production-ready solutions. System Design: Architect scalable, reliable ML pipelines that integrate seamlessly with existing backend systems. Innovation & Applied Research: Stay ahead of the curve by exploring emerging algorithms, technologies (such as LLMs and LLM-based applications), and frameworks — critically eva

pythonmachine learningai
View job →
L
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. With a billion rides per year and counting, Lyft is solving hard problems in a rapidly growing domain with a lot of data and creative solutions in Rider, Marketplace, Growth, and beyond. While traditional approaches to optimization and problem decomposition are sufficient to disrupt transportation, building a next-generation platform for low-cost, ultra-immersive transportation to improve people's lives warrants modern ML utilizing peta-byte scale data. Our highly motivated Machine Learning Engineers work on these challenging problems and define solutions to directly impact various aspects of our core business. If you are a critical thinker with experience in machine learning workflows and LLMs, passionate about solving business problems using data and working in a dynamic, creative, and collaborative environment, we are searching for you. We are seeking a Senior Machine Learning Engineer to join the Rider Applied AI team and lead the design, development, and deployment of state-of-the-art machine learning and artificial intelligence systems. This role requires a strategic thinker who can balance high-level system architecture with hands-on technical implementation. You will collaborate across teams to shape the future of ride-sharing by leveraging AI, Machine learning and Data science. Responsibilities: Model Development & Research: Design, build, and deploy machine learning models for real-time applications, including translating state-of-the-art research into production-ready solutions. System Design: Architect scalable, reliable ML pipelines that integrate seamlessly with existing backend systems. Innovation & Applied Research: Stay ahead of the curve by exploring emerging algorithms, technologies (such as LLMs and LLM-based applications), and frameworks — critically eva

pythonmachine learningai
View job →

About the Team At OpenAI, our User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We translate real-world user and operational signals into timely decisions, practical interventions, and improvements to our products and systems. This role will take on new, ambiguous, or underdeveloped operational risks and help mature them into scalable capabilities. We work across USRO and partner closely with Product, Engineering, Data Science, Product Policy, Legal, Safety, Support, and external vendors or partnership stakeholders. About the Role We are seeking a Senior Operations Analyst to take on complex, ambiguous safety and risk problems and turn them into practical operational solutions that can scale. This is a senior individual-contributor role for a versatile operator who is comfortable moving between queues, investigation, analysis, workflow design, hands-on execution, and cross-functional leadership. Depending on team needs, the role may focus on emerging-risk incubation, cloud deployment partnerships, or other new operational areas. You will be expected to move quickly, work hands-on, and create structure without waiting for perfect requirements or a large support team. The work starts with the problem, not a prescribed process. You may investigate unstructured user signals, stand up a lightweight workflow, build an AI-assisted tool, improve an existing operation, or help a new launch become operationally ready. The goal is to produce durable systems that other people can run, not simply complete a series of individual tasks. The portfolio will change with company priorities and may span established harm areas, emerging-risk incubation, cloud deployments and partnerships, device safety, or new product launches. Some hires may focus primarily on cloud deployment operations, including launch readiness, partner coordination, safety workflows, and operational monitoring. You will ty

sqlawsrest
View job →

Location Details: Pune, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a hybrid position. You’ll divide your time between working remotely from your home and an office, so you should live within commuting distance. Hybrid teams may work in-office as much as a few times a week or as little as once a month or quarter, as decided by leadership. The hiring manager can share more about what hybrid work might look like for this team. Join our Team Our team builds and operates the foundational infrastructure platforms that power GoDaddy's engineering organization. We own critical services including secrets management, software distribution, host security controls, and live patching for thousands of Linux systems running on OpenStack. This role sits at the intersection of Linux engineering, platform engineering, reliability engineering, and security. You will help define how core infrastructure services are designed, operated, automated, and scaled across the enterprise! What you'll get to do... Design, build, and operate highly available, scalable, and secure infrastructure platforms supporting large-scale Linux environments, with a focus on reliability, resiliency, and operational efficiency Lead the architecture, implementation, and operation of infrastructure services, including OpenStack, enterprise secrets management, package management, software promotion pipelines, and platform lifecycle management Develop and maintain automation solutions using infrastructure-as-code, Ansible, Python, Go, and self-service capabilities to improve efficiency and reduce operational overhead Build and improve observability and reliability practices through monitoring, logging, alerting, dashboards, managing incidents, analyzing underlying causes, disaster recovery, and service health reporting

pythongitlinux
View job →
🔔

Get new senior systems design and integration specialist jobs by email

Daily job updates · Unsubscribe anytime