Jobiba hiring network

Senior Systems Design And Integration Specialist Jobs

7,101 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior systems design and integration specialist jobs. Use filters to narrow by work mode, employment type, experience and date posted.

I
Instacart
📍 BC, Canada• Full-time• Remote• From C$168K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview About the Role We are currently seeking a Senior Software Engineer to join our Agentic Analytics Platform team — the team responsible for the AI-for-Data charter inside Instacart's Data Infrastructure org. You'll design and build LLM-powered systems that transform how data practitioners (data scientists, data engineers, analysts, PMs) interact with data at Instacart — from natural-language data access and AI-assisted SQL, to automated metadata generation, to embedding intelligent capabilities across our broader data infra ecosystem. This is a hands-on role at the frontier of applied AI inside a large, modern data stack. About the Team The mission of the Instacart Self-Serve organization is to improve the productivity of data practitioners through easy-to-use, self-serve tools. Agentic Analytics is the team chartered with bringing AI and LLMs into that miss

REMOTEpythonsqlai
View job →
P
Pagerduty
📍 Lisbon• Full-time
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. About the role PagerDuty’s Operations Cloud runs on a platform that ingests billions of signals and turns them into real-time action for thousands of customers. We’re looking for a Senior AI/ML Engineer who lives at the intersection of two disciplines: large-scale distributed systems and applied AI. In this role you will design and ship AI systems that run in production at PagerDuty’s scale — powering Incident Management AI Agents, event intelligence, and the LLM-powered capabilities embedded across our platform. You’ll own the full lifecycle, from framing the problem to serving reliably at scale. We are looking for a candidate who is genuinely passionate about building with modern AI — LLMs, agents, and retrieval — but grounded in the realities of building resilient, high-throughput systems. What you’ll do Design and build AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time event streams, from problem framing through production deployment and monitoring. Architect and own the systems behind them: agent and prompt orchestration, retrieval pipelin

awsazuregcp
View job →
LA
18 days ago

ROLE SUMMARY We are looking for a Senior Data Scientist to lead complex data science engagements that combine traditional statistical modelling with Generative AI. You will work hands-on with very large datasets across disparate systems and formats, translate ambiguous business problems into rigorous analytical solutions, and present those solutions clearly to C-level stakeholders. This is a delivery-first role with a fast track into technical leadership: alongside your own project work, you will help guide junior data scientists and shape how Lynx builds and ships data science solutions. KEY RESPONSIBILITIES Solution Design & Delivery Design and deliver end-to-end solutions for defined data science problems, combining classical modelling, data transformation, and Generative AI / LLM techniques. Work hands-on with very large datasets across disparate stores and formats, from ingestion and transformation through to modelling and validation. Apply statistical and machine learning methods to business problems such as customer retention, campaign management, and commercial performance optimisation. Client Communication & Leadership Present results and prepare client-ready materials for project stakeholders, including C-level audiences, translating technical work into clear business narratives. Lead smaller data science workstreams, with support from internal leadership and the PMO, including day-to-day guidance for junior team members. Partner with delivery and account teams to scope problems, set realistic timelines, and manage stakeholder expectations. Knowledge Building Create reusable documentation, presentations, and code libraries during projects so future engagements can build on prior work. Participate in internal education, research, and knowledge-sharing initiatives that raise the technical bar across the practice. SKILLS, QUALIFICATIONS AND EXPERIENCE 8+ years of overall experience in data science, with a track record of leading analytic

pythonsqlaws
View job →

What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.

I
Instacart
📍 United States• Full-time• Remote• From $240K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview As a Senior Machine Learning Engineer II on the Ads Response Prediction team, you will lead the design and development of core ML models that power Instacart’s ads ecosystem. This is a research-leaning role focused on theoretical problem formulation, training methodology, and model quality rather than infrastructure or full-stack engineering. You will tackle fundamental challenges in pCTR modeling such as mitigating selection bias, position bias, and optimizer’s curse in training data, improving model calibration across surfaces and domains, and advancing our multi-task learning and sequence modeling capabilities. You will also have the opportunity to shape our next-generation foundation model approach for ads ranking and contribute to cutting-edge retrieval systems like TIGER (Transformer Index for Generative Recommenders), Semantic ID and domain language

REMOTEpythonsqlmachine learning
View job →
A
Asana
📍 Warsaw• Full-time• $372K – $432K/yr
1mo ago

We're looking for a Senior Infrastructure Engineer who brings strong software engineering skills and a deep understanding of production systems. This role is a good fit for someone who enjoys building systems that make infrastructure more scalable, reliable, and easy to operate – using code, not runbooks. You'll work with a highly collaborative team to design and build the internal platforms that power all of Asana, from product features to AI systems to offline analytics. Our tech stack includes: AWS, Kubernetes (EKS), MySQL (RDS), OpenSearch, DynamoDB, Redis, Terraform, Datadog, TypeScript, Scala, Go, and Python. We’re especially interested in people who think like backend engineers but care deeply about systems – things like failure modes, operational cost, debuggability, and performance. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you’ll achieve: Design and build frameworks, tools, and services that improve the reliability, observability, and scalability of Asana’s infrastructure. Lead end-to-end projects, from scoping and design through to rollout, across multiple systems and teams. Improve the operability of stateful infrastructure like MySQL, OpenSearch, and DynamoDB – and help drive Asana’s long-term vision for storage reliability. Debug production issues across the stack. Yes, there’s an on-call rotation – but this isn’t a pager monkey role. You’re here to fix things properly and make sure they don’t break again. Partner with product teams to shape a service-oriented architecture that enables fast, reliable development. Share knowledge through code reviews, design discussions, and mentorship. Abou

typescriptpythonsql
View job →
D
1mo ago

This role will join Datadog’s Data Visualization organization, a team responsible for the visualization experiences that power dashboards, notebooks, investigations, and product workflows used across the platform. The team is a highly product-oriented organization, building AI-native experiences that help customers understand, investigate, and interact with complex operational data. As a Staff Software Engineer, you will provide technical leadership in applying AI technologies to customer-facing product experiences, helping shape how users interact with Datadog through agents, conversational interfaces, and intelligent investigation workflows. You will partner across engineering and product teams to develop reliable, scalable, and trustworthy AI-powered experiences while helping establish AI engineering expertise within the broader organization. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the design and delivery of AI-powered product experiences across Datadog’s visualization and investigation surfaces. Develop systems that combine deterministic product capabilities with LLM-powered experiences to deliver trustworthy and explainable customer outcomes. Drive innovation in context engineering, prompt engineering, evaluation frameworks, and AI application reliability. Partner with product and engineering teams to improve investigation workflows and help customers discover insights more efficiently. Build experiences that enable Datadog capabilities to operate within third-party AI platforms, agents, and conversational environments. Provide technical leadership and mentorship while helping establish AI engineering best practices across the Data Visualization organization and broader Graphing group. Who You Are: You have extensive softw

aigorust
View job →
M
Mongodb
📍 San Francisco• Full-time• From $126K/yr
1mo ago

Join the Atlas Search team to design and develop the next generation of Semantic and Vector Search infrastructure. Atlas Search is a growing cloud service that allows users to execute complex search and vector search queries using the MongoDB Query Language. Our users can focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team is building a cloud-based distributed system responsible for the core components of search including data ingestion, performance, query language, query execution, for both relevance-based search and vector search. Our product is being adopted quickly and there are many interesting projects. This is a technical role where you will be responsible for the infrastructure and features enabling our at-scale cloud service powering vector and semantic search. We are looking to speak to candidates who are based in the San Francisco Bay Area for our hybrid working model. What You’ll Do Lead complex projects across the MongoDB ecosystem, for instance, development of a new Search deployment framework within the MongoDB managed cloud Set project level strategy, architect features, and lead projects to successful execution Identify, design, and implement features enhancing our reliability, performance, security and efficiency Perform code reviews with peers and make recommendations on how to improve our software development processes Influence and grow team members through active mentoring and leading by example What We Look For 5+ years experience in data management/search systems, ideally with a strong distributed systems and infrastructure background Experienced in the development and maintenance of concurrent, stateful services Eager to shape the technological direction of a complex system and have the ability to lead initiatives through collaboration with others Experienced in writing features, debugging and optimizing multithreaded applications written in Java Familiarity with LLM

javamongodbaws
View job →

Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability. Key Responsibilities MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and

pythonawsazure
View job →
M
Mongodb
📍 San Francisco• Full-time• From $126K/yr
1mo ago

Join the Atlas Search Query team to design and develop the next generation of Search query architecture, optimization, and execution. Atlas Search is a growing cloud service that allows users to execute complex search and vector search queries using the MongoDB Query Language. Our users can focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team is building a cloud-based distributed system responsible for the core components of search including data ingestion, performance, query language, query execution, for both relevance-based search and vector search. Our product is being adopted quickly and there are many interesting projects. This is a technical role where you will be responsible for the success of complex Search Query feature development. We are looking to speak to candidates who are based in San Francisco, CA for our hybrid working model. What You’ll Do Lead complex projects across the MongoDB ecosystem, for instance, development of a new Search aggregation framework within the MongoDB aggregation framework. Set project level strategy, architect features, and lead projects to successful execution Identify, design, and implement features enhancing our query language, performance, and operability Perform code reviews with peers and make recommendations on how to improve our software development processes Influence and grow team members through active mentoring and leading by example What We Look For 5+ years experience in data management/search systems, ideally with a strong query processing and optimization background Experienced in the development and maintenance of stateful distributed systems Eager to shape the technological direction of a complex system and have the ability to lead initiatives through collaboration with others Experienced in debugging and profiling multithreaded applications written in Java and Rust Bonus: experience with designing high-volume query engines, such as a datab

javamongodbaws
View job →
P
17 days ago

About Us Pearl is AI for professional services at global scale, combining advanced AI with verified human expertise to deliver help that is accurate, accountable, and fast. Since 2003, our network has connected millions of customers with licensed professionals across 196 countries, making real expertise available anytime, anywhere. Our Values Data driven: Start with truth, measure what matters. Courageous: Bias to action; run toward hard problems. Innovative: Seek novel, elegant solutions. Lean: Do more with less. Build, ship, learn fast. Humble: Strong opinions, lightly held. About the Role The future of analytics isn't dashboards. It's intelligent systems that anticipate questions, surface insights, and help people make better decisions. We're looking for a Senior Analytics Engineer to help build that future at Pearl. In this role, you'll design and develop AI-powered analytics experiences that combine trusted enterprise data with modern AI capabilities, enabling business users to interact with data conversationally and uncover insights faster than ever before. You'll build production AI agents, create scalable semantic data models, develop intelligent analytics applications, and establish best practices for responsible AI across the Analytics organization. Working closely with Product, Engineering, and business leaders, you'll turn emerging AI technologies into real business capabilities that improve decision making across the company. This is an opportunity to help define how AI transforms analytics at Pearl while working on some of the most exciting technologies in data, LLMs, and agentic AI. What You’ll Do Build proactive analytics and AI solutions that surface actionable business insights, anticipate business needs, and enable smarter decision-making. Design and develop AI-powered analytics tools, including conversational interfaces that allow business users to query data using natural language. Build semantic data models and reusab

pythonsqlazure
View job →
JT
1mo ago

Senior Consultant-We are seeking a seasoned ITES Solution Architect to lead the design and execution of enterprise-grade IT infrastructure projects. -"Key Responsibilities: Infrastructure Architecture & Design • Design and implement end-to-end IT infrastructure solutions including compute, storage, network, and security. • Architect high-availability systems with failover clustering, load balancing, and redundancy. • Define and implement disaster recovery strategies with clear RTO/RPO objectives. Data Center & DR Strategy • Design and manage primary and secondary data center environments. • Plan and execute DR site configurations, ensuring seamless failover and recovery. • Implement real-time or scheduled data synchronization between DC and DR using replication technologies (e.g., SAN replication, DFS-R, Veeam, Zerto). Backup & Recovery Planning • Develop and maintain enterprise backup strategies using tools like Commvault, Veeam, or NetBackup. • Ensure secure, encrypted backups with retentio

Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. Location- Chennai Team: Engineering Enablement Group As a Senior Software Engineer in our Engineering Enablement Group, you will lead the re-design and evolution of our Mobile Branding framework — the system that enables customers to create custom-branded versions of the Appian mobile application for both iOS and Android. You will drive the architectural modernization of the end-to-end branding pipeline, from the customer-facing Forum application and provisioning tools to the backend build service running on Mac EC2 runners in AWS. By leveraging modern microservices, CI/CD automation, and cloud-native infrastructure, you will transform the current system into a more reliable, scalable, and maintainable platform that reduces manual intervention and accelerates customer delivery. We are looking for a technical leader who can bridge the gap between complex Ruby/Bash-based tooling, Appian process models, and AWS infrastructure to deliver a seamless mobile branding experience. Primary Qualifications: 6-9 Strong working experience with Android and iOS frameworks and mobile application development workflows. Familiarity with mobile build systems (Fastlane, Xcode, Gradle) and code-signing workflows. Experience with proficiency in Python, with experience in Ruby, Bash, or Go being a plus. Advanced experience with AWS infrastructure (S3, Lambda, EC2) and CI/CD pipeline design. Strong end-to-end knowledge of pipeline creation, deployment automation, and infrastructure-as-code (Terraform). Familiarity with monitoring, observability, and performanc

pythonawsci/cd
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

REMOTEkubernetesaigo
View job →

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with 5–8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role—leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet. What You'll Do Architectural Leadership: Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management. Platform Orchestration: Lead complex, large-scale infrastructure shifts. Operational Excellence: Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies. This Role Requires Experience: 5–8 years in a DevOps, Site Reliability, or Infrastructure Engineering role. Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators. Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation. Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through auto

awsazurekubernetes
View job →
🔔

Get new senior systems design and integration specialist jobs by email

Daily job updates · Unsubscribe anytime