About the Team The ChatGPT organization at OpenAI supports our mission by bringing advanced AI capabilities to hundreds of millions of users worldwide. The Image Generation team is responsible for one of the fastest-growing experiences in ChatGPT, enabling users to create, edit, and transform images through natural language. Recent advances in our multimodal image models have dramatically improved image quality, instruction following, editing precision, consistency, and text rendering, unlocking entirely new creative and professional workflows. We work at the intersection of research, infrastructure, and product to build the systems that power image generation at global scale. Our team partners closely with researchers, product engineers, designers, and platform teams to bring state-of-the-art image capabilities to millions of users while continuously pushing the boundaries of what AI-powered creation can do. About the Role We are looking for an experienced Backend Engineer to join the Image Generation team and help build the systems that power image creation and editing across ChatGPT. You'll work on the core backend infrastructure that enables users to generate, edit, and iterate on visual content using cutting-edge multimodal AI models. This includes building highly scalable services, orchestration systems, APIs, storage platforms, and distributed infrastructure that support billions of image generations and editing workflows. You'll partner closely with product, research, and mobile teams to transform breakthrough AI capabilities into reliable, performant experiences used by millions around the world. In this role, you will: Design, build, and operate backend systems that power image generation and image editing experiences in ChatGPT. Develop scalable APIs, services, and infrastructure that support multimodal AI workflows. Optimize reliability, latency, throughput, and cost across large-scale distributed systems. Partner with researchers to productionize new im
Jobiba hiring network
Distributed Systems Engineer Data Platform Delivery Database Retrieval Jobs
1,301 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current distributed systems engineer data platform delivery database retrieval jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Contribute to the development of GoDaddy’s eCommerce and SSO infrastructure and Kubernetes systems on AWS. On a day-to-day basis you will be working on the team who designs, writes, tests and deploys the infrastructure and application management software for GoDaddy’s eCommerce applications. Expect to learn every day. What you'll get to do... Work as a polyglot engineer, writing and maintaining Infrastructure as code with frameworks/ ecosystems such as Java, Unix CLI, and NodeJS Build and operate infrastructure workflows and deployment pipelines using Kubernetes, Argo Workflows, Argo CD, and GitOps practices Design, build, and own services and APIs in Java, running on Kubernetes-based platforms across AWS and distributed systems Develop and support application and infrastructure delivery pipelines, enabling reliable releases of eComm, Auth and Infrastructure services Collaborate closely with other GoDaddy departments to help advance security and technical standards, maintain regulatory compliances while operating eComm & Auth platforms Your experience should include... 5+ years of strong backend software engineering experience in Java Hands-on experience with Kubernetes, including Helm, Kustomize, or equivalent tools to deploy and manage backend services Experience building and operating high-volume, mission-critical production systems on AWS with continuous deployment (CD) practices Strong experience with infrastructure as code, supporting backend applications and services Experience with observability and l
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: As a New Grad Software Engineer, you'll join a team of exceptional builders working on products that are reshaping how the world creates software. You'll have the opportunity to work on everything from our AI-powered development platform to the distributed systems that enable real-time collaboration for millions of developers. This is a chance to define your career while defining the future of software development. You'll work on problems that matter, with the autonomy to drive solutions and the support to grow into a technical leader. What you will build: Product features that delight users and make it possible for anybody to create software AI coding agent that understands intent and generates production-ready applications Cloud infrastructure that provides instant, powerful development environments at global scale Platform features that enable one click deployments and scale to millions of users Required skills and experience: Recent graduate (2027) with a degree in Computer Science, Computer Engineering, or related field Strong programming skills in a modern language (JavaScript/TypeScript, Python, Go, Rust) Full-stack capabilities with experience in React, Node.js, and database technologies Growth orientation - eager to learn new technologies and take on increasing responsibility Collaborative spirit - you work well in cross-functional teams and value diverse perspectives What we value : Problem-solving mindset: Ability to approach complex operational challenges systematically and devise effective solutions Self-directed and autonomous: Capable of working independently while collaborating effectively with cross-functional teams Strong communication skills: Ability to explain complex technical conce
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are looking for an experienced Principal Software Engineer to work on our next-generation Imports Platform team. Imports Platform team is leading a strategic initiative to modernize Okta's identity lifecycle management capabilities by architecting and migrating from a legacy monolithic system to a highly scalable, distributed microservices platform. This critical service orchestrates the importing, syncing, and provisioning of identities and access policies—users, groups, roles, entitlements—from external directory services including Active Directory, Office 365, and LDAP-based systems. As a Principal Software Engineer on the Imports Platform team, you will be a cross-team technical leader who takes difficult, ambiguously defined problems and drives them from ideation through production impact without oversight. You will own projects from zero to landing—defining scope, planning execution, making architectural decisions, and articulating measurable impact across the group. You will generate novel solutions to complex distributed systems challenges, guide the team's technical direction, and get stakeholder buy-in on architectural strategy spanning multiple teams. Your sphere of influence extends beyond the Imports Platform team to adjacent teams within the group and cross-functional partners in Product, Design, and SRE. You will participate in group-level strategy, break down strategic initiatives into actionable technical milestones, and drive cross-team
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Staff Software Engineer on the Core Automation team within the Platform group, you'll architect and build the Agentic AI systems that are transforming how Coinbase operates. This team is reimagining customer support and compliance processes for a fully AI-driven world, designing intelligent agents, orchestration frameworks, and measurement systems that deliver delightful customer experiences at scale. You'll own the technical direction for production AI systems, working across cross-functional teams to bring this vision to reality while building primitives that scale automation across the company. What you'll do: Architect and build Agentic AI systems that power Coinbase's compliance automation and other Operations, from intelligent agents through orchestration and guardrails Design foundational APIs and measurement frameworks that ensure AI agents are grounded, relevant, and reliably deliver customer delight with minimal hallucination Lead technical direction for distributed systems underpinning AI automation, defining architecture patterns and strategic roadmaps in partnership with engineering leadership Build reusable primitives and orchestration solutions that enable AI-powered automation to scale across multiple domains beyond the initial customer support and compliance focus Mentor engineers on AI system design techniques, coding standards, and production-
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the role. We’re building the next generation of agentic AI infrastructure at Cohere. This team sits at the intersection of ML systems, distributed infrastructure, and developer experience, creating the platform that powers autonomous AI agents at scale. You’ll work on hard, forward-looking problems with few established patterns, including secure code execution, agent state management, model routing, identity and authentication, and resource management for long-running agent workflows. This role is a strong fit for someone who combines systems depth with ML intuition. You should be comfortable building reliable infrastructure, thinking through distributed systems tradeoffs, and understanding how emerging agentic capabilities shape platform design. What you’ll work on. Secure execution environments for agent-generated code Identity, authentication, and trust boundaries for agents Model routing and orchestration across different model types and environments Rate limiting, quotas, and resource management for agent workflows State management, memory, and filesystem abstractions for agents. In this role you will: Turn emerging M
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: Join a team that builds robust, real-time distributed systems for a cutting-edge database. We care about performance, reliability, scalability, and most of all learning and having fun together. Whether you’re a seasoned coder or just getting started, if you’re passionate about technology and eager to learn, you’ll fit right in. Who we are: We show up to work, ready to collaborate and build technologies that make a difference, with people who genuinely care. We chase improvements such as tail latencies, bytes throughput, cache hit rate, and operational cost efficiency. We believe learning is ongoing and that even the most complex problems can have simple solutions. What You’ll Do: Collaborate with teammates to design and build database features that power AI applications. Learn how to tune performance and support reliability in distributed systems (don’t worry, we’ll guide you). Help Pinecone run smoothly on popular cloud providers. Take ownership of your work and grow your skills every day. Have fun. Who You Are: 5+ years of work experience - programming in Rust, Go, C++, or a comparable language. You’re genuinely curious about distributed systems and eager to dive deep into technical challenges. You approach problems with creativity and persistence, and you’re comfortable asking thoughtful questions or seeking feedback. You’re excited to learn, value constructive feedback, and appreciate mentorship. Bonus Points: You have hands-on experience with cloud platforms (AWS, GCP, Azure) or have demonstrated an ability to pick u
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know the team The Developer Platform team at Auth0 (an Okta company) owns the platform that developers build identity on. Increasingly they build alongside AI agents, and that raises the bar on everything underneath: interfaces have to hold up whether a person or an agent is calling them, and the systems behind them have to stay reliable and coherent as usage grows. We move fast, we own problems end to end, and we care deeply about the platform we put in front of the developers and agents who depend on it. The opportunity We're hiring a Principal Engineer (P5) to serve as the technical leader and compass for the Developer Platform team. You'll work across the breadth of the platform, tackling the highly complex, vaguely specified problems that span it and turning them into clear technical direction the team can execute against, without day-to-day oversight. Above all, you'll own how the platform is architected to scale: the distributed systems behind it, the reliability and consistency guarantees developers depend on, and the coherence that keeps it easy to build on as usage grows. You'll champion the team's technical execution, raise the engineering bar, mentor the people around you, and partner with tech leads across teams to keep the wider platform aligned. You'll have real influence over how our platform holds up in a world where developers and agents are both first-class consumers. What you'll be doing Own the platform architecture: Set the tech
Location Details: Canada, Remote At GoDaddy, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) , and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Contribute to the development of GoDaddy’s eCommerce and SSO infrastructure and Kubernetes systems on AWS. On a day-to-day basis you will be working on the team who designs, writes, tests and deploys the infrastructure and application management software for GoDaddy’s eCommerce applications. Expect to learn every day. What you'll get to do... Work as a polyglot engineer, writing and maintaining Infrastructure as code with frameworks/ ecosystems such as Java, Unix CLI, and NodeJS Build and operate infrastructure workflows and deployment pipelines using Kubernetes, Argo Workflows, Argo CD, and GitOps practices Design, build, and own services and APIs in Java, running on Kubernetes-based platforms across AWS and distributed systems Develop and support application and infrastructure delivery pipelines, enabling reliable releases of eComm, Auth and Infrastructure services Collaborate closely with other GoDaddy departments to help advance security and technical standards, maintain regulatory compliances while operating eComm & Auth platforms Your experience should include... 5+ years of strong backend software engineering experience in Java Hands-on experience with Kubernetes, including Helm, Kustomize, or equivalent tools to deploy and manage backend services Experience building and operating high-volume, mission-critical production systems on AWS with continuous deployment (CD) practices Strong experience with infrastructure as code, supporting backend applications and services Experience with observability a
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. We are hiring a Software Engineer I to join our Engineering team. You will get your hands on full-stack services, dig into distributed systems problems, and work right at the intersection of traditional software engineering and AI-driven development. We are not just dabbling in AI here. We use it to write better code, catch bugs faster, and troubleshoot smarter, and we want engineers who want to get good at that too and help their teammates do the same. You will ship real features, sit in on architectural conversations that actually shape the product, and grow alongside a team that cares as much about doing good work as they do about doing meaningful work. You Will: Build scalable front-end and back-end services for the next generation of applications at Smartsheet (Kotlin, Java, Typescript, React) Solve challenging distributed systems problems and work with modern cloud infrastructure (AWS, Kubernetes) Take part in code reviews and architectural discussions as you work with other software engineers and product managers Forge a strong partnership with product management and other key areas of the business Enhance existing application code with new features and strike a balance when making technical decisions (build vs refactor vs simplify) Actively use AI tools to improve personal and team efficiency across coding, testing, design, and troubleshooting, exploring AI integration opportunities within team processes and features, and coaching others on effective AI use You Have: 1+ years software development experience build
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. As a Software Engineer II at Smartsheet, you will help build the next generation of our platform while tackling meaningful distributed systems challenges alongside a talented engineering team. You will collaborate closely with product managers and fellow engineers through code reviews and architectural discussions, and you will have the opportunity to mentor junior engineers as you grow in your own career. This role is a strong fit for someone with a couple of years of hands-on development experience who is ready to take on more ownership, sharpen their skills in cloud infrastructure, and help shape how the team uses AI tools to work smarter and faster . You Will: Build scalable back-end services for the next generation of applications at Smartsheet (Kotlin, Java) Solve challenging distributed systems problems and work with modern cloud infrastructure (AWS, Kubernetes) Take part in code reviews and architectural discussions as you work with other software engineers and product managers Mentor junior engineers on code quality and other industry best practices Forge a strong partnership with product management and other key areas of the business Actively use AI tools to improve personal and team efficiency across coding, testing, design, and troubleshooting, exploring AI integration opportunities within team processes and features, and coaching others on effective AI use You Have: 2+ years software development experience building highly scalable, highly available applications 2+ years of programming experience with full st
This new team will architect the platform foundation and infrastructure that enables a "build once, run anywhere" model, ensuring our AI-powered modernisation suite operates seamlessly regardless of a client's security or network constraints. We are looking for engineers to join this high-visibility initiative, where you will solve unique distributed systems puzzles and help shape the future of how global enterprises leverage GenAI. We are looking for an experienced software engineer who thrives on solving infrastructure constraints and building enterprise facing platforms. The ideal candidate is a hands-on technical leader who can architect complex distributed systems, mentor engineers, and collaborate with product teams to deliver a platform that minimizes deployment friction and meets customers’ compliance requirements. This role can be based out of our Gurgaon office or Remotely in India. The ideal candidate for this role will have 8+ years of software development and operations experience, with a focus on building platforms and distributable software infrastructure 2+ years of experience leading, coaching, and mentoring a team of engineers to achieve high-impact results Deep experience designing distributable applications that run in air-gapped or highly restricted network environments Strong proficiency in containerization and orchestration, with the ability to design systems where the host executes containerized applications under strict security monitoring Experience building Service Host architectures that provide shared routing, proxy layers, and API gateways for multiple underlying services Experience managing persistent storage solutions (blob/file storage) within distributed systems Understanding of security-first design, specifically regarding the execution of untrusted code, fine-grain access control, and rigorous input validation Experience designing auto-update mechanisms for software that cannot directly access the public internet Curiosity,
We are seeking a Senior Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Senior Site Reliability Engineer on this new team, you will be responsible for enabling deployment at scale of AI applications and improving the performance, scalability, and reliability of the distributed systems infrastructure for this new product. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing Identify and configure key metrics to detect incidents and quantify service health, availability, and performance Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Mentor early-career SREs and contribute to the team’s operational practices as it grows Qualifications Strong background in software development and operating distributed systems 6+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language Experience operating Kubernetes in production and debugging below the abstraction layer, including scheduling, cluster networking, and node-level issues Expertise in cloud infrastructure platforms, in
Get new distributed systems engineer data platform delivery database retrieval jobs by email
Daily job updates · Unsubscribe anytime