Jobiba hiring network

Production Tech Jobs

3,233 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current production tech jobs. Use filters to narrow by work mode, employment type, experience and date posted.

P
Pendo
📍 Raleigh• Full-time• $105K – $130K/yr
1mo ago

The team + the role Pendo's Applied AI team turns AI infrastructure into real business outcomes across Sales, Marketing, and Customer Engineering. AI here isn't a feature we bolt on — it's how we scale GTM capability across the company. We build AI that makes GTM teams measurably faster and more effective, and we measure success by whether those teams are actually using what we ship and getting real value from it. As an Applied AI Engineer, you'll own the full lifecycle of AI-powered solutions — from problem definition and prompt design through to production deployment and ongoing iteration. You'll work directly with GTM stakeholders to identify high-value problems, set realistic expectations about what AI can and can't do, and ship solutions that stick. The best person for this role has a strong engineering instinct, deep curiosity about how businesses operate, and the judgment to know when AI is the right tool — and when it isn't. This role is based in Raleigh, NC and follows Pendo's hybrid model: in-office 3 days per week. What this looks like day-to-day Own AI solutions end-to-end: problem definition, prompt design, production deployment, monitoring, and iteration — you ship, you watch, you improve. Build and manage GTM workflow automations that reduce manual work across Sales, Marketing, and Customer Engineering systems, including Slackbots and other integrations. Partner directly with GTM stakeholders to surface high-value problems, validate solutions, drive adoption, and set honest expectations about AI capabilities and limitations. Develop and maintain prompt management practices that make AI outputs reliable, auditable, and improvable over time — treat prompts as production code, not experiments. Work with the Data Platform and Systems teams to identify foundational tooling gaps and contribute clear, actionable requirements based on what you encounter in production. Share reusable AI workflow patterns and tool findings with the broader team — your impact sh

S
Sendbird
📍 Seoul• Full-time
1mo ago

Sendbird is building AI agents for customer experience. Our platform already powers billions of conversations every month across chat, voice, video, and messaging APIs. We are now using that foundation to build agents that understand customer context, reason over business data, and take reliable action in production. We are looking for a Machine Learning Engineer to research, build, and productionize new capabilities for those agents. This role sits at the intersection of agent product development, applied AI research, and production engineering. You will work on systems that enterprise customers depend on every day, not demos or isolated prototypes. About Sendbird and delight.ai Sendbird has spent more than a decade building communication infrastructure for in-app chat, voice, video, and messaging APIs. More than 4,000 brands use our platform, including DoorDash, Match Group, Noom, Yahoo Sports, and Rakuten. Our systems support more than 7 billion messages every month. In 2024, we made a strategic shift toward AI-first customer experience. In 2025, we launched our enterprise AI agent product, delight.ai. Delight.ai helps businesses deliver customer support and engagement that is faster, more contextual, and more personal. Unlike simple FAQ bots, our agents are built to remember customer context, use tools, retrieve relevant knowledge, connect across channels, and handle real customer workflows with accuracy and control. The Role As a Machine Learning Engineer, you will design, build, evaluate, and ship new capabilities for our AI agents. You will work across agent architecture, retrieval, memory, planning, tool use, workflow automation, voice, evaluation, data pipelines, model adaptation, inference, and production integration. This is a hands-on engineering role for someone who can turn AI research and product ideas into reliable customer-facing features. Some problems will require training, fine-tuning, or adapting models. Others will require better retrieval, bet

pythonawsazure
View job →
C
Cloudflare
📍 Hybrid• Full-time• Hybrid• $230K – $281K/yr
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. About the team and the role The Developer Tooling team builds internal tools and platforms that improve the velocity and developer experience of Cloudflare's engineering teams. We own developer productivity (code commit to production), developer insights and ADLC metrics, developer infrastructure (CI/CD, build systems), AI-assisted development, and Cloudflare-on-Cloudflare dogfooding. This team is responsible for AI-assisted development across all

typescriptawskubernetes
View job →
T
Twilio
📍 - US• Full-time• Remote• $155.5K – $194.4K/yr
1mo ago

Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Machine Learning Engineer. About the job This position is needed to drive innovation and the development of cutting-edge products that serve developers, builders, and operators within Twilio’s Data & Observability Substrate organization. This is a hands-on, builder-focused engineering role that bridges Product, Design, and Engineering to develop, evaluate, and maintain scalable, low-latency, ML-based systems for real-time applications. You will lead rapid research-to-production cycles that translate business ideas into solutions for complex problems—such as streaming anomaly detection, recommendation systems, predictive modeling, and agentic AI frameworks—with the goal of delivering personalized customer experiences. You will collaborate closely with a cross-functional team of engineers, architects, product managers, UI/UX designers, and ML/data science partners to deliver robust, reliable solutions that power c

REMOTEpythonjavasql
View job →
D
Datadog
📍 Massachusetts• Full-time• From $234K/yr
1mo ago

The ML Observability team builds cutting-edge tools to monitor, explain, and improve AI systems in production, particularly those leveraging Large Language Models (LLMs) and generative AI. We provide robust, scalable observability for AI workloads, including drift detection and model evaluation, and behavior tracing, enabling customers to ship AI with confidence. As a Staff Engineer, you’ll lead the development of new features and foundational capabilities within Datadog’s LLM Observability product. You will shape product direction, drive experimentation, and apply your deep understanding of both AI systems and software engineering to solve open-ended problems in the fast-moving AI landscape. Your work will directly impact how our customers monitor, troubleshoot, and optimize LLM-based applications in production. Join us in building the foundational tools that make AI systems observable, understandable, and reliable in the real world. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Drive design and implementation of LLM observability features. Ideate, prototype, and scale new product features to provide insights and drive improvements for generative AI systems Work cross-functionally with other eng teams, product, UX, and applied science to iterate fast and find product-market fit Develop and extend tools for tracing, evaluating, and debugging LLMs Influence architecture decisions and mentor engineers to build resilient, high-performance systems Stay close to customer pain points and use those insights to guide product and engineering priorities Stay current with industry trends and advancements in machine learning and observability, driving innovation within the team Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or r

machine learningaigo
View job →
I
Instacart
📍 BC, Canada• Full-time• Remote• From C$168K/yr
1mo ago

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacart’s AI Productivity team builds AI-powered platforms and tools that help our engineers and operators move faster, reduce toil, and deliver higher-quality experiences for customers, shoppers, retailers, and brand partners. As a Senior Software Engineer on this team, you will design, build, and operate production systems that bring large language models and intelligent automation into everyday workflows across Instacart. You’ll partner closely with product, developer platform, ML platform, security, and data teams to ship reliable, secure, and measurable solutions that improve developer velocity and operational efficiency. You’ll join a collaborative group of approximately 10 engineers who value ownership, iteration, and pragmatic problem solving. If you enjoy rolling up your sleeves, navigating ambiguity, and turning cutting-edge AI into real, scala

REMOTEtypescriptpythonjava
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $186.1K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Senior Software Engineer, Backend - Platform (Tokens & Wrapped Assets) We're hiring a Senior Software Engineer to join the Tokens & Wrapped Assets team within the Platform organization. This team builds and operates the platform that issues, wraps, and transforms every asset at Coinbase, including flagship wrapped assets like cbBTC, spanning Ethereum, Base, and Solana. You'll own complex projects end-to-end, from smart contract through APIs and production rollout, designing and operating the Go services that mint, burn, bridge, and reconcile wrapped assets across chains. What you'll do: Own end-to-end delivery of complex token and wrapped asset initiatives, from smart contract deployment through API design, SLOs, and production rollout across multiple chains. Build and operate the Go services and smart contracts that mint, burn, bridge, and reconcile wrapped assets, ensuring correctness across on-chain and off-chain systems. Drive improvements to the asset transformation and migration pipeline, making it dramatically faster and safer to launch, upgrade, and migrate tokens at scale. Partner across Trading, Custody, USDC, Compliance, and the broader crypto stack to ship customer- and revenue-impacting work that advances onchain innovation. Lead on-call ownership for team services, building observability, runbooks, and operational improvements that keep T

REMOTEawsaigo
View job →
S
Stripe
📍 Toronto• Full-time
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Our Applied ML team aims to reform how our users interact with Stripe. We are doing so by (a) automating the easy tasks, and (b) assisting our users in the difficult tasks. Some examples include helping our users resolve issues with Stripe faster or making it easier for our users to sign up and navigate Stripe. We are using the latest LLMs as well as fine-tuning our own models. We're an end-to-end team going from ideas to models to shipping in production. You can learn more about our team’s work from this recent talk . What you’ll do As a machine learning engineer, you will be responsible for analyzing opportunities, proposing ideas, training & evaluating ML models, running experiments, and deploying everything to production. You will also have the opportunity to contribute to and influence ML architecture at Stripe as well as be a part of a larger ML community. Responsibilities Our team operates fluidly and here are some problems you may tackle: How do we evaluate a system offline & online? How do we improve performance to match (and beat) humans? How do we ensure model quality doesn’t degrade online? Does fine-tuning an LLM give us better performance? What are the right OSS and in-house platforms we should invest in? And in the process you will: Develop pipelines and automated processes to train and evaluate models in offline and online environments Integrate ML models into production systems and ensure their scalability and reliab

machine learningaigo
View job →
T
Taskrabbit
📍 San Francisco• Full-time• $175K – $225K/yr
1mo ago

About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! This is a hybrid role that will require two days in-office each week on Tuesdays and Wednesdays at our SF location on 130 Sutter Street. About the Role Every business function at Taskrabbit — Marketing, Customer Support, Finance, Operations — is a "customer" with real workflows, real data, and real friction. Your job is to embed with them, scope their use cases, and build the Claude-powered agent, automation, or tool that solves them. This is an internal-facing role — there is no external customer or product work. Reporting to the Director of AI Strategy and Enablement, you'll operate the way an FDE operates at a high-growth AI company: full ownership of a deployment from discovery through production and direct accountability for whether what you ship actually changes a metric. In most cases you'll own a build end-to-end solo; in some functions you may partner with that team's own subject-matter expert to pair domain depth with

O
1mo ago

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

pythonawsazure
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. We operate some of the largest Spark compute fleets in production; design, and build data lakes and metadata systems on Iceberg and Delta with a vision toward exabyte-scale architecture; run high throughput streaming platforms on Kafka and Flink; provide orchestration with Airflow; and support ML feature engineering tooling such as Chronon. Our mission is to deliver reliable, secure, and efficient data access at scale and accelerate intelligent, AI assisted data workflows. Join us to build and operate these core platforms that underpin OpenAI products, research, and analytics. We’re not just scaling infrastructure – we’re redefining how people interact with data. Our vision includes intelligent interfaces and AI-assisted workflows that make working with data faster, more reliable, and more intuitive. About the Role This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You’ll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You’ve supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You’re well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per wee

awsrestmachine learning
View job →
O
1mo ago

About the Team ChatGPT is a rapidly evolving system: new capabilities ship continuously, product surfaces change quickly, and usage patterns shift week-to-week. Supporting that pace requires infrastructure that can handle real production constraints—high concurrency, unpredictable traffic patterns, complex dependency graphs, and frequent change. The ChatGPT Infrastructure team builds and operates the platforms that enable fast iteration without compromising performance or reliability. We design shared systems, data paths, rollout mechanisms, and reliability guardrails that teams rely on to ship changes to ChatGPT at scale. We focus on high-leverage infrastructure: primitives and “golden paths” that incorporate operational lessons as defaults, so engineers don’t need to rediscover failure modes, latency pitfalls, or integration issues each time they build something new. About the Role We’re hiring Senior and Staff Engineers to design and build infrastructure systems that underlie ChatGPT and multiply the effectiveness of teams building user experiences. This is not a support-only role. It’s a platform-building role: you’ll define interfaces, develop core abstractions, and create tooling to make safe, fast iteration the norm. Your work will reduce friction, prevent regressions, improve performance, and ensure systems scale gracefully as the product grows. Where You Can Have Impact You might work on one or more of the following areas (without being restricted to any single area): Platform foundations & frameworks: Core libraries, service frameworks, and shared components that standardize system building, integration, and evolution. Scalability & performance primitives: Patterns and infrastructure that reduce tail latency, improve throughput, and keep costs predictable as demand increases. Reliability guardrails: Mechanisms that prevent outages by design—rate limiting, load shedding, dependency isolation, backpressure, safe fallbacks, and robust regression contr

redisawsrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users. That means working across the systems that make Codex actually function as an agent in the real world: the production performance envelope around tokens, latency, reliability, cost, and capacity; the core execution loop and interfaces that turn models into useful behavior; the shared infrastructure that enables other teams to build on Codex; and the feedback loops that turn real-world usage into better models and better agent behavior over time. About the Role We’re looking for engineers to build the infrastructure that powers Codex agents in production. This role focuses on the systems that let models safely execute code, interact with tools, complete long-running tasks, and operate reliably and efficiently at scale. You’ll design and operate the infrastructure behind sandboxed execution, orchestration, stateful workflows, app-server and SDK boundaries, and model rollouts. You’ll work at the intersection of distributed systems, developer tooling, and AI, building primitives that make Codex faster, safer, more reliable, and easier for the rest of the organization to build on. What You’ll Do Design and build execution environments for AI agents, including sandboxing, isolation, and reproducibility. Develop systems for agent orchestration across multi-step, tool-using workflows. Build infrastructure for running, testing, and debugging code generated by models. Create state and memory systems that allow agents to persist context across long-running tasks. Optimize tokens, latency, reliability, and cost across Codex’s production fleet. Support model rollouts, capacity planning, and the core tradeoffs between quality, speed, and economics to manage a fleet of frontier agents at scale. Build shared platform capabilities that unblock product teams, partner teams, and open source Codex. Yo

awsci/cdrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users. That means working across the systems that make Codex actually function as an agent in the real world: the production performance envelope around tokens, latency, reliability, cost, and capacity; the core execution loop and interfaces that turn models into useful behavior; the shared infrastructure that enables other teams to build on Codex; and the feedback loops that turn real-world usage into better models and better agent behavior over time. About the Role We’re looking for applied AI engineers to help bring Codex agents from impressive demos to dependable tools. This role is about improving agent performance on real software engineering tasks and closing the gap between research capability and real-world usefulness. You’ll work closely with research, infrastructure, and product to ensure agents are not just powerful, but useful, steerable, and reliable in practice. The job is not only to improve model behavior in isolation, but to turn those improvements into measurable gains in solve rate, usefulness, and economic value for users. What You’ll Do Design and iterate on agent behaviors across real-world coding tasks and long-horizon workflows. Work closely with research to develop and run evals to measure agent performance, regressions, failure modes, and edge cases. Improve performance through prompting, tool-use strategies, context construction, and model-facing experimentation. Analyze failures in production and systematically improve robustness and reliability. Build feedback loops and data systems that get better real-task data into evaluation and research. Work with product teams to shape user-facing agent experiences and the interfaces the agent depends on. Help define what “good” looks like for agents completing complex tasks end-to-end. You Might Be a Good Fit If You Ha

pythonawsrest
View job →
🔔

Get new production tech jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More production tech opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.