NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! This is an outstanding opportunity for someone with HPC and performance background to advance the state of the art in this space. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters. Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available Triage and root-cause performance issues reported by our customers Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information <li
Jobs in United States
Docker in United States
158 active opportunities · Updated October 2026
Showing
15 jobs
Explore current docker jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Forward Deployed Engineer at Baseten, you will partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten’s platform. You’ll own the journey with customers from initial exploration to production deployment, translating ambiguous business goals into reliable, observable services with clear quality, latency, and cost outcomes. This role is a great fit for entrepreneurial engineers who want a front-row view into how modern companies adopt AI at scale and who enjoy working across product, software development, performance engineering, and customer-facing implementations. To be clear, this is an engineering role with hands-on coding and software development that also includes aspects of product management, technical customer success, and pre-sales solution engineering mixed in. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Develop and maintain software systems and product features using one or more general-purpose programming languages in a production-level environment, with a preference for Python due to its relevance in ML projects. Drive customer impact by designing, implementin
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Are you passionate about advancing the application of artificial intelligence? We are looking for a Software Engineer focused on ML performance to join our dynamic team. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM Inference. If you are a backend engineer who thrives on making things faster and is excited about open-source ML models, we look forward to your application. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Model Performance team: Baseten Embeddings Inference: The fastest embeddings solution available The Baseten Inference Stack Driving model performance optimization RESPONSIBILITIES Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure. Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues. Apply and scale optimization techniques across a wide range of ML models, particularly large language models. Collaborate with a diverse team to design and implement innovative solutions. Own projects from idea to production. REQUIREMENTS Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field. Experience with one
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten's engineers want to work in an AI-first way. What's missing isn't enthusiasm — it's the platform underneath it. Today everyone assembles their own agent config, context files, and MCP servers, so the good patterns stay trapped in individual setups instead of becoming defaults everyone inherits. You'll build that platform: the agent configurations tuned to our monorepo, the context and tooling layer that makes agents competent in our codebase, the evals that tell us which approaches actually work, and the rollout mechanics that get a new engineer productive with agents in week one. You are not here to mandate how engineers use AI — you're here to make the good path the easy path. Success looks like teams adopting what you build because it beats what they'd cobble together themselves, not because a policy requires it. Platform engineer, not AI evangelist. Ship infrastructure, measure it, kill what doesn't work, let adoption be the referee. The playbook for AI-first SDLC doesn't exist at any company yet. You'll write ours. WHAT YOU'LL BUILD Agent substrate — Repo-level context infrastructure that makes agents competent in our codebase ( CLAUDE.md/AGENTS.md conventions, architecture and domain context, and the tooling to keep it accurate as code moves). Internal MCP servers giving agents scoped access to CI, observability, incident tooling, deployment state, and docs. Shared skills, subagents, and hooks th
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Forward Deployed Engineer at Baseten, you will partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten’s platform. You’ll own the journey with customers from initial exploration to production deployment, translating ambiguous business goals into reliable, observable services with clear quality, latency, and cost outcomes. This role is a great fit for entrepreneurial engineers who want a front-row view into how modern companies adopt AI at scale and who enjoy working across product, software development, performance engineering, and customer-facing implementations. To be clear, this is an engineering role with hands-on coding and software development that also includes aspects of product management, technical customer success, and pre-sales solution engineering mixed in. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Develop and maintain software systems and product features using one or more general-purpose programming languages in a production-level environment, with a preference for Python due to its relevance in ML projects. Drive customer impact by designing, implementin
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Platform team within the Infrastructure organization. As a senior member of Baseten's Platform Team, you will own the systems that let every engineer at Baseten prove their code works before it reaches production. Our product runs mission-critical AI inference for customers who measure downtime in dollars per second, which means our internal bar for correctness, performance, and failure tolerance has to be exceptional. Your focus is the full testing stack: fast and reliable unit test tooling, integration harnesses that spin up realistic environments on demand, load and performance testing for GPU-backed inference workloads, and resilience testing that deliberately breaks things so our customers never have to find out what happens when a node dies mid-request. This is a builder role with org-wide leverage. You won't be writing tests for other teams — you'll be building the frameworks, harnesses, and feedback loops that make writing good tests the path of least resistance, and you'll set the standards for what "well-tested" means at Baseten. RESPONSIBILITIES Own Baseten's testing strategy end to end — define the standards, the tiers, and the tooling that engineering teams build against. Build and maintain unit, integration, load and performance testing frameworks Design end to end test infrastructure that provisions realistic dependencies
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Voice is becoming the internet’s next interface, but a production-grade Voice AI system is "hard to build" . You’ll join a small founding team of Baseten Voice AI, focused on bringing state-of-the-art open source models into production for Voice AI customers across productivity, customer service, clinical conversation, creator tools, education, and more. You’ll make a meaningful impact on people’s daily lives and help reshape these industries. This is a high-impact, high-ownership role. You will be the primary owner of Baseten Voice AI - our in-house inference stack to power Voice AI models - from product roadmap through engineering implementation. You’ll partner closely with Forward Deployed Engineers, Model Performance Engineers, and sister engineering teams to push the boundaries of Voice AI. EXAMPLE INITIATIVES: Develop world-class model serving stack for state-of-the-art open-source voice models - reduce end-to-end and tail latency (p95/p99), increase throughput, and improve GPU efficiency via profiling, runtime tuning, and server-level optimizations. Build large-scale, real-time infrastructure for multi-model voice agents - orchestrate STT, TTS, and agent components with streaming I/O to meet customer SLOs. Design tight training and inference iteration loops for voice model customization - enable fast evaluation, safe rollout, and rapid experimentation for custom voice model development. Past projects:
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Engineering Manager (Player & Coach), you will lead and mentor a team of Forward Deployed Engineers focused on building, scaling, and optimizing LLM inference workloads for Baseten customers. Applying both hands-on technical ownership and managerial leadership, you will guide your team through the processes of designing, deploying, and managing high performance, low latency AI applications on Baseten’s platform. FDE at Baseten is not a sales function – we are a mix of engineering, product, and customer architects who contribute to the core Baseten codebase, drive large portions of our feature roadmap, and execute on complicated customer engagements. You will also partner with product, infrastructure, and other customer engineering teams to ensure that large language models (LLMs) and other generative AI systems deliver best-in-class performance, reliability, and cost efficiency in production environments. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Leadership & Team Management Lead, mentor, and grow a team of Forward Deployed Engineers, providing guidance on technical direction, project execution, and professional deve
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides tools that help developers find and fix issues in their applications. The Developer Infrastructure team owns the systems that make every engineer at Sentry more effective at shipping high quality software: developer environments, CI/CD for our open source and closed source codebases, the golden path for building new services, our SDK and library publishing tooling, and the metrics and dashboards that keep it all healthy. Our goal is straightforward: developers should spend their time thinking about the software they build for customers, not the software they use to build it. That means everything from local and cloud development environments, to the CI/CD pipeline that gets code out safely and quickly, to the tooling that catches flaky tests before they cost someone a day. As AI coding agents become part of how engineers work, we're also investing in making sure our environments, CI, and deployment systems support that shift. As a Software Engineer on Dev Infra, you'll help build and scale this infrastructure end to end. In this role, you will Build and maintain the tooling that powers local and cloud-based developer environments Improve CI for our codebases and CD for our deployment pipeline, keeping both fast and reliable as the org and its infrastructure grow Define and evolve the golden path for building new services and libraries at Sentry Contribute to our SDK and library publishing tooling and release processes Build the metrics, dashboards, and flaky test detection that give the org visibility into deployment health and CI reliability Build tooling that gives engineers, and the AI codin
From $85K/yr
The Team As Datadog’s in-house product experts, the Technical Escalation Engineering (TEE) team plays a critical role in driving our global success. We enable our customers, from the world’s most innovative startups to the largest enterprises. Through deep technical expertise, relentless problem-solving, and exceptional customer engagement, we educate, guide, and troubleshoot, delivering high-impact solutions that shape the customer experience. Whether through hands-on technical call, in-depth fact findings meeting, or complex investigations, we set the gold standard for technical excellence and customer advocacy. As part of our TEE team, you’ll tackle the most challenging technical problems, collaborate directly with Engineering and Product to refine and evolve our platform, and mentor teams worldwide, elevating the technical bar at every level. Role Summary As part of the Technical Escalation Engineering (TEE) team, you’ll operate at the heart of Datadog’s ecosystem, working at the intersection of Technical Solutions, Engineering, Product, and our Customers. Every challenge you take on will directly impact the performance, scalability, and success of both our clients and our platform. You’ll be in an environment that moves fast, challenges you daily, and rewards curiosity, ownership, and technical excellence. This is your chance to shape the future of observability and security, driving innovation, mentoring teams, and influencing product direction while witnessing your expertise make an immediate and lasting impact. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Develop deep technical expertise and continuously learn as the product evolves. Investigate complex escalations, lead high-stakes technical calls, and drive solutio
Note: if you are an intern, new grad, staff, frontend or fullstack applicant, please do not apply using this link and visit our jobs page for those specific postings. Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the most ambitious startups — use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the Organization Payments and Risk Sub-orgs: Link and OCS, Payments Acceptance, Risk The Payments organization focuses on developing products and platforms that enable users to accept payments from customers efficiently. This includes building APIs for processing payments, enabling regional, non-card payment options, and extending Stripe's capabilities to make it easy for businesses to accept in-person payments. Optimized Checkout and Link teams work to create best-in-class checkout experiences that enhance customer satisfaction and drive merchant conversion rates. The Risk Engineering team develops products that minimize financial and regulatory risks while ensuring a seamless user experience, thereby safeguarding Stripe's brand and financial stability. Team Matching: Exact team matching for one of the sub-teams will begin during final stages. Please note we may also consider you for different orgs based on your experience, location, etc. What you'll do We're looking for backend engineers who want to make an impact on managing money at a global scale with a passion for building ergonomic APIs. Our team collaborates with many cross-functional teams — from Infrastructure to Product — to deliver innovative solutions that address evolving user needs. Responsibilities
Note: if you are an intern, new grad, staff, front-end, or full-stack applicant, please do not apply using this link and visit our jobs page for those specific postings. Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team MaaS sub-organizations include Accounts and Connect, Money Movement and Storage (MMS), Banking as a Service (BaaS), and Crypto. Money as a Service (MaaS) oversees a diverse portfolio of core Stripe products and platforms. These offerings facilitate the global movement and management of funds for users. The teams that fall under the MaaS umbrella include Accounts and Connect, Money Movement and Storage (MMS), Crypto, and Banking as a Service (BaaS). Together, these teams work to ensure Stripe users have the robust financial infrastructure and tools they need to power their businesses on a global scale. Exact team matching for one of the subteams begins during final stages. Please note we may also consider you for different orgs based on your experience, location, etc. Find more information on our team matching process here. What you'll do We're looking for Backend Engineers who want to make an impact on managing money at a global scale with a passion for building ergonomic APIs. You'll play a key role in extending our balance management platform and in building out a new funds accessibility platform that enterprises and SMBs use. Our team collaborates with many cross-functional teams—from Infrastructure to Product—at Stripe to deliver innovative solutions that address evolvin
$83K – $127.8K/yr
Software Developer in Test Description - This role is responsible for ensuring quality, reliability and performance of software applications throughout the development lifecycle primary through software test automation. The role designs, codes, and implements software test automation using appropriate programming languages, frameworks, and tools. The role works closely with cross-functional teams to gather requirements, provide technical insights, and ensure the successful execution of test automation with the main goal of improve quality of the solution. The role also creates and executes comprehensive test plans, test cases, and test scripts based on project specifications. The role sets and provides design guidance to other developers and SQA engineers regarding test automation and the test framework. *Onsite in Ft Collins 4-days a week Responsibilities • Designs quality assurance and test processes for portions of end-user video conferencing/collaboration application, systems software running on android hardware, local, networked, and Internet-based platforms. • Analyzes design and determines test scripts, coding, automation, and integration activities required based on specific objectives and established project guidelines. • Designs and maintains Test automation framework • Executes and writes portions of testing plans, protocols, and documentation for assigned portion of application; identifies and debugs issues with code and suggests changes or improvements. • Identifies opportunities for performance improvements and optimizes code and application performance. • Utilizes latest AI tools and technologies in speeding up test automation • Executes test cases depending on the needs of the project • Provides valuable input into the development of user stories and acceptance criteria, shaping a quality-oriented d
The Applications Development Technology Lead Analyst is a senior level position responsible for building robust, high-performance, large-scale applications. The overall objective of this role is to lead applications systems analysis and programming activities. Responsibilities: * Partner with multiple management teams to ensure appropriate integration of functions to meet goals as well as identify and define necessary system enhancements to deploy new products and process improvements * Resolve variety of high impact problems/projects through in-depth evaluation of complex business processes, system processes, and industry standards * Provide expertise in area and advanced knowledge of applications programming and ensure application design adheres to the overall architecture blueprint * Utilize advanced knowledge of system flow and develop standards for coding, testing, debugging, and implementation * Develop comprehensive knowledge of how areas of business, such as architecture and infrastructure, integrate to accomplish business goals * Provide in-depth analysis with interpretive thinking to define issues and develop innovative solutions * Serve as advisor or coach to mid-level developers and analysts, allocating work as necessary * Appropriately assess risk when business decisions are made, demonstrating particular consideration for the firm's reputation and safeguarding Citigroup, its clients and assets, by driving compliance with applicable laws, rules and regulations, adhering to Policy, applying sound ethical judgment regarding personal behavior, conduct and business practices, and escalating, managing and reporting control issues with transparency. Qualifications: * 5 to 8 years of relevant experience in Apps Development or systems analysis role * Extensive experience system analysis and in programming of software applications * Hands-on experience in
The Applications Development Technology Lead Analyst is a senior level position responsible for building robust, high-performance, large-scale applications. The overall objective of this role is to lead applications systems analysis and programming activities. Responsibilities: * Partner with multiple management teams to ensure appropriate integration of functions to meet goals as well as identify and define necessary system enhancements to deploy new products and process improvements * Resolve variety of high impact problems/projects through in-depth evaluation of complex business processes, system processes, and industry standards * Provide expertise in area and advanced knowledge of applications programming and ensure application design adheres to the overall architecture blueprint * Utilize advanced knowledge of system flow and develop standards for coding, testing, debugging, and implementation * Develop comprehensive knowledge of how areas of business, such as architecture and infrastructure, integrate to accomplish business goals * Provide in-depth analysis with interpretive thinking to define issues and develop innovative solutions * Serve as advisor or coach to mid-level developers and analysts, allocating work as necessary * Appropriately assess risk when business decisions are made, demonstrating particular consideration for the firm's reputation and safeguarding Citigroup, its clients and assets, by driving compliance with applicable laws, rules and regulations, adhering to Policy, applying sound ethical judgment regarding personal behavior, conduct and business practices, and escalating, managing and reporting control issues with transparency. Qualifications: * 6 to 9 years of relevant experience in Apps Development or systems analysis role * Extensive experience system analysis and in programming of software applications * Hands-on experience in
Other cities to consider
More places hiring for this role
Get new docker jobs in United States by email
Daily job updates · Unsubscribe anytime