Jobiba hiring network

Remote Software Engineer Inference Performance Optimization Specialist Manager Jobs

15 active opportunities · Updated for September 2026

Market range: $156.6K – $262.5K/yr

Fresh results

15 shown

Explore current remote software engineer inference performance optimization specialist manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.

C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Large Language Models (LLMs) continue to push the boundaries of what AI systems can do — but inference is still the bottleneck. The Model Efficiency team is responsible for pushing the limits of LLM inference efficiency across our foundation models. We explore and ship breakthroughs across the model execution stack, including: model architecture and MoE routing optimization decoding and inference-time algorithm improvements software/hardware co-design for GPU acceleration performance optimization without compromising model quality Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. As a Staff Research Engineer, you will develop, prototype, and deploy techniques that materially improve how fast and efficiently our models run in production. You may be a good fit

gitrestmachine learning
View job →
A
1mo ago

About Anyscale At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure. As part of this role, you will Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices We'd love to hear from you if you have Familiarity with running ML inference at large scale with high throughput and low latency Familiarity with deep learning and deep learning frameworks (e.g. PyTorch) Solid understanding of distributed systems, ML inference challenges Bonus points

machine learningai
View job →
C
Coinbase
📍 - USAFull-timeRemoteFrom $218K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Staff Software Engineer on the Staking Platform team within the Platform group, you'll serve as Coinbase's definitive Solana staking technical authority, owning strategy across validator operations, staking integrations, and protocol evolution. Coinbase operates one of the largest Solana staking operations in the world, managing approximately 9.25% of all staked SOL. You'll combine deep Solana protocol mastery with hands-on engineering execution and external ecosystem influence to shape Coinbase's Solana staking trajectory. What you'll do: Own Coinbase's multi-year technical strategy for Solana staking across validator performance, protocol participation, and product integration, connecting engineering decisions to yield optimization, cost efficiency, and customer growth. Lead the engineering effort to achieve industry-leading APY through validator optimization, including vote accuracy, block production, MEV strategies, commission tuning, and stake distribution tooling. Serve as Coinbase's foremost authority on the Solana runtime, consensus mechanism, staking economics, and validator client landscape (Agave, Firedancer), evaluating protocol upgrades (e.g., SIMD proposals) and proactively positioning Coinbase for changes before they land. Partner with Retail and Institutional Staking product and engineering teams to architect scalable staking integrations across

REMOTEawsrestai
View job →
A
Anyscale
📍 RemoteFull-time$156.6K – $262.5K/yr · Jobiba est.
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About Ray Data Team: Ray Data is Python-native data processing engine that is a one stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting edge AI frameworks using both multi-modal and structured data. The Ray Data team currently develops and maintains Ray Data . We are a team of engineers passionate about building a Data processing engine which is a one-stop shop for all of your ML/AI needs. We are looking for exceptional engineers to build, optimize, and scale Ray for modern and increasingly complex AI workloads. As part of this role, you will: Improve the performance of Ray Data and multi-modal batch inference use cases. Ensure efficient scaling across different stages of the Data pipeline in a heterogeneous environment. Building data loading solutions for production training workloads. Focus on stability and fault tolerance at high scale Working with customers and new age AI native companies in scaling their AI workloads. We'd love to hear from you if have: At least 3-4 years of relevant work experience Solid background in building scalable and fault-tolerant distributed systems Experience with data processing, database internals. Passionate about large

pythonmachine learningai
View job →
EA
EnCharge AI
📍 IndiaFull-time$156.6K – $262.5K/yr · Jobiba est.
3 days ago

AI Software Engineer, Agent Harness Location: Bengaluru, Karnataka (or throughout India remote-friendly with travel) About EnCharge AI EnCharge AI is building the next generation AI platform. Our novel in-memory-computing architecture delivers a 10x step-function improvement in compute energy efficiency and performance for AI inference workloads. As the demands of artificial intelligence move beyond today's models, we believe fundamental underlying infrastructure must evolve. We are an experienced team of AI researchers, silicon & systems engineers, and architects backed by leading investors, poised to become the essential platform for the next wave of AI innovation. The Opportunity We serve open-weight models and our own bespoke checkpoints on EnCharge hardware. The models change often, and the harness around them needs to keep up. You own this layer that runs agents against files, tools, documents with permissions, memory, unattended execution, and real outputs. It will be assembled from a combination of open-source and bespoke code. Key Responsibilities Own the harness architecture end to end — agent loop, safe execution, context management, knowledge base, memory, permissions, orchestration, outputs, interfaces, observability — one component per layer, with clear interfaces so layers can be swapped. Build the pieces with no open-source equivalent e.g. session semantics, enforced permissions, memory in a human-editable file, orchestrator, and outputs. Keep pace with the models: adapters, prompt formats, tool-call schemas, stop conditions, benchmarking and evaluation. Make tool use reliable across models of uneven tool-calling quality — validation, repair, retries, fallbacks. Develop agents, tools, and MCP servers for internal and customer use cases, and review them for security before they ship. Build the evaluation harness: task suites, regression runs on every model or harness change, cost and latency per task alongside quality. Define the interfaces:

pythonaic++
View job →
T
Tenstorrent
📍 AustinFull-time$100K – $500K/yr
3 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We’re looking for a Staff Forward Deployed Engineer who’s excited to build with the engineers using the AI computers Tenstorrent makes. You will create continuity between customers, engineering, and AI inference service products. This is an engineering role first: you contribute production code, operate deployments, and you can explain a trade-off to customer leadership as clearly as to core engineering teams. This is a high-autonomy role with direct customer impact. This role is remote, based out of North America, with preference near one of our main hubs: Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You understand how accelerator compute, memory, and networking topology constrain AI workloads, and don't treat hardware as a black box. You're an early adopter of AI for your work from coding to building agentic workflows that multiply your impact. You work directly with customers to understand their challenges and provide effective solutions. You are comfortable debugging across the full inference stack: from failing requests, through the serving layer, down to OOMs or kernel dispatch if n

awskubernetesmachine learning
View job →
EA
3 days ago

EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing. EnCharge’s robust and scalable next-generation in-memory computing technology provides orders-of-magnitude higher compute efficiency and density compared to today’s best-in-class solutions. The high-performance architecture is coupled with seamless software integration and will enable the immense potential of AI to be accessible in power, energy, and space constrained applications. EnCharge AI launched in 2022 and is led by veteran technologists with backgrounds in semiconductor design and AI systems. Senior Emulation Engineer Location: India - Remote Job Description: At EnCharge AI, we are building the next generation of AI compute silicon — purpose-built for high-performance, low-power, and scalable AI inference. As an Emulation Engineer, you will play a critical role in validating complex AI accelerator architectures on emulation platforms before tape-out. This position is ideal for someone passionate about bridging the gap between hardware and software in fast-paced, deep tech environments. Responsibilities: • Set up and maintain Siemens Veloce emulation and prototyping platforms • Adapt SoC designs for Emulation and Prototyping • Develop and debug emulation testbenches and system-level environments • Support pre-silicon validation, power/performance analysis, and early software bring-up. Participate in silicon bring-up and validation. • Collaborate with design and verification teams to isolate design issues and accelerate debug. • Optimize performance of the emulation workloads and reduce turnaround time. • Work with firmware/software teams to enable use of emulators for OS and driver testing. Required Background: • BS/MS/Ph.D. in EE, CS, or related field with 7+ years of SoC design experience. • Experience with emulation platforms (Veloce, Palladium, or ZeBu) and FPGA-based prototyping systems (proFPGA, HAPS, or Protium) • Experience with emula

pythongitai
View job →
O
OpenAI
📍 San FranciscoFull-timeRemote$156.6K – $262.5K/yr · Jobiba est.
28 days ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new

REMOTEpythonawsrest
View job →
A
Amplitude
📍 RemoteFull-time$156.6K – $262.5K/yr · Jobiba est.
1mo ago

About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We’re entering a world where AI agents don’t just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents’ ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova’s throughput, correctness, and operational rigor grows dramatically. We’re looking for a Staff Software Engineer who wants to go deep on both the engine internals and the infrastructure underneath it. You’ll work across the full stack of a modern OLAP system — query planning and execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — while driving meaningful improvements to performance, cost-efficiency, and reliability at scale. You’ll influence technical direction through your work, your design reviews, and your mentorship of other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of customers. What You’ll Do Build and evolve core query engine infrastructure Work across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Design and implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Design for high-throughput automated quer

pythonjavaredis
View job →

About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We're entering a world where AI agents don't just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents' ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova's throughput, correctness, and operational rigor grows dramatically. We're looking for a Senior Software Engineer who wants to go deep on the engine internals and the infrastructure underneath. You'll own significant components of a modern OLAP system — across query execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — and drive meaningful improvements to performance, cost-efficiency, and reliability. You'll grow your technical influence through the quality of your code, your design contributions, and your collaboration with other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of enterprise customers. What You'll Do Build and improve core query engine components Contribute across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Help ensure Nova's components support

pythonjavaredis
View job →
M
Mongodb
📍 New York City; United StatesFull-timeFrom $106K/yr
1mo ago

Join the MongoDB Networking & Observability team and help build the core of a distributed database! Our team focuses on creating and enhancing components which facilitate communication between distributed processes and make these processes, and their communication, easily observable. Networking Observability’s responsibilities include improving MongoDB networking, improving the efficiency of resource utilization, and building low-overhead observability features. Our team includes engineers located in New York City and fully remote engineers. We operate close to the bottom of the stack, and have a lot of influence over the availability, performance, and robustness of our open source database. Recently, we’ve improved connection handling, explored new networking architectures, and integrated OpenTelemetry to make issues easier to diagnose and connect MongoDB to modern observability tools. We are planning to further improve our networking’s stack performance, availability and scalability as well as further enhance our observability stack using open observability frameworks. Are you excited to help the MongoDB engineering team build a better database? We are! Join us today, and we can build a faster, more reliable, exceptionally observable, database system together. This role can be based out of our New York City office or remotely within the United States and Canada. Candidate Profile 3+ years of experience building distributed systems Passionate about delivering and deploying a product with cross-team stakeholders Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Hands-on experience with building production-level code. Experience in C++ is required Interest in furthering their knowledge of networking, observability and how computer architecture and internals impact the availability of SaaS Solid verbal and written communication skills and highly motivated to collaborate with colleagues Po

mongodbawsazure
View job →
M
Mongodb
📍 AlbertaFull-timeFrom C$122K/yr
1mo ago

Join the MongoDB Networking & Observability team and help build the core of a distributed database! Our team focuses on creating and enhancing components which facilitate communication between distributed processes and make these processes, and their communication, easily observable. Networking Observability’s responsibilities include improving MongoDB networking, improving the efficiency of resource utilization, and building low-overhead observability features. Our team includes engineers located in New York City and fully remote engineers. We operate close to the bottom of the stack, and have a lot of influence over the availability, performance, and robustness of our open source database. Recently, we’ve improved connection handling, explored new networking architectures, and integrated OpenTelemetry to make issues easier to diagnose and connect MongoDB to modern observability tools. We are planning to further improve our networking’s stack performance, availability and scalability as well as further enhance our observability stack using open observability frameworks. Are you excited to help the MongoDB engineering team build a better database? We are! Join us today, and we can build a faster, more reliable, exceptionally observable, database system together. This role will be based remotely in Canada. Candidate Profile 3+ years of experience building distributed systems Passionate about delivering and deploying a product with cross-team stakeholders Solid computer science fundamentals, with strong competencies in data structures, algorithms, and software design/architecture Hands-on experience with building production-level code. Experience in C++ is required Interest in furthering their knowledge of networking, observability and how computer architecture and internals impact the availability of SaaS Solid verbal and written communication skills and highly motivated to collaborate with colleagues Position Expectations Understand and improve the current funct

mongodbawsazure
View job →
TA
3 days ago

About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine

pythonkubernetesci/cd
View job →
C
Coinbase
📍 - USAFull-timeRemoteFrom $253.9K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Staff Software Engineer on the Data Platform team within Platform , you'll define and lead the technical strategy for Coinbase's data infrastructure, spanning ingestion, transformation, warehousing, streaming, and serving systems. This is a foundational role at the intersection of distributed systems, data engineering, and AI-readiness, reporting to the Senior Director of Engineering. You'll set architectural direction, drive multi-quarter roadmaps, and transition the organization from managed-service dependency toward engineering-built, platform-grade infrastructure that powers everything from fraud detection to modern multi-agent AI architectures. What you'll do: Own the technical strategy and architecture for Data Platform, setting direction across data ingestion, transformation, warehousing, streaming, and serving systems while driving engineering-led cost reduction at the infrastructure layer. Architect data infrastructure to natively support AI and ML workloads, ensuring pipelines, data lake systems, and compute can power ML training, feature stores, real-time inference, and multi-agent AI architectures at scale. Drive the evolution to near-real-time data availability, enabling downstream teams across Coinbase to act on fresher data for fraud detection, financial reporting, and analytics. Build alignment and secure commitment from senior leadership

REMOTEawsaigo
View job →
A
1mo ago

About the Role & Team We’re looking for an Engineering Manager to lead the Data Infrastructure team within Statsig Experiment at Amplitude. You will lead a multidisciplinary team of software engineers, data engineers, and data scientists responsible for the systems that power experimentation at scale. The team owns three critical areas: Data ingestion: Collecting and importing experiment exposures, custom events, OpenTelemetry data, and real user monitoring data across SDKs, streaming systems, cloud storage, and customer data warehouses. Data computation: Building distributed computation systems that transform raw data into accurate, timely experiment results. Stats engine: Developing and productionizing the statistical methods that help customers make trustworthy decisions from their experiments. This is not a traditional data engineering management role. We are looking for a leader with a solid data science and statistical foundation who can connect advances in experimentation methodology with scalable production systems. You will help set our technical and scientific direction, translating new statistical methods and machine learning research into capabilities that customers can use reliably at scale. You’ll partner closely with data scientists, engineers, product managers, and customers to advance the state of experimentation. The ideal candidate is equally comfortable discussing causal inference and statistical power with data scientists, distributed computation architectures with engineers, and experimentation strategy with customers. What You’ll Do Lead and grow the team responsible for Statsig’s data ingestion, experiment computation, and stats engine. Define the technical and scientific strategy for advancing experimentation across both Statsig Cloud and warehouse-native deployments. Partner with data scientists and engineers to turn new statistical and causal inference methods into scalable, reliable product capabilities. Evolve our data and computatio

restmachine learningai
View job →
🔔

Get new remote software engineer inference performance optimization specialist manager jobs by email

Daily job updates · Unsubscribe anytime