About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine
Jobiba hiring network
Scheduler Jobs
1,211 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current scheduler jobs. Use filters to narrow by work mode, employment type, experience and date posted.
We are developing advanced multi-rack, multi-tenant AI/ML datacenters with NVIDIA GB200, and upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the cloud-native stack for datacenter products like GB200. In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex scheduling challenges across racks, tenants, and clouds as part of the CSP engagements team. What you’ll be doing: Perform deep-dive debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime issues, device-plugin crashes, RDMA/IB fabric anomalies, etc. Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom micro-services that expose new GPU capabilities. Drive joint architecture reviews and “whiteboard” sessions with CSP and internal platform teams; convert findings into RFCs and upstream pull requests. Create reproducible testbeds (Helm/Ansible/Terraform) that mirror customer environments; automate validation and benchmark suites. Deliver technical collateral-design docs, how-to guides, demo scripts-and present at customer on-sites, KubeCon, and SlurmUG. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Strong source-level expertise in Kubernetes internals (scheduler, CRI/CNI/CSI, operators) and Slurm (federation, power-save, plugins). Hands-on experience integrating next-gen GPUs (Blackwell/GB200/GB300) or comparable accelerators into containerized clusters. Proven track record debugging large-scale, cloud-native stacks across ne
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. The PBM Batch Operations role is responsible for providing 24x7 operational support for Pharmacy Benefit Management (PBM) production processing environments. The position monitors, controls, and supports enterprise batch workloads, mainframe systems, iSeries environments, and associated operational processes to ensure critical pharmacy and business applications execute successfully and on schedule. Schedule: WorkDays: TBD Hours: 7 AM to 7PM AZ time Shift Structure: Three 12-hour shifts Additional Requirement: Must be available to work overtime as needed to provide coverage PBM Batch Operations Functional Responsibilities The PBM Operations environment includes responsibility for: Monitoring and supporting IWS (IBM Workload Scheduler) batch processing. Batch job interventions (restart, hold, kill, force complete). Mainframe IPL support. Mainframe console monitoring across multiple LPARs. RxClaim and iSeries batch monitoring. PBM Disaster Recovery support. Vendor escort activities and data center operational support. Data center security ticket processing. MIR3 paging and incident notifications. ServiceNow ticket management. Procedure verification and operationa
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. The PBM Batch Operations role is responsible for providing 24x7 operational support for Pharmacy Benefit Management (PBM) production processing environments. The position monitors, controls, and supports enterprise batch workloads, mainframe systems, iSeries environments, and associated operational processes to ensure critical pharmacy and business applications execute successfully and on schedule. Schedule: WorkDays: Wednesday through Saturday Hours: 7:00 PM to 5:00 AM AZ time Shift Structure: 10-hour shifts Additional Requirement: Must be available to work overtime as needed to provide coverage PBM Batch Operations Functional Responsibilities The PBM Operations environment includes responsibility for: Monitoring and supporting IWS (IBM Workload Scheduler) batch processing. Batch job interventions (restart, hold, kill, force complete). Mainframe IPL support. Mainframe console monitoring across multiple LPARs. RxClaim and iSeries batch monitoring. PBM Disaster Recovery support. Vendor escort activities and data center operational support. Data center security ticket processing. MIR3 paging and incident notifications. ServiceNow ticket management. Procedure ver
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Registered Nurse Case Manager – Field in Suffolk, Brooklyn, Kings and Queens and surrounding counties, NY Position Summary: This position is a field‑based Registered Nurse (RN) role responsible for conducting member assessments for new NY Better Health enrollees, annual reassessments, semiannual reviews, and change‑in‑condition visits. The nurse will complete 2–3 home visits per day, with all documentation expected within 24 hours of each assessment. The role requires extensive travel (up to 75%) across assigned Suffolk and Surrounding Counties, NY and close collaboration with the Scheduler/CM Assistant who coordinates Work hours are Monday–Friday, 8:30 a.m.–5:00 p.m. EST. Key Responsibilities: Initial Assessments — Conduct member assessments for new NY Better Health applicants. Reassessments — Perform annual, semiannual, and change‑in‑condition assessments per regulatory and plan requirements. Clinical Documentation — Submit complete, accurate documentation within 24 hours of each visit. Care Coordination — Communicate findings to interdisciplinary teams to support care planning. Travel & Scheduling — Maintain punctuality and reliability for 2–3 daily home visits coordinated by the scheduling team. Independent Work — Operate effectively in a remote/telephonic environment while collaboratin
Leidos is seeking an experienced Program Analyst to support large information technology programs for our Utility Technologies Solutions team. In this role, you will provide program and project management support across a portfolio of IT projects supporting electric utility operations, technology modernization, and evolving grid initiatives for our utility client. You will help ensure projects remain organized, on track, and aligned with established governance and delivery standards while adhering to Leidos' Mission, Vision, and Values. The Program Analyst will support a Senior Program Manager/Portfolio Manager across a portfolio of utility IT projects, with a focus on project governance, document control, portfolio planning, reporting, and operational coordination. This role will work closely with Project Managers, IT Managers, Project Controllers, Schedulers, EPMO, and other business and technical stakeholders to help ensure projects meet established requirements and remain on track throughout the project lifecycle. The role will support stage-gate activities, maintain project and program documentation, coordinate portfolio information, assist with long-range planning, and help improve program processes and controls. The ideal candidate is highly organized, detail-oriented, proactive, and comfortable working across multiple projects and stakeholders in a fast-paced environment. Location: This position is based in South Plainfield, New Jersey, and requires onsite work three days per week, Tuesday through Thursday. Qualified candidates must reside within a reasonable commuting distance and be able to work onsite on these days. Salary: This position pays a salary in the range of $90,000 to $110,000 annually. Responsibilities: Provide day-to-day program man
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Enterprise Operations Engineer Overview The Multimedia Services team is looking for a Senior Enterprise Operations Engineer to join the team in Pune, India. The role is focused on lifecycle refresh, transformational initiatives, events support, and service delivery as well as supporting the global team as needed. The ideal candidate is highly motivated, analytical, outcomes oriented, and is a good communicator. With the increased focus on flexible workplace of the future, you will be a key contributor to a dynamic Employee Digital Experience team to enable, connect, and empower your colleagues globally. Role Working with the Regional Operations Manager: • Participate in the planning and implementation of strategic refresh projects within the region o Collaborate with local stakeholders (e.g. end users, Real Estate Services, Network Engineering, and vendors) to simplify and modernize audio visual experiences o Facilitate final acceptance of solutions and projects to the Multimedia Services Operations Team leveraging the standard processes and documentation • Be the product owner for an audio-visual technology (e.g. video conferencing, wireless content sharing, room schedulers, or electronic bulletin boards/digital signage) • Lead resolution and problem management for video
Job Overview: We are seeking a high-caliber Rust Systems Engineer to design, build, and optimize high-throughput, low-latency backend systems and infrastructure. In this role, you will go beyond web APIs—tackling low-level performance bottlenecks, complex state management, parallel compute algorithms, and concurrent data processing pipelines. If you thrive on deterministic memory management, zero-cost abstractions, and writing safe, blazingly fast concurrent systems, this role is for you. Key Responsibilities: High-Performance Architecture: Design and implement zero-cost, high-throughput, low-latency engine components and data processing pipelines in Rust. Concurrent & Thread-Safe Systems: Build thread-safe, lock-free, or fine-grained locked data structures and state machines capable of scaling across multi-core architectures without race conditions. Priority & Real-Time Threading: Design custom thread pools, task schedulers, and execution queues with priority-based task scheduling and resource allocation controls. Algorithmic Optimization: Implement complex computational logic, including high-efficiency recursive algorithms, tail-call optimizations, and dynamic cache-friendly data structures. Resource & Memory Management: Leverage Rust’s ownership model, lifetime annotations, custom allocators, and non-blocking I/O to achieve predictable low-latency profiling (minimizing allocations and cash misses). System Profiling & Benchmarking: Conduct continuous benchmarking (criterion), flame graph analysis, memory profiling (Val grind/heap track), and CPU SIMD/vectorization optimizations. Key Skills: Core Rust & Functional Programming: Advanced Rust Mastery: Deep experience with Rust internals (stdsync, stdcell, custom Drop, unsafe Rust boundaries, and macro systems). Closures & Higher-Order Functions: Mastery of Rust’s functional traits (Fn, FnMut, FnOnce), capturing environments, move semantics within closures, and passing unboxed closures for zero
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a talented engineer to join our CPU design team to define and implement RTL for high-performance CPUs. You’ll work on a CPU based on RISC-V ISA, collaborating with DV, PD, and performance teams to deliver a functional, timing, and power-converged design. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced in CPU microarchitecture with expertise in Rename, Scheduler, ROB, Load Store, Branch Prediction, Cache or Datapath. Skilled in RTL coding (Verilog/VHDL) and familiar with industry-standard tools for simulation, synthesis, and power analysis. Proficient in debugging RTL/logic across multiple design hierarchies and pre/post-silicon environments. Background in microarchitecture definition, design specification, and performance-driven trade-off analysis. What We Need Own RTL design and microarchitecture development for a portion of a CPU block of a high-performance RISC-V CPU. Collaborate closely with DV, PD, and performance engineers to meet functional, timing, and power goals. Use innovative techniques to optimize power, performance, and
Forward Deployed Senior Software Engineer (Migration Tooling – RunMyJobs) OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leader in full-stack automation fabric solutions for mission-critical business processes. Our flagship SaaS platform, RunMyJobs (RMJ) , is the first composable automation platform specifically built for ERP environments. We enable organizations to orchestrate, manage, and monitor workflows across applications, services, and infrastructure — in the cloud or on premises. Our global team of automation experts, engineers, and customer success professionals work together to deliver seamless automation transformations. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT We are seeking a Forward Deployed Software Engineer focused on migration tooling and customer onboarding to RunMyJobs (RMJ) . This role is part of an exciting new Forward Deployed Engineering team within the Global Professional Services team, at the intersection of Product Engineering and Go To Market teams. In this role, you will operate at the intersection of engineering and delivery. Your primary focus will be designing, building, and enhancing migration frameworks, tooling, and automation accelerators that enable customers to smoothly transition from legacy schedulers and automation platforms into RMJ. You will work closely with: Migration Architects to design scalable and reusable migration patterns Professional Services to enable efficient customer onboarding Engineering & Product to improve platform capabilities based on field learnings Customers (occasionally) to validate requirements, troubleshoot edge cases, and ensure successful implementations This is a forward-deployed engineering role — highly technical, impact-driv
By submitting your resume, you’re expressing interest in our 2027 RDSS (Research and Development Substitute Services) program. Please confirm your eligibility with the local district office before applying the role. NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are looking to grow our company, and grow our teams with the smartest people in the world. What you’ll be doing: You will work with ground breaking technologies for the Tegra SoC and various NVIDIA embedded platforms Implement power and thermal management software features in Linux Kernel and user space Collaborate with power architects, hardware and software engineers on usecase power estimation, power and performance optimization What we need to see: MS in CS, CE, EE, Systems Engineering or related software/hardware engineering major Software development experience with a significant focus on Linux Excellent C programming/debugging skills within Linux kernel and user space software Background with working on embedded systems and ARM processor specific System-level debugging experience and problem-solving skills Excellent communication skills Ways to stand out from the crowd: Experience in working with the Linux and open-source software communities Understanding of the Linux power management features (scheduler, dynamic frequency scaling, runtime power management, su
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Game Engine team at Roblox works on the systems that power the experiences at the heart of the metaverse. Our code is the technical foundation in our client, editor, and our simulation servers. The Game Engine team is broadly split into departments — Audio-Video-Communication, Avatar, Core AI, Digital Matter, Productivity, and Systems. This position is for the Systems Runtime pod. The Runtime pod builds and maintains the foundational C++ components that power the entire Roblox engine and Studio stack. We own the runtime layer — the performance, efficiency, and usability of the core primitives that every other engineering team relies on. Think of Runtime as the “standard library” and execution engine for Roblox's C++ world. Our ownership centers on three pillars: Concurrency System — our task scheduler and fiber runtime power how work is parallelized across CPU cores, letting teams scale features across platforms while keeping code understandable and debuggable. Memory System — we own the engine's memory allocator stack, tracking pipelines, and observability tooling to make memory behavior predictable and prevent leaks and fragmentation. Profiling and Observability — we build and evolve
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: Optimizing performance of large-scale workloads on Ray Stability and stress testing infrastructure Improving fault tolerance (HA) As part of this role, you will: Leading cross-team projects while mentoring junior team members Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements
Get new scheduler jobs by email
Daily job updates · Unsubscribe anytime