The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
Jobiba hiring network
Scheduler Jobs
1,211 active opportunities · Updated for October 2026
Fresh results
11 shown
Explore current scheduler jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a Hardware Chips Programs Manager at OpenAI, you will help bring our chips hardware roadmap to life, navigating an array of technical and partnership challenges. We’re looking for people excited to push the frontiers of computing by navigating technical explorations and are passionate about building. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Manage the design and implementation planning of our ML acceleration hardware, working across technical, cross-functional and external stakeholders Lead planning and scheduling of chip hardware designs with our strategic partners and vendors Coordinate and marshal internal resources and communication for efficient interaction with partners and vendors. You might thrive in this role if you: Have experience as a technical program manager for data center hardware products (server, GPU, TPU, networking, storage and so on) Know the whole end-to-end system program management from concept, design, production, deployment into the data center Have some experience with System SW programs through NPI Want to help design some of the world’s largest supercomputing systems, working at the edge of complex hardware challenges Enjoy working with and enabling world-class AI Researchers and Engineers Are passionate about the technical program function, and enjoy independently owning and delivering on your tea
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will develop and evolve the tooling ecosystem that hardware engineers rely on every day — from hardware compilers and IR transformations to simulation, debugging, and automation infrastructure. The work spans software engineering, compiler concepts, and practical hardware workflows, with direct impact on how quickly and effectively we design next-generation AI systems. You’ll collaborate closely with architects, RTL designers, and verification engineers to translate real engineering friction into durable, scalable tooling solutions. In this role you will: Build and improve the software tooling that makes hardware teams faster: compilation, IR transforms, RTL generation, simulation, debug, and automation. Extend and integrate hardware compiler stacks (frontends, IR passes, lowering, scheduling, codegen to Verilog/SystemVerilog) and connect them to real design workflows. Improve developer experience and reliability: reproducible builds, better error messages, faster iteration loops, and dependable CI and regression infrastructure. Work closely with designers and verification engineers to turn real pain points into durable tools. Dive into RTL when needed: read and reason about Verilog/SystemVerilog to debug issues, validate tool output, and improve debuggability. Be willing to go all the way down the stack when necessary, including gate-level views, synthesis results, and implementation artifacts. Help enable PPA optimization loops by building analysis and au
About the Role The Engineering Acceleration team builds and operates the foundational systems that engineers use to build, test, and ship ChatGPT, the API, and OpenAI's infrastructure. We are looking for an engineer to help evolve OpenAI's build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, and software quality. You will work on the systems that determine how quickly and confidently engineers can move: Bazel-based builds, Buildkite pipelines, test selection, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to make OpenAI one of the most productive engineering organizations in the world while preserving a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping useful systems instead of fighting infrastructure. In This Role, You Will Own and evolve Bazel-based build and test workflows across a large, polyglot monorepo. Design and maintain Starlark rules, macros, toolchains, and integrations that make builds reproducible, hermetic, and easy for product teams to adopt. Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, test sharding, retry behavior, and flake isolation. Build systems that reduce unnecessary CI work through affected-target detection, dependency graph analysis, test selection, caching, batching, and smarter scheduling. Improve local development workflows so engineers can reproduce CI behavior, debug build failures, and iterate quickly without learning every detail of the build stack. Operate and optimize build infrastructure across Docker/OCI images, Kubernetes-based runners, cloud resources, and remote cache/exec
About the Team Our Executive Operations team includes Executive Business Partners and Administrative Business Partners, who serve as trusted advisors and collaborators to OpenAI's executives and leaders, focused on strong communication and operational excellence across teams. With a focus on elevating the impact and efficiency of leadership, we anticipate needs, streamline processes, and provide comprehensive support to ensure our executives can focus on high-impact initiatives. We are pivotal in driving success and achieving key milestones by cultivating strong relationships and leveraging our deep understanding of business objectives. With a commitment to excellence and a proactive approach, we are dedicated to empowering our executives and contributing to the overall growth and success of the company. Our leadership team reflects OpenAI’s culture and core values and is a mission-driven, kind, and thoughtful group. We take pride in creating a work environment that fosters collaboration, open communication, and authenticity, making OpenAI an excellent place to work for highly accomplished professionals. About the Role: This posting is part of a shared hiring process for Executive Business Partner and Administrative Business Partner opportunities at OpenAI. Rather than hiring for a specific team, we consider candidates across multiple opportunities and identify the best fit based on your experience, interests, and business needs as you progress through the interview process. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Manage complex calendars, balancing competing priorities while ensuring leaders’ time is aligned with business needs. Coordinate internal and external meetings, resolve scheduling conflicts, and facilitate effective communication across stakeholders. Plan and manage domestic and international travel, ensuring seamless logis
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches. About the Role In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs. In this role, you will Build and refine performance models that translate microbenchmark results into cost-to-serve estimates. Analyze inference workloads end to end across applications, models, and fleet infrastructure. Enhance tooling to identify bottlenecks across layers for latency and throughput. Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference. You might thrive in this role if you: Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency. Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling. Have deep expertise with performance profiling, benchmarking, analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve o
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Our teams develops in-house connectors to connect with external systems and data warehouses. We are working to connect Snowflake with Fortune 500 companies - in the cloud. In this area you will be working on exposing API, blueprints and frameworks for other developers to create their own connectors. YOUR RESPONSIBILITIES Design and develop data integration and processing applications. These applications replicate data from various data sources (including relational databases, saas, data streams) into Snowflake, following CDC patterns. Develop and extend a robust connector platform to standardise and accelerate the development of connectors, either developed by Snowflake or by third parties. Typical example includes developing a scheduling service ensuring timely execution of tasks and optimal resource allocation. Optimize performance of the ingestion, meet with customers & troubleshoot issues and secure data transfer from external systems. Collaborate with teams across the organization and roles. Create design documents and present them to architects and other stakeholders, including company founders. Lead and mentor a group of engineers Coordinate synchronous and asynchronous communication to ensure goals are met. Collaborate with PMs and customers to understand busine
About Paytm Europe Merchants deserve better from their payments provider: fairer pricing, faster settlement, easier tools, and support from a partner that understands their business. Paytm Europe Payments S.A., backed by One97 Communications Limited, one of India's largest fintech groups, is building a merchant payments business in Europe. We are starting with Luxembourg and are currently in the early stages of shaping our product, go-to-market approach, merchant outreach, and local market understanding. We are looking for motivated interns who are curious, proactive, and excited to be part of building something from the ground up. About the Internship Duration: 6 months Location: Luxembourg Type: Internship As a Market Intelligence Intern, you will help Paytm Europe better understand Luxembourg's merchant market: visiting local businesses, speaking with merchants, collecting information on how they accept payments today, and helping the team spot opportunities. This is a hands-on role, split between the field visits and desk research. What You'll Be Doing ● Visit merchants across Luxembourg: restaurants, cafés, retail shops, convenience stores, fuel stations, and other small businesses. ● Ask merchants how they currently accept payments, which provider or card machine they use, and what works well and what does not. ● Record findings in a structured format: merchant name, location, business type, payment method, provider, and key pain points. ● Support desk research on Luxembourg's payments market: local providers, pricing models, card machines, and merchant reviews. ● Identify common challenges: payment costs, settlement delays, onboarding friction, hardware or reporting issues. ● Explore where AI or tech tools could help merchants: marketing, invoicing, inventory, reporting, loyalty, scheduling. ● Write findings into simple weekly summaries for the team. ● Coverage areas: Luxembourg City, Kirchberg, Cloche d'Or, Esch-sur-Alzette, Differdange, Dude
About Paytm Group: Paytm is India's leading mobile payments and financial services distribution company. Pioneer of the mobile QR payments revolution in India, Paytm builds technologies that help small businesses with payments and commerce. Paytm’s mission is to serve half a billion Indians and bring them to the mainstream economy with the help of technology. About the Role *We are looking for a detail-oriented and proactive CLM Operations Executive to join the Paytm. *Campaign Lifecycle Management (CLM) team. In this role, you will be responsible for the end-to-end execution of off-deck and on-deck campaigns across multiple channels using CleverTap. *You will collaborate closely with Category, Growth, and Traffic teams to ensure campaigns are set up accurately, approved on time, and monitored for performance. Key Responsibilities Campaign Execution & Management • Create, configure, and publish campaigns on CleverTap across Push (Override, Lifecycle, Nodedupe, Event-based), Email, SMS, Chat, and WhatsApp channels. • Set up and manage CleverTap Journeys, including event-based triggers, audience segmentation, and multi-step flows. • Set up and validate context strings for both Override and Lifecycle on-deck campaigns (Native Display). Data Fetching & Validation • Fetch and validate campaign base sizes using CleverTap's segment builder and cross- reference against push campaign base size checks. • Validate Banner IDs, deeplinks (via QR code generation), widget types, slot IDs, and image specs before campaign scheduling. • Monitor campaign delivery metrics post-launch; flag anomalies such as unusual drop in reach or delivery rates (expected benchmark: 50–60% delivery/acquiry). • Pull and verify data from the Campaign Tracker, Traffic Approval Sheet, and Allocation Team Sheet to ensure alignment. • Ensure all campaigns have a valid Label ID from the CLM Category List, and that CategoryId and CategoryName key-value pairs are correctly mapped (case-sensit
Get new scheduler jobs by email
Daily job updates · Unsubscribe anytime