Jobs in United States

Systems Administrator in United States

5,046 active opportunities · Updated October 2026

Explore current systems administrator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 New York, NY, United States· Full-time
✓ High-confidence listingCompany trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Emerging Talent has changed drastically. Ramp is building it into one of the most important recruiting functions at the company: a way to meet exceptional people earlier than everyone else and turn that connection into the next generation of Ramp talent. We are looking for a true builder to lead the full Emerging Talent function. Reporting to Ramp’s Head of Talent, you will set the strategy and lead the team responsible for the internship programme, early-career recruiting, campus and community relationships, events, candidate programming and conversion of top interns into full-time employees. You will have the backing of the executive team and a real mandate to build something category-defining. What You'll Do Own the strategy, operating model and results for Ramp’s Emerging Talent function, from early identification through full-time conversion. Build a world-class internship experience, including seasonal programmes, events, programming, manager partnership and intern-to-full-time conversion. Find original ways to identify and attr

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $184K/yr

Quick readStrong listing-quality and freshness signals

TPMs at Datadog see the problems hiding between teams, engineer away the work that shouldn’t require humans, and drive the company’s most technically complex and consequential bets to completion. Technical Program Management at Datadog operates at the intersection of engineering depth and organizational reach by driving high priority, cross-functional programs that are too complex and consequential for any single team to own. We partner with engineering on solving deeply technical problems at scale by connecting the people, decisions, and context to move Datadog's most important work forward. We build the systems and automation that make entire classes of program work self-executing. We are in the architecture conversation early, earning trust through technical judgment. We use AI to surface risks earlier, accelerate program execution plans, and find cross-team patterns that would otherwise stay hidden. The faster teams move, the more essential it is to have someone who can operate across them. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What We Expect: These are the expectations we hold for every TPM at Datadog. Technical depth, product domain expertise, and AI systems literacy; knowing how AI solutions work, where they fail, and the scope and impact of those failures. AI brings more complexity into the picture - the technical bar is higher, not lower. Build the systems that reduce the need for coordination Identify what matters before anyone asks, and automate the rest Engineer program lifecycles end-to-end See what no single team can see and own the solution Drive the company's most technically complex and consequential bets through cross-functional agreement, organizational visibility, and influence Build AI powered automation tools and

RestAIGoRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s acquisition of io marks our entry into consumer hardware and our ambition to define the next human–computer interface. Success in hardware requires strong financial stewardship across the full product cost stack—from early design and sourcing decisions through manufacturing, logistics, inventory, returns, and warranty. Hardware Finance works across Product, Supply Chain, Operations, Accounting, Systems/Data, and Finance to connect business decisions to product cost, inventory, cash, COGS, and margin. About the Role We are seeking a Hardware Finance Manager to own an assigned area of hardware COGS and inventory end to end. The initial assignment will depend on business priorities and the successful candidate’s expertise. It may include BOM and product cost, manufacturing variance analysis, inventory planning, logistics, returns and warranty, customer support, or another connected set of hardware-finance responsibilities. This is an individual-contributor role with broad scope. Prior hardware experience and deep, hands-on expertise in at least two relevant domains are required. The person will be expected to operate independently, build reusable processes and analytical workflows, and remain accountable for the analysis, judgment, and recommendations. In this role, you will: Own an assigned area of hardware COGS and inventory end to end. Own forecasting, close, and business variance analysis for the assigned scope. Provide hardware leadership with clear variance explanations, trend analysis, and forward-looking signals that connect business and supplier decisions to inventory, cash, COGS, and margin. Partner with business teams and Finance Platforms to establish the financial data, systems, and dashboards needed to support analysis. Ensure data integrity and governance through clear definitions, ownership, validation checks, controls, and review processes. Improve forecasting, reporting, systems, and finance processes so they remain reliable an

AWSRestAIGo
S
📍 United States· Full-time
✓ Quality checkedCompany trend -95.9%

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Risk Operations Technology Enablement team sits within the Risk Operations org and is responsible for partnering across Risk teams to shape the technical infrastructure that enables us to scale our Risk operations and achieve our dual goals of preventing bad actors from exploiting Stripe’s systems and delivering an exceptional user experience for legitimate users. More specifically, our team partners with Risk product and engineering teams to identify, scope, prioritize, pilot, and roll out technical infrastructure and tooling improvements to protect Stripe and serve our users. From the application of AI to our reviews, to agent diagnostic and resolution tools, to review and support case routing, to our user-facing risk support surfaces, our team’s mandate is expansive and a major lever in both protecting Stripe and its users from risk as well as improving the user experience. What you’ll do You’ll lead and develop a high-performing team of Engineering Program Managers who deliver technical programs with step-change impact for Stripe. You’ll build strong partnerships and operating mechanisms with Risk Engineering and Product, co-develop the strategy for Risk’s technical infrastructure grounded in a deep understanding of the Risk technology stack, and own mission-critical technical processes (e.g., Risk Ops tooling-stack prioritization). Responsibilities Shape and drive alignment on Risk Operations’ technical strategy that enables us

O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Plugin Developer Platform team builds the APIs, SDKs, and tools that let people extend ChatGPT and Codex. We work on plugins, connectors, the Model Context Protocol (MCP), and interactive apps. We want anyone to be able to turn a useful workflow into a plugin, share it, and have other people use it. A plugin can package instructions and skills with connections to the tools and data it needs. Our work covers plugin creation and publishing, the systems that run plugins across our products, and open standards that developers can build on. About the Role We’re looking for platform-minded engineers who know what it takes to build a platform developers want to use. You’ll work across developer-facing interfaces, APIs, and backend systems. You’ll own features from the first developer conversation through implementation and release. You’ll talk directly with developers, partners, and the open-source community. Their experience will inform the APIs and abstractions you design, the problems you prioritize, and the tradeoffs you make. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. What You’ll Do Design and ship APIs, SDKs, and services that developers use to extend ChatGPT and Codex. Make plugins easier to create, test, publish, update, and share. Improve compatibility and consistency across ChatGPT and Codex, including interactive app experiences. Contribute to MCP and other open standards, bringing practical developer needs into their design. Work with developers and partners to understand recurring problems and improve the platform, tooling, and documentation. Work with Product, Research, Security, and Trust & Safety on permissions, compatibility, and safe, reliable execution. You Might Thrive Here If You Have built software that other developers use. Your experience might include an open-source project, an API or SDK, a developer platform, internal too

AWSRestAIRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Cloud Platform Engineer, you'll envision and build robust systems and processes that ensure our infrastructure is scalable, reliable, and efficient. This can range from automating deployments and monitoring systems to optimizing performance and managing incidents. We all work closely with our users, learning from their past struggles in operationalizing ML, onboarding them onto our platform, and turning our learnings into ideas for improving Baseten. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Infrastructure team: Multi-cloud capacity management Inference on B200 GPUs Multi-node inference Fractional H100 GPUs for efficient model serving RESPONSIBILITIES Build and maintain scalable infrastructure to support the deployment and operation of machine learning models. Establish standards and best practices for reliability and performance across the infrastructure. Automate processes when relevant, particularly for managing CI/CD pipelines. Own products and projects end-to-end, functioning as both an engineer and a project manager, with a focus on user empathy, project specification, and end-to-end execution. Collaborate with cross-functional teams to understand project requirements and translate them into technical solutions. Mentor junior team members and contribute to knowledge sharing within the organization. Navigate ambiguity and exercise good judgment on tradeoffs and

KubernetesCI/CDGitMachine Learning
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently. You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. EXAMPLE INITIATIVES You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services. Building AI-assisted tooling for incident triage and response. RESPONSIBILITIES Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking. Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code. Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution. Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations. Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, con

KubernetesGitMachine LearningAI
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for PhD research interns with strong research experience in reinforcement learning, machine learning, and foundation models, including large language and multimodal models, to join our research team. This internship is well suited to candidates interested in improving existing methods and developing new techniques for large-scale model training, optimization, and inference, extending models to long-context and long-horizon tasks, and improving inference-time efficiency, reliability, and robustness in high-stakes real-world deployments. Preferred Qualifications: Currently pursuing a PhD in computer science, machine learning, or a related field. A demonstrated record of research in reinforcement learning, machine learning, foundation models, or related areas. Experience developing and evaluating large-scale models or machine learning systems. Familiari

RestMachine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for a Customer Marketing Manager who can own the full customer evidence motion at Baseten: building the systems that capture customer stories, running the co-marketing programs that amplify them, and developing the channels and assets that get those stories in front of the right people. Our customers are ML engineers and AI teams deploying serious workloads — and the stories they tell about what they've built matter. We've earned trust with some of the most demanding technical teams in the industry, and this role exists to turn that trust into evidence. RESPONSIBILITIES Co-Marketing Execution Serve as the DRI for every customer co-marketing launch end to end — managing timelines, coordinating internal and external stakeholders, and driving the process from first outreach to final publication Own the single source of truth for what's in flight across all customer co-marketing activity Coordinate with design, social, and sales to ensure every asset is built, approved, and distributed correctly Customer Evidence & Asset Library Own the customer evidence library: written case studies, video stories, customer quote repository, logo library, and sales snippets ensuring all assets stay current and are tagged and accessible for sales and marketing use Run the monthly operating rhythm: new logo additions from closed-won opportunities, asset updates, and customer health checks Programs & Channels I

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Platform team within the Infrastructure organization. As a senior member of Baseten's Platform Team, you will own the systems that let every engineer at Baseten prove their code works before it reaches production. Our product runs mission-critical AI inference for customers who measure downtime in dollars per second, which means our internal bar for correctness, performance, and failure tolerance has to be exceptional. Your focus is the full testing stack: fast and reliable unit test tooling, integration harnesses that spin up realistic environments on demand, load and performance testing for GPU-backed inference workloads, and resilience testing that deliberately breaks things so our customers never have to find out what happens when a node dies mid-request. This is a builder role with org-wide leverage. You won't be writing tests for other teams — you'll be building the frameworks, harnesses, and feedback loops that make writing good tests the path of least resistance, and you'll set the standards for what "well-tested" means at Baseten. RESPONSIBILITIES Own Baseten's testing strategy end to end — define the standards, the tiers, and the tooling that engineering teams build against. Build and maintain unit, integration, load and performance testing frameworks Design end to end test infrastructure that provisions realistic dependencies

PythonDockerKubernetesCI/CD
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab

KubernetesMachine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use. As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently. This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users. EXAMPLE INITIATIVES Blog Posts https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/ https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/ https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid RESPONSIBILITIES Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference Work across the stack, from customer-facing features to low-le

KubernetesCI/CDRestMachine Learning
B
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for a Marketing Operations Manager who can own and harden the systems layer of Baseten's Marketing engine. Marketing at Baseten is scaling fast — more spend, more campaigns, more model launches, more inbound. The systems underneath (Our ESP, CRMs, forms, tracking, routing, alerting) need an owner who treats them like production infrastructure that cannot go down. When something breaks, it costs us time, pipeline, and trust in the data. This role exists so it doesn't break. You’ll simultaneously build for the future and re-think assumptions about our tech stack in the age of agents. This is an offensive play that gives the rest of the team leverage and superpowers to hit our ambitious goals. This is NOT an IT or service role. This is a core member of the marketing team who implements technology to achieve outcomes. RESPONSIBILTIES Own the marketing tech stack end-to-end: ad platforms, email systems, tracking, pixels, forms, connectors. Build defense-in-depth on inbound: spam/bot protection, rate limiting, email/domain validation, sync gating — and the alerting to catch anomalies before they hit sales or leadership dashboards. Enforce data integrity: UTM governance, campaign membership, lifecycle stages, lead scoring and routing logic, field-level hygiene, canonical metric definitions. Operationalize the web request pipeline with our dev agency: structured briefs, tickets, SLAs, and launch-day runb

PythonRestMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Engineering Manager for Baseten's Cloud Platform team, you will directly manage a team of cloud platform engineers responsible for building the systems and processes that keep our infrastructure scalable, reliable, and efficient — from automated deployments and monitoring to performance optimization and incident response. You are a people-first leader with a strong cloud infrastructure background. You set a high bar for reliability and operational excellence, engage credibly in technical discussions and code reviews, and know how to build a culture of ownership and accountability. You'll spend most of your time close to the work: unblocking your team, shaping technical direction on day-to-day decisions, and developing your engineers. At Baseten, we work closely with our users to understand their struggles operationalizing ML — you'll keep your team connected to that mission and translate user learnings into better infrastructure. RESPONSIBILITIES Recruit, hire, and grow a high-performing team of cloud platform engineers; provide ongoing coaching, feedback, and career development through regular 1:1s. Set clear performance expectations, hold a high bar, and create an environment where engineers do their best work. Foster a culture of ownership, accountability, and continuous improvement. Drive day-to-day technical decisions through design reviews, code reviews, and architectural discussions; translate th

KubernetesCI/CDGitMachine Learning
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're building a platform that covers the whole life of an LLM: training it, deploying it, and observing it in production. We already run multi-node training, elastic inference, sandboxes, and distributed volumes, and we control the infrastructure underneath. We’re looking for research depth in post-training to sit alongside our systems and product work. What you'll do: We are looking for research scientists with a strong track record in reinforcement learning, machine learning, and foundation models, including large language and multimodal models, to join our research team. This role is well suited to candidates interested in improving existing methods and developing new techniques for large-scale model training, optimization, and inference, extending models to long-context and long-horizon tasks, and improving inference-time efficiency, reliability, and robustnes

RestMachine LearningAIGo
🔔

Get new systems administrator jobs in United States by email

Daily job updates · Unsubscribe anytime