The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
Jobs in United States
Experienced Field Service Representative in United States
8,527 active opportunities · Updated October 2026
Showing
15 jobs
Explore current experienced field service representative jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team Business Systems / Enterprise Platform Technology builds the internal systems, data foundations, workflow infrastructure, and enterprise platforms that help OpenAI operate at scale. The EPT AI Pod builds AI-native internal apps, MCP connectors, multi-agent workflows, and reusable platform capabilities across Finance, People, and GTM. About the Role As an Enterprise Applied AI Engineer, you will build internal apps for enterprise operations and the shared platform components those apps run on. This includes MCP connectors, multi-agent orchestration, data architecture, evals, monitoring, auditability, and governance. We’re looking for a hands-on engineer who is strong in Python, system design, enterprise integrations, data architecture, and applied AI systems. You should be excited to turn ambiguous business workflows into reliable internal products and shared infrastructure. In this role, you will: • Build internal apps for enterprise operations across Finance, People, and GTM • Build MCP connectors and enterprise integrations with strong auth, permissions, idempotency, retries, and rate-limit handling • Design end-to-end multi-agent workflows with tool routing, human approvals, audit trails, and safe action boundaries • Design data architecture for operational AI systems, including ingestion, schemas, quality checks, lineage, and governance • Build evals, monitoring, metrics, and regression tests for agentic workflows • Create reusable infrastructure, patterns, and components that other enterprise teams can build on • Partner with system owners and business owners to turn messy enterprise workflows into reliable internal products You might thrive in this role if you: • Have strong Python engineering skills for backend services, MCP connectors, agent/tool workflows, eval harnesses, and data ingestion jobs • Have strong system design skills across shared infrastructure, app architecture, reliability, and scaling • Have experience building internal apps,
This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea
$293K – $405K/yr
About the team Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security that could scale to an extreme level of severity. Our work involves: Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework , and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. About the role As AI agents become more capable at software engineering, and automate more of our internal work, they could become a dangerous cyber threat. People in this role will help OpenAI prepare for security threats from advanced AI agent insiders. In this role, you will: Identify paths by which capable future internal AI agents could compromise OpenAI. Design security controls - focusing on measures with long lead times that benefit from advanced preparation. Stress-test defenses with AI agent evaluations and penetration tests You might thrive in this role if you: Are deeply technical across security and modern infrastructure, and are comfortable digging into the details of operating systems, cloud, containers, CI/CD, or distributed systems. Have strong software engineering skills and enjoy building prototypes yourself. Are interested in engaging with stakeholders and can do so effectively. Bonus: have experience securing cloud infrastructure, and are deeply familiar with core components of the AI stack. Compensation Range: $293K - $405K USD About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy
About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need
About the team OpenAI’s Forward Deployed Engineering team partners with customers to turn research breakthroughs into production systems. We operate at the intersection of customer delivery and core platform development. About the Role We are looking for customer-focused software engineers to build effective custom software that leverages OpenAI’s APIs to solve real customer problems. As an FDSWE, you will work with our customers and OpenAI Forward Deployed Engineers to design and implement scalable solutions that solve their most difficult problems. You will design abstractions to solve customer problems, and then use them to scale our speed and quality of delivery across all Forward Deployed engagements. You will collaborate closely with Sales, Solutions Engineering, Solutions Architects, and Customer Success Managers who work on the same account. You will also work with our Research and Applied Product and Engineering teams to provide insightful customer feedback. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. Travel up to 50% is required. In this role, you will: Embed deeply with strategic customers to understand their business challenges and technical requirements in detail. Design, architect, and develop full-stack solutions using an experiment-driven, iterative approach. Prepare detailed scopes of work and project plans for both proof-of-concept prototypes and full production deployments. Work hands-on with customers' technical teams as a technical expert and trusted advisor, coding side-by-side to drive projects to completion on their infrastructure. Collaborate with Product, Research and Applied teams to ensure seamless customer experiences, project success and actionable product feedback Contribute to internal knowledge bases, codifying best practices and sharing insights gained from customer engagements to scale the Forward Deployed Engineering function. You’ll thrive in th
About the Team OpenAI's Training team is responsible for producing the large language models that power our research, our products, and ultimately bring us closer to AGI. Achieving this goal requires combining deep research into improving our current architecture, datasets and optimization techniques, alongside long-term bets aimed at improving the efficiency and capability of future generations of models. We are responsible for integrating these techniques and producing model artifacts used by the rest of the company, and ensuring that these models are world-class in every respect. Recent examples of artifacts with major contributions from our team include GPT4-Turbo, GPT-4o and o1-mini. About the Role As a member of the architecture team, you will push the frontier of architecture development for OpenAI's flagship models, enhancing intelligence, efficiency, and adding new capabilities. Ideal candidates have a deep understanding of LLM architectures, a sophisticated understanding of model inference, and a hands-on empirical approach. A good fit for this role will be equally happy coming up with a creative breakthrough, investing in strengthening a baseline, designing an eval, debugging a thorny regression, or tracking down a bottleneck. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, prototype and scale up new architectures to improve model intelligence Execute and analyze experiments autonomously and collaboratively Study, debug, and optimize both model performance and computational performance Contribute to training and inference infrastructure You might thrive in this role if you: Have experience landing contributions to major LLM training runs Can thoroughly evaluate and improve deep learning architectures in a self-directed fashion Are motivated by safely deploying LLMs in the real world Are well-versed in the state of the art tran
$295K – $380K/yr
About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Senior Software Engineer, ML Systems & Training Infrastructure, you will be a deeply hands-on engineering force multiplier for the robotics team. You will help keep the training framework and surrounding infrastructure healthy, review and improve code quickly, debug failures across ML systems and infrastructure, and unblock researchers and engineers when the path from idea to working training job gets rough. We’re looking for people who love writing, reading, reviewing, and fixing code; who can get productive quickly in unfamiliar systems; and who bring strong practical judgment without a lot of ego or process overhead. This role will be based in San Francisco, CA and be expected in office 5 days per week and offer relocation assistance to new employees. In this role, you will: Review, improve, and clean up code across training frameworks and adjacent infrastructure. Identify risky or low-quality changes before they land, and raise the code quality bar without slowing the team down. Debug issues across ML training systems, GPUs, clusters, networking, and related infrastructure. Help researchers and engineers unblock broken training jobs, flaky workflows, and brittle internal tooling. Improve the reliability, maintainability, and usability of the robotics team’s training framework. Move quickly on practical engineering problems that directly affect team velocity. You might thrive in this role if you: Have strong software engineering fundamentals and excellent code review judgment. Have experience with ML systems, training fr
About the Team OpenAI’s Finance and Revenue Operations organization builds the commercial infrastructure that enables the business to scale with speed, discipline, and financial integrity. Deal Desk operates as the commercial strategy and governance function at the intersection of Sales, Partnerships, Legal, Technical Revenue, Finance, Order Management, Billing Operations, Product, and GTM Systems. We architect complex enterprise transactions, turn ambiguity into executable decisions, and create the guardrails that let the business move quickly with operational and financial discipline. As OpenAI’s enterprise business grows in scale and complexity, Deal Desk defines how novel commercial motions become durable operating capabilities. We convert precedent-setting deal decisions into repeatable policy, controls, workflows, and systems requirements. About the Role We are hiring a Strategic Deals & Commercial Architecture Lead to lead the structuring and governance of OpenAI’s most complex enterprise transactions. This is a senior individual-contributor leadership role for someone with exceptional enterprise deal judgment, operational rigor, and a builder mindset. You will serve as the commercial architect for high-stakes opportunities, translating ambiguous requirements into coherent deal structures, approval strategies, and executable quote-to-cash plans. Your work will shape more than individual transactions. You will establish decision principles and precedent, clarify tradeoffs, and turn recurring patterns into scalable guidance, controls, systems requirements, and enablement. You will own the commercial decision and governance layer that helps strategic opportunities move decisively while managing downstream risk across contracting, billing, revenue recognition, provisioning, reporting, controls, auditability, and customer experience. The right candidate can move seamlessly between advising on a single high-value transaction and improving the operating model be
About the Team At OpenAI, we’re building the connective tissue between our mission and our people. People Innovation Labs is a fast-moving engineering team embedded in the People organization, focused on rethinking how we find and retain the best talent and empower everyone to do their best work. From recruiting to culture, we’re designing systems that give our People Team a significant edge by infusing OpenAI’s models and first-principles thinking into every aspect of our work. Our projects range from greenfield 0-1 products like OpenHouse (our internal knowledge hub) to AI-powered automations and scalable recruiting tools. We’re defining the future of work at OpenAI, creating a blueprint for how AI can supercharge productivity, culture, and innovation. About the Role We are looking for a self-starter engineer who loves building new products in an iterative and fast-moving environment. This team is for full stack product engineers who are deeply curious about culture, recruiting and people development, and want to know everything from the business strategy and metrics down through the code that gets us there. In this role, you will work with members of the People Team and leaders across the company to build software focused on HR, culture and recruiting from the ground up. You’ll also innovate on how we apply LLMs in these domains. In this role, you will: Own the full product development lifecycle for new people products end-to-end Talk to internal stakeholders to understand their problems and design solutions to address them Work with the research team to share relevant feedback and iterate on applying their latest models Collaborate with a cross-functional team of engineers, HRBPs, recruiters, researchers, product managers, designers, and people in operations to create cutting-edge products Your background might look something like: 4+ years of professional engineering experience (excluding internships) in relevant roles at tech and product-driven companies Forme
About The Team The Data Understanding team is responsible for creating the high quality datasets and their quantized representation for OpenAI. This includes synthesizing data, building VQ representations, and processing, filtering, deduplication, quality control, and tokenization so it can be used effectively in big model training runs. About The Role We're looking to advance how OpenAI builds and understands pretraining data at scale. You'll treat data quality and curation as core research problems: developing new methods to select, combine, and transform data; creating datasets that improve model capabilities; and designing rigorous experiments to understand how data choices and interventions affect model learning and downstream behavior. You'll work closely with frontier models and web-scale data to build evidence for which approaches work and why, then translate successful research into scalable data processing pipelines We Expect You To Have a strong track record of new or improved ML ideas, through publications, projects, or applied research. Own and drive a research agenda, from choosing the right problems to carrying long-running work through to impact. Be excited by OpenAI’s empirical, collaborative approach to research. Nice To Have Thoughtfulness about AI’s impact, including privacy, provenance, and data quality. Experience building high-performance deep learning or large-scale data processing systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer
About the Team OpenAI's Research Data Team exists to accelerate the evaluation, safety and capabilities of our models and products. Made up of technical operators and software engineers, we design the methods in which we acquire and create data. About the Role As a Research Program Manager (RPM), Data Acquisition, you will partner with research, engineering, and operations to design and implement pragmatic solutions for acquiring data. You will be a key interface between our research roadmap and external data offerings. This role is based in our San Francisco HQ and will be part of a team of RPMs pushing the frontier of data acquisition. In this role, you will: Partner deeply with research: Work with researchers to scope data needs, define success criteria, and translate priorities into clear execution plans. Shape the data acquisition pipeline: Identify, evaluate, and advance high impact data opportunities - balancing research value, feasibility, quality, and responsible execution. Unblock yourself: Move work forward even when the path is unclear — using technical judgement, creative problem solving, and scrappy execution to make progress while longer-term solutions are still forming. Build lightweight systems and visibility: Use SQL, Python, dashboards, and simple tooling to track performance, quality, and blockers. Drive technical roadmaps: Collaborate with engineers to enhance data platforms, resolve blockers, and ensure security best practices such as access management. Scale your impact: Equip vendors and internal teams with the context, standards, and operating rhythms needed to focus on the most important problems. You’ll thrive in this role if you: Are proficient in SQL and Python for analysing datasets, querying databases, building dashboards, and generating actionable insights. Are comfortable using APIs, automation, and AI tools such as Codex to accelerate workflows, remove manual overhead, and upskill quickly in unfamiliar technical areas.Experience sou
About the Team The Storage Infrastructure team builds and operates the storage foundation behind OpenAI’s most demanding workloads. We work directly with research to design storage systems for rapidly evolving experiments, while also powering production at scale. We own the platform end to end: backend systems, user-facing services and APIs, and the control planes that manage how data is placed, moved, and retained over time. Our stack spans cloud and in-house object stores across very different workload profiles, from GPU-attached systems to dedicated storage hardware. We also build the federation layer that unifies these backends behind a simple interface and routes each workload to the right storage solution. About the Role You will help build the storage platform that powers OpenAI’s research and production systems. This is a hands-on infrastructure role for engineers who want to work on deeply technical systems at scale and own them in production. You’ll work across object storage, cross-region data movement, lifecycle management, and the federation layer that provides a unified interface across multiple backends. Much of our stack runs on Kubernetes, and we primarily build services in Rust. In this role, you will: Build and operate storage services that underpin OpenAI’s research infrastructure Develop object storage systems across cloud and in-house environments Build systems for cross-region data movement, replication, and recovery Design lifecycle management capabilities that keep data durable, available, and cost-effective Evolve the federation layer that unifies multiple backend systems behind a simple interface Improve performance, reliability, and operational excellence across the platform Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads You might thrive in this role if you: Have experience building or operating distributed systems in production Have worked on storage infrastructure, object stores, dist
About the Team Codex is OpenAI’s first-party developer product focused on agentic software engineering. We’re building tools that help engineers design, write, test, and ship code faster—safely and at scale. We partner tightly with research and product to translate model advances into tangible developer productivity. About the Role As a Data Scientist on Codex, you will measure and accelerate product-market fit for AI developer tools. You’ll define what “developer productivity” means for our product, run experiments on new coding models and UX, and pinpoint where the model helps or hurts across languages and tasks. Your insights will directly shape how an entire industry builds software. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Embed with the Codex product team to discover opportunities that improve developer outcomes and growth Design and interpret A/B tests and staged rollouts of new coding models and product features Define and operationalize metrics such as suggestion acceptance, edit distance, compile/test pass rates, task completion, latency, and session productivity Build dashboards and analyses that help the team self-serve answers to product questions (by language, framework, repo size, task type) Diagnose failure modes and partner with Research on targeted improvements (model quality signals, user feedback, evals) You might thrive in this role if you have 5+ years in a quantitative role at a developer-facing or high-growth product Fluency in SQL and Python; comfort with experiment design and causal inference Experience defining product metrics tied to user value Ability to communicate clearly with PM, Eng, and Design—and to influence product direction You could be an especially great fit if you have Strong programming background; ability to prototype, run simulations, and reason about code quality Familiarity with IDE/extensi
About the team OpenAI’s Forward Deployed Engineering team partners with customers to turn research breakthroughs into production systems. We operate at the intersection of customer delivery and core platform development. About the Role We are looking for customer-focused software engineers to build effective custom software that leverages OpenAI’s APIs to solve real customer problems. As an FDSWE, you will work with our customers and OpenAI Forward Deployed Engineers to design and implement scalable solutions that solve their most difficult problems. You will design abstractions to solve customer problems, and then use them to scale our speed and quality of delivery across all Forward Deployed engagements. You will collaborate closely with Sales, Solutions Engineering, Solutions Architects, and Customer Success Managers who work on the same account. You will also work with our Research and Applied Product and Engineering teams to provide insightful customer feedback. This role is based in NYC. We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. Travel up to 50% is required. In this role, you will: Embed deeply with strategic customers to understand their business challenges and technical requirements in detail. Design, architect, and develop full-stack solutions using an experiment-driven, iterative approach. Prepare detailed scopes of work and project plans for both proof-of-concept prototypes and full production deployments. Work hands-on with customers' technical teams as a technical expert and trusted advisor, coding side-by-side to drive projects to completion on their infrastructure. Collaborate with Product, Research and Applied teams to ensure seamless customer experiences, project success and actionable product feedback Contribute to internal knowledge bases, codifying best practices and sharing insights gained from customer engagements to scale the Forward Deployed Engineering function. You’ll thrive in this role if
Other cities to consider
More places hiring for this role
Get new experienced field service representative jobs in United States by email
Daily job updates · Unsubscribe anytime