Jobs in United States

Platform Operations Specialist in United States

3,022 active opportunities · Updated October 2026

Explore current platform operations specialist jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

B
📍 Berkeley, United States
✓ High-confidence listingCompany trend +515.8%
Quick readStrong listing-quality and freshness signals

Cloud Analyst (Mid-Level or Senior) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations is currently seeking a Cloud Analyst (Mid-Level or Senior) **Sign on Bonus Potential** to join the team in Berkeley, MO, Dayton Beach, FL; or Seattle, WA . The Cloud Analyst plays a key role in supporting the testing and execution of automated scripts across a range of cloud technologies, including Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP). In this position, the selected candidate will help ensure smooth, reliable cloud operations while providing thoughtful recommendations to improve performance, efficiency, and automation. This role also requires strong customer relationship skills, with a focus on clear communication, responsiveness, and delivering high-quality support that builds trust and confidence. Join a team driving innovation in cloud technology and automation, and help shape secure, efficient, and reliable digital solutions. Position Responsibilities: Develop and maintain comprehensive documentation that may include detailed technical documents Create process flows, business requirements, functional specifications, and user guides, with a strong emphasis on cloud-based systems Collaborate with stakeholders to gather, analyze, and validate business and technical requirements related to cloud infrastructure and cost management Design and document process flows to support cloud service management, price tracking, and operational improvements Support testing activities by creating and executing test plans, test cases, and automa

AWSAzureGCPSap
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $131K/yr

Quick readStrong listing-quality and freshness signals

At Datadog, People Operations is more than just human resources—it’s a data-driven team dedicated to constantly improving the way we hire, develop, and support our most valuable asset: our people. Our People Operations team are strategic problem solvers who work closely with leadership and employees to ensure that Datadog keeps scaling smoothly and remains a great place to work. Datadog is seeking a HRIS Manager who will be responsible for managing and supporting projects within People Technology. This hands-on technical role demands excellent knowledge of HR business processes and methodologies along with a strong analytical and reporting background. A successful candidate will have a solid understanding of Workday and the ability to focus on one or more of the functional areas in People Operations. This will include the ability to assess systems and business processes, coordinate with peers, project managers, and management on impacts to other Datadog systems or business processes. You will play a critical role in the continued deployment of new functionality, developing solutions and enabling the continued growth of Datadog. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the end-to-end product lifecycle for your pod's systems and processes—from gathering requirements and designing solutions through testing, delivery, and adoption. Serve as the Workday subject matter expert for your domain, advising stakeholders on configuration best practices, optimization opportunities, and governance. Assess when Workday is the right tool and when a specialized third-party platform better serves the business. You'll partner with stakeholders on those decisions and own the requirements that follow. Understand integration concepts

R
📍 San Fransisco, California, United States· Full-time
✓ High-confidence listingCompany trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Team & Role Ramp's procurement product is changing how companies buy. We've built the fastest intake-to-pay platform on the market — 3x faster than traditional P2P workflows, with AI that automates sourcing, contract extraction, approval routing, and invoice matching in one place. Customers save an average of 16% on vendor spend annually and eliminate 46+ hours of manual purchasing work every month. As Manager of Procurement Activation, you're the person who makes that real at scale . You'll lead a team of Procurement Activation Specialists working with Mid-Market customers, typically 100 to 999 employees, $50M to $100M in revenue, from discovery and configuration through go-live. Where your ICs work directly with CFOs and controllers to redesign how individual companies buy, you're building the system that makes those implementations excellent, repeatable, and fast. This is a founding role . There's no playbook to inherit — you'll write it. That means developing your people, building the motion from scratch, and becoming the primary p

ReactRestAIGo
R
📍 San Mateo, CA, United States· Full-time
✓ Quality checkedCompany trend -100%

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox is seeking a Hardware and Testing Lab Tecnician (Tech Support Specialist III) to support our automated testing device farm, a rapidly scaling on-premises hardware environment that enables automated testing across mobile, desktop, console, and specialty devices. This infrastructure is critical to Roblox's ability to ship quality software rapidly. This role sits within Corporate Engineering's Client Services team and is responsible for the physical and operational lifecycle of hundreds of test devices, from procurement and rack installation through daily health maintenance, break/fix support, and end-of-life recycling. You will work closely with Engineering to understand device configurations and testing requirements, while owning the hands-on execution that keeps these labs running reliably. Success in this role requires someone who takes full ownership of their work, holds themselves accountable, and can be counted on to follow through. This is a fully on-site role in our San Mateo, CA offices. You Will: Receive, asset-tag, inventory, and physically install devices (phones, tablets, Macs, PCs, and potentially consoles/VR) into IDF rack environments. Configure devices at t

AWSCI/CDGitAI
M
📍 Minnesota, United States of America, United States
✓ High-confidence listingCompany trend +1850%
Quick readStrong listing-quality and freshness signals

We anticipate the application window for this opening will close on - 2 Oct 2026 Careers that change lives start here. Medtronic is a global leader in healthcare technology with a Mission to alleviate pain, restore health, and extend life. Our 95,000 employees work across more than 150 countries to put patients first — developing innovative medical technologies that improve the lives of 72&#43; million patients each year. Your unique talents will help shape the future of healthcare while building a career grounded in purpose, growth, and impact. A Day in the Life We are seeking a committed professional to join our team, required to reside within the territory and drive to multiple accounts throughout the region. A valid driver's license is essential for this role, which also involves travel outside the territory, presenting opportunities for broader engagement. As one of three comprehensive portfolios at Medtronic, Neuroscience is dedicated to improving the lives of people living with neurological disorders, spine conditions, and chronic pain. Guided by our Mission—to alleviate pain, restore health, and extend life—we develop technologies and therapies that help people regain function, reduce pain, and return to the activities that matter most. Our Cranial & Spinal Technologies (CST ) operating unit advances surgical care for spine and cranial conditions through an integrated ecosystem of implants, navigation, robotics, imaging, and planning tools. Platforms like AiBLE enhance precision, efficiency, and outcomes for complex procedures worldwide. Check us out on LinkedIn: Medtronic CST At Medtronic, the Clinical Specialist, CST </

LinuxRecruitmentHR
D
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $214K/yr

Quick readStrong listing-quality and freshness signals

Datadog is seeking a strategic, visionary, and results-oriented Senior Director, Growth Marketing – Organic Growth (SEO/GEO/PLG) to lead our organic acquisition strategy across traditional search engines and emerging AI/LLM platforms. This leader will own the vision, strategy, and operating model responsible for driving measurable growth in organic traffic and inbound pipeline through content programs, off-page authority, and AI/LLM discoverability initiatives. In this role, you will lead a team of organic growth specialists, while partnering closely with Website Experience, Product Marketing, and Engineering teams. You will define Datadog's long-term organic growth strategy, establish investment priorities, and ensure the organization is positioned to win across an increasingly complex discovery ecosystem. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Own Datadog’s content-led SEO/GEO/PLG strategy, defining the topics, formats, and ecosystems that drive measurable growth in traffic and pipeline Lead, coach, and develop a high-performing team Identify and prioritize high-impact content opportunities across product areas and influence cross-functional teams to bring that content to life Map and optimize Datadog’s presence across the full ecosystem of LLM-ingested content (e.g., YouTube, Reddit, review sites) to improve AI-driven discoverability Design and execute a comprehensive off-page strategy, including link acquisition, digital PR, and authority-building initiatives Partner closely with the Website Experience team, who owns technical SEO, to ensure content is effectively surfaced, indexed, and performant Create and contribute to high-impact content (e.g., flagship pieces, new formats, or experimental channels), setting the standard for qualit

SQLGitAIGo
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors, software, and data center systems that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our U.S. engineering teams contribute to the hardware and software platforms that support the next generation of AI systems. The Opportunity We are looking for a recent graduate or early-career engineer to join the BMC Engineering team as a Graduate Firmware Engineer. You will develop low-level and embedded firmware that supports the operation, control, monitoring, and validation of advanced compute systems. You will work with experienced firmware, hardware, systems, and software engineers throughout the development lifecycle. The role combines hands-on implementation with automated testing, lab-based debugging, hardware bring-up, and analysis of interactions between firmware and the underlying platform. Start: September, 2027 Location: Austin, Texas, USA What You Will Do Design, implement, test, and maintain system and embedded firmware in C, C++, or Python. Take ownership of defined firmware features and deliver them from requirements and design through implementation, validation, and documentation. Develop and debug firmware in a Linux-based engineering environment using appropriate diagnostic tools and techniques. Create automated tests and scripts that improve firmware validation, test coverage, and engineering efficiency. Contribute to continuous integration and delivery workflows for firmware development and testing. Plan and conduct engineering experiments, analyze test data, and communicate findings clearly. Support lab setup, system configuration, hardware bring-up, and firmware validation on development platforms. Investigate firmware behavior and hardware-software

PythonLinuxArtificial IntelligenceAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world's most transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the future of AI computing. The Opportunity As Technical Services Director, you will lead the teams that operate and evolve Graphcore's engineering labs, high-performance computing (HPC) platforms, and data center environments globally. You will be accountable for reliable, secure, cost-effective infrastructure that supports demanding engineering, AI, silicon-development, and validation workloads. This role combines people leadership, infrastructure strategy, operational excellence, capacity and financial planning, procurement, and program delivery. You will partner with Engineering, Information Technology, Security, Finance, Facilities, Supply Chain, customers, and external suppliers. The position is based onsite in Austin and requires travel to company facilities, data centers, and supplier locations, including international travel. What You'll Do Lead, recruit, mentor, and develop the systems administration, lab operations, and technical services teams responsible for the facility supporting global Engineering and Research and Development. Own the reliability, efficiency, protection, safety, supportability, and continuous improvement of engineering labs, HPC systems, and infrastructure facilities. Establish service levels, operating standards, escalation paths, performance measures, monitoring, observability, automation, ticketing, and configuration-management practices. Translate engineering and customer requirements into infrastructure roadmaps, capacity p

LinuxAIGoExcel
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe

PythonAIGoDevOps
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Mechanical and Thermal Laboratory Technician Position Summary Graphcore is a globally recognized leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data center hardware that provide the specialized processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Austin, which will play a central role in Graphcore's work building the future of AI computing. Responsibilities The Mechanical and Thermal Laboratory Technician is a hands-on technical role supporting the development and validation of advanced AI hardware systems for data center environments. Working as part of a cross-functional engineering team, this individual will be responsible for executing mechanical and thermal laboratory testing, supporting product validation activities, prototype fabrication and assisting with troubleshooting and root-cause analysis of complex hardware systems. Requirements Associate degree in Mechanical Engineering Technology or a related technical field preferred. Equivalent combinations of education, training, and relevant experience will be considered, including experienced non-degreed candidates or candidates with degrees in unrelated disciplines. Minimum of 5 years of experience working in mechanical laboratories, machine shops, test labs, or similar technical environments. Experience with server hardware platforms and data center equipment. Knowledge of Direct Liquid Cooling (DLC) systems and their implementation in server environments. Experience operating forklifts, pallet jacks, and other material-handling equipment. Key Responsibilities Execute mechanical and thermal test plans to validate hardware designs a

AISEMTraining
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto

PythonCI/CDGitRest
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. The Mandate: Industry Disruption and 10x Scale: Vision: Product Foundry is Replit's internal engine for innovation . We move Replit beyond being a collaborative development environment to becoming the foundational operating system for the entire next generation of software, where AI Agents are the primary actors. The 10x Goal: Our purpose is to launch high-risk, full-stack 0→1 initiatives that define and establish entirely new, multi-billion-dollar product categories. We are seeking non-linear growth opportunities, fundamentally aiming to 10x Replit's value and addressable market by proving out unprecedented technical primitives and disruptive Go-To-Market strategies. The Audience: We build for the next generation of creators and high-leverage users and enterprises, equipping them with tools that enable them to build anything, anywhere . Candidates that do well here will certainly go on to build their own companies in the future! This is a high visibility role reporting to Execution Model: High-Agency Founding Teams Structure: We operate as a collective of in-house technical founders —not just specialized engineers. Initiatives are run by lean, autonomous squads built for velocity and maximum technical leverage. This model is centered around an Engineer DRI (Directly Responsible Individual) who maintains total ownership over the initiative's technical, product, and launch success, supported by fractional PM and Design resources. Cadence: We enforce rapid iteration and rapid market validation via 3-week sprints per initiative. This cadence forces fast deployment, immediate user feedback, and tight alignment with the internal betting table process, mirroring the intensity and speed of a lean startup. Required skills and

TypeScriptReactNode.jsAI
R
📍 Foster City, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -85.9%
Quick readStrong listing-quality and freshness signals

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are looking for an AI Agent Security Architect to function as the primary technical authority for Replit’s autonomous and AI agent security blueprint. In this critical role, you will design, implement, and maintain the runtime defense systems, guardrail frameworks, and sandboxing architectures that govern AI agents executing code, invoking tools, and reasoning across our platform. You will be a key technical contributor—leading high-impact AI security initiatives and bridging the gap between non-deterministic AI behavior and rigorous cybersecurity controls for both engineering and executive leadership. What You'll Do AI Agent Security Strategy & Technical Execution AI Agent Security Blueprint: Define the long-term vision and architectural patterns for securing autonomous agent workflows, Model Context Protocol (MCP) integrations, multi-turn reasoning loops, and multi-agent coordination. Runtime Guardrails & Policy Enforcement: Architect and deploy dynamic input/output guardrail systems, semantic firewalls, and real-time intent verification filters to prevent goal hijacking, system prompt leaks, and indirect prompt injections. Agent Execution & Tool Sandboxing: Partner with Infrastructure and AppSec teams to design secure, short-lived, micro-isolated environments (e.g., microVMs, WebAssembly, container sandboxes) where agents can dynamically execute code, run shell commands, and interact with host operating systems safely. Agentic Threat Modeling & Red Teaming: Conduct specialized threat modeling against non-deterministic systems. Lead automated and manual AI red-teaming initiatives to uncover vulnerabilities in RAG context pipelines, vector stores, and tool-calling interfaces. Identity

JavaScriptTypeScriptPythonJava
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors, software, and data center systems that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our U.S. engineering teams contribute to the hardware and software platforms that support the next generation of AI systems. The Opportunity We are looking for a computer engineering, electrical engineering, or computer science student or recent graduate to join the BMC Development team as a Firmware Engineering Intern. You will work with experienced engineers on low-level and embedded firmware that supports the operation, control, and manageability of advanced compute systems. This internship provides hands-on experience in firmware development, test automation, engineering experiments, and lab-based system testing in a Linux development environment. You will own clearly defined technical tasks with guidance from the team and contribute to production-quality engineering work. Type: 12-week summer internship Timing: May - August (exact dates to be confirmed) Commitment: Full-time What You Will Do Contribute to the design, implementation, and testing of system and embedded firmware. Develop and maintain firmware and supporting software in C, C++, or Python. Support firmware development and debugging in a Linux-based engineering environment. Create automated tests and scripts that improve firmware validation, test coverage, and engineering efficiency. Contribute to continuous integration and delivery workflows for firmware development and testing. Plan and conduct well-defined engineering experiments, record results accurately, and draw conclusions from test data. Support lab setup, system configuration, hardware bring-up, and firmware testing. Use debugging and diagnostic techn

PythonGitLinuxArtificial Intelligence
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui

PythonAWSAzureGCP
🔔

Get new platform operations specialist jobs in United States by email

Daily job updates · Unsubscribe anytime