Jobs in United States

Platform Operations Specialist in United States

3,618 active opportunities · Updated October 2026

Explore current platform operations specialist jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

S
📍 San Francisco, CA, United States· Full-time
✓ Quality checkedCompany trend -96.2%

Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. The Role: SoFi's Cyber Defense organization is looking for an Offensive Security Lead to mature and grow our Penetration Testing and Red Team functions. This is a hands-on leadership role for someone who has spent years both doing the work and building the program around it — someone equally comfortable running a red team engagement against a critical banking platform and designing the operating model that lets a small team of offensive operators keep pace with a fast-growing fintech. A defining part of this role is modernizing how the team scales. We're looking for a leader who has already built and implemented AI-assisted penetration testing and red teaming programs — using AI tooling to accelerate reconnaissance, exploit development, attack-path analysis, and reporting — and who can bring that experience to bear on the program. You'll own the strategy, staffing, tooling, and execution quality of both disciplines, report into Cyber Defense leadership, and act as a trusted advisor to engineering, product, and risk partners across the company. What You’ll Do: Lead and unify Penetration Testing and Red Team into a single, cohesive Offensive Security function — shared standards, shared tradecraft, shared reporting, distinct missions. Set and execute the offensive security roadmap, aligning testing scope

AWSAzureGCPRest
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -88%

About the Team The Ona team at OpenAI is helping build the software factory for the enterprise. We build infrastructure that enables AI agents to work in secure, customer-controlled cloud environments, with the context, tools, and controls they need to make progress across the software lifecycle—beyond a single developer’s laptop or active session. Our focus is helping enterprises move from experimenting with agents to using them reliably in production. That means solving challenging problems in cloud environments, orchestration, security, and collaboration, while making the experience straightforward for the people directing and reviewing the work. We’re a team that values initiative, close relationships with customers, and exceptional engineering craft. We take ownership, learn quickly, and communicate directly and kindly. About the Role We’re hiring backend-focused Product Engineers across our platform and security product teams. You’ll build infrastructure and customer-facing workflows that let developers and AI agents work reliably in parallel. You’ll work primarily in Go on APIs, complex networking, development environments, and orchestration for long-running tasks. You’ll own outcomes from understanding a user’s problem and choosing an approach through shipping, operating, and improving the solution, working closely with frontend, infrastructure, and security engineers. In this role, you will: Work directly with customers to build developer and security workflows, from getting a project running to investigating findings, reviewing agent-generated changes, and verifying fixes. Build Go services and APIs for provisioning cloud environments, running agents in customer infrastructure, and integrating with source control, CI, and other developer tools. Design reliable orchestration for long-running, parallel work, including durable state, retries, cancellation, and recovery. Build security into execution workflows through clear permissions, credential handling, is

AWSAzureGCPKubernetes
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -94.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake runs large scale cloud infrastructure to deliver its own service — production and internal deployments, Kubernetes fleets, CI/CD, etc. Our cloud spend is in billions of dollars per year. The Cloud Efficiency team builds a unified, self-serve cloud efficiency platform along with AI skills and agents that makes spend observable, attributable, governable while driving recommendations and optimization of our cloud spend. AS A SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Design, develop, and maintain scalable platform for resource ownership registry, usage attribution, utilization measurement, and cost modeling. Build AI agents, tools and automation to enhance system monitoring, alerting, and root cause analysis. Improve and optimize data ingestion, storage, and query efficiency for cloud utilization, cost and efficiency data at scale. Collaborate with teams across Snowflake to understand attribution and observability needs and implement solutions that improve operational visibility. Contribute to open-source and industry best practices in monitoring and distributed systems monitoring. Ensure high availability, reliability, and performance of team-managed platforms by participating in on-call rotations and incident management. Partner with Finance, Product and Engineering

PythonJavaAWSAzure
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend 0%

From $224K/yr

Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Solutions Engineer (Enterprise Pre-Sales) Secure Every Identity, from Human to AI Agent About the Role At Okta, we believe that identity is the foundation of security and digital transformation. As a Senior Solutions Engineer, you will be the trusted technical advisor to our largest, most complex enterprise customers and a critical driver of our sales organization. We are looking for a highly strategic Pre-Sales professional whose consultative expertise, emotional intelligence, and enterprise sales acumen are their defining strengths. While technical agility is required, your primary focus will be owning the technical sales cycle by anchoring technical features to positive business outcomes. You will partner closely with Enterprise Account Executives to uncover top-of-mind business challenges—specifically around mitigating risk, reducing costs, and driving operational efficiency. By establishing value-based conversations, you will prove how Okta’s independent, neutral, and end-to-end identity platform can transform their architecture and secure the "Tech Win" on large-scale deals. What You'll Be Doing Master the Discovery Process: Leverage exceptional active listening to dig deep into customer pain points. You will uncover the 'why' behind the initiative, focusing on how Okta's vast pre-built integrations and scalable architecture can solve their most complex, diverse environmental challenges. Navigate Complex Organizations: Translate highly technica

AWSGitRestMachine Learning
N
📍 New York, New York, United States· Full-time· Remote
✓ Quality checkedCompany trend -88.6%

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: As a Forward Deployed Architect on Notion’s Services team, you will help customers turn AI ambition into operational reality and unlock the full potential of our platform. You will partner directly with customers to design, build, and launch AI-native workflows that drive adoption, usage, and measurable business outcomes. This role sits at the center of Notion’s workflow-led GTM motion. You will partner directly with customers to understand their business goals, operating processes, technical environment, governance needs, and adoption barriers, then translate that context into scalable Notion architectures that drive measurable value. You will be hands-on across discovery, solution design, workspace architecture, workflow build, AI implementation, launch planning, adoption and handoff. You will work deeply with customers to define what Notion should own, where it should integrate with existing systems, and how Notion can become a durable operating layer for their teams. The best candidates combine strong consultative discovery, systems thinking, technical fluency, AI-native workflow design, implementation judgment an

N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

We are developing advanced multi-rack, multi-tenant AI/ML datacenters with NVIDIA GB200, and upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the cloud-native stack for datacenter products like GB200. In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex scheduling challenges across racks, tenants, and clouds as part of the CSP engagements team. What you’ll be doing: Perform deep-dive debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime issues, device-plugin crashes, RDMA/IB fabric anomalies, etc. Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom micro-services that expose new GPU capabilities. Drive joint architecture reviews and “whiteboard” sessions with CSP and internal platform teams; convert findings into RFCs and upstream pull requests. Create reproducible testbeds (Helm/Ansible/Terraform) that mirror customer environments; automate validation and benchmark suites. Deliver technical collateral-design docs, how-to guides, demo scripts-and present at customer on-sites, KubeCon, and SlurmUG. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Strong source-level expertise in Kubernetes internals (scheduler, CRI/CNI/CSI, operators) and Slurm (federation, power-save, plugins). Hands-on experience integrating next-gen GPUs (Blackwell/GB200/GB300) or comparable accelerators into containerized clusters. Proven track record debugging large-scale, cloud-native stacks across ne

PythonKubernetesArtificial IntelligenceAI
I
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -90.5%

From $141K/yr

Quick readStrong listing-quality and freshness signals

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Why this role is on the menu Instacart's Talent Team is at an inflection point in how we manage external workforce: our contingent population has grown steadily, and there is now an opportunity to build and mature the governance, policy, and program infrastructure to match that growth. We're hiring a Lead Contingent Workforce Program Manager to inherit our newly launched MSP/VMS platform (Beeline + Monument) and turn it into a fully governed, scalable program — with clear policies, a trusted supplier network, and a compliance-first operating model. This role sits at the intersection of vendor management, workforce planning, and risk, and reports to the Head of Talent, working closely with Finance, Legal, Accounting, and IT. Twelve months from now, this person will have built a comprehensive playbook giving Instacart even greater visibility and insight into its cont

AIGoRustSap
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Requirements: 5+ years of experience writing high-quality production code Experience building high-performance distributed systems at a large scale (the more battle scars, the better) Strong cloud skills Strong knowledge of low-level operating system foundations (Linux kernel, file systems, containers, etc.) Experience with performance engineering (tell us a story of when you shaved off a few milliseconds!) Ability to work in-person in our NYC or SF office. Prior experience with Rust is nice to have, but not required. Ability to participate in on-call rotation and respond to production incidents.

LinuxRestAIGo
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a Detection & Response Engineer to build the systems that help us identify, investigate, and respond to threats across our platform. This is an engineering role focused on automation. You'll build detections, investigation tooling, and response capabilities that scale with our infrastructure, using AI where it meaningfully improves signal, investigation speed, and operational effectiveness. You'll work closely with infrastructure, platform, and security engineers to ensure every incident makes the platform more resilient. What You'll Work On: Detection Engineering Design and build high-fidelity detections for attacks, abuse, and anomalous behavior across our infrastructure and production systems Continuously improve detections based on telemetry, threat intelligence, and lessons learned from incidents Improve visibility across cloud infrastruc

SQLKubernetesGitLinux
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc

AWSAzureGCPKubernetes
C
📍 Washington, Washington, DC, United States· Full-time
✓ High-confidence listingCompany trend -79.2%

From £230K/yr

Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Senior Account Executive, Federal Defense and Intelligence (Washington, DC) Why this role? As a Senior Account Executive at Cohere you will have ownership of the full sales cycle - from identifying leads to closing deals, with Defense and Intelligence customers in the United States. In this role you will be focussed on business development activity within the Department of War, including Combatant Commands and some elements of the broader national security community. We’re looking for an approachable and compelling communicator who loves working with operators and analysts to uncover their needs and feels comfortable developing tailored value propositions around how Cohere’s platform can help them drive mission outcomes. You’ll lay the foundation for Cohere’s growth by owning your territory and collaborating with teammates across Customer Success, Sales Development, Marketing, and Sales Engineering. You’ll be the voice of the field and help our Product and Engineering teams prioritize the Cohere roadmap with customer-centric care. In this highly self-directed role, you will thrive in a quickly evolving environment. As a Senior A

GitAIGoRust
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by leading the design of high-performance, scalable and reliable machine learning systems? Do you want to set technical direction and help shape the next generation of AI platforms powering advanced NLP applications? We are looking for a Lead Member of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will provide technical leadership across multiple teams, driving the architecture and strategy for deploying optimized NLP models to production in low latency, high throughput, and high availability environments. You will serve as a key point of contact for customers, leading the design of customized deployments to meet their specific needs, and mentoring engineers to raise the technical bar across the team. You may be a good fit if you have: 8+ years of engineering experience running production infrastructure at a large scale, with a track record of technical leadership Demonstrated experience leading the architecture

AWSAzureGCPKubernetes
C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 We are seeking a highly technical and experienced Engineering Manager to one of our business critical engineering team. This role is ideal for a hands-on leader who thrives in a fast-paced environment, has a deep understanding of collaborative document editing technologies, and is passionate about building scalable, high-performance software. As the Manager, you will oversee the development and delivery of innovative features, ensure technical excellence, and mentor a team of talented engineers. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our product. Operational Excellence : Ensure the stability and performance of the Docs platform, proactively addressing technical debt and optimizing system architecture. Qualifications: Technical Expertise : Proficiency in

Node.jsAWSMachine LearningAI
P
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -75.7%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. About the Team: The Security Governance, Risk, and Compliance (GRC) team is part of Plaid’s security organization, focused on enabling the business by proactively managing information security risks and maintaining effective controls. Our mission is to reduce the likelihood and impact of security risks while operating a robust assurance program that builds trust with our customers, consumers, and data partners. We partner closely across the company to ensure Plaid’s platform remains secure, resilient, and aligned with industry and regulatory expectations. The Security Contracts workstream is a core part of our Security Assurance and Trust Enablement program — ensuring Plaid's contractual security obligations with customers and data partners are defensible, consistent, and never a bottleneck to deal velocity, all while building trust. About the Role You will own Plaid’s Security Contracts workstream end-to-end—the DRI for how security contract reviews get done, how fast they move, and how the program improves over time. You will review security provisions in customer MSAs, DPAs, and security addenda, identify unacceptable clauses, and provide Legal and GTM with the actionable security feedback they n

AWSAIGoRust
P
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -75.7%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely. Release Engineering owns the path from merge to production, including Plaid's zero-touch deployment system, progressive rollouts, metric-gated analysis, and automatic rollback. Our goal is to make safe shipping the default for every product team. As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale Plaid's reliability practices across product engineering. You'll architect our SLO and error-budget programs, drive the adoption of progressive delivery, and ensure new products are production-ready. By partnering across product and platform teams, you'll translate complex production needs into intuitive, self-service tooling. This is a hands-on technical leadership role where you'll shape the future of our deployment systems—ensuring they remain fast and safe even as AI-assisted development increases code velocity. What excites you Lead the expansion of reliability standards across product engineering, converting foundational infrastructure into lasting operational habits and tooling. Architect and manage the SLO and error-budget

AWSKubernetesAIGo
🔔

Get new platform operations specialist jobs in United States by email

Daily job updates · Unsubscribe anytime