Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an
Jobs in United States
Cloud Platform Administrator Sign On Bonus Potential in United States
698 active opportunities · Updated October 2026
Showing
15 jobs
Explore current cloud platform administrator sign on bonus potential jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
What you’ll do Own and evolve the Quality Management System (QMS) to support a regulated medical device development program, including design controls and DHF maintenance. Establish and enforce requirements traceability: user needs → design requirements → verification/validation artifacts and change control. Define and run the program-level V&V strategy (verification, validation, and test coverage), including test plans, protocols, reports, and acceptance criteria. Drive risk management activities (e.g., DFMEA / PFMEA, hazard analyses) and ensure mitigations are reflected in requirements and verification. Lead document control: reviews, approvals, training, retention, and audit readiness. Partner with engineering to make quality “native” to the dev workflow (automated testing, release gates, software configuration management). Prepare the program for audits and inspections, including hands-on audit leadership. What we’re looking for Senior experience leading quality for complex hardware + software products in a regulated environment. Deep familiarity with design controls, DHF, document control, risk management, and verification planning. Strong systems thinking and the ability to translate ambiguous product intent into testable requirements. Comfortable collaborating directly with multidisciplinary engineering (recon/ML, embedded, mechanical, EE, cloud). Useful experience Regulated product quality leadership (ISO 13485 / 21 CFR 820 or equivalent), including audit readiness and FDA-facing work. eQMS + document control fluency (e.g., Greenlight Guru) that integrates cleanly with modern engineering workflows.
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. The Product Security team is responsible for managing the security processes, policies and controls to secure Plaid’s developer and consumer facing products.The product security team is focused on areas like Application Security, Vulnerability Management, Secure Development Lifecycle, Penetration Testing and Cloud Security. We build the services and components that protect Plaid’s products. We move security "left" by engineering common libraries, modules, and workflows that make the secure path the easiest path for all Plaid engineers. Plaid is looking for a Product Security Engineer who is a builder to join our Product Security team. Unlike traditional Product security roles, this position is for a Senior software engineer who wants to solve security challenges at scale by designing and building production-grade services, libraries, and frameworks. Our goal is to make the "secure path" the only path for Plaid developers. The Role You will lead, design and develop security capabilities to manage vulnerabilities lifecycle and automate workflows to reduce KTLO toil. You will own, maintain, and build Plaid’s VM Orchestration service and build solutions to eliminate the entire vulnerability classes. You
From $204K/yr
The opportunity Datadog’s Infrastructure products help engineers understand and operate the systems their applications depend on. Our customers work in complex environments like Kubernetes and serverless, where infrastructure changes constantly, information is dense, and decisions about reliability, performance, and cost are closely connected. We’re looking for a Staff Product Designer to join Modern Compute, with an initial focus on Containers Autoscaling. Autoscaling helps engineering teams make better decisions about how their applications and infrastructure use resources. Designing these experiences requires making deeply technical systems understandable, helping customers act with confidence, and fitting into the tools and workflows they already use. The team is rethinking how workload and cluster autoscaling come together as a more coherent product experience. This includes how customers get started, understand recommendations, evaluate value, and safely apply changes across their environments. The work also connects to other parts of Datadog, including observability, Cloud Cost Management, permissions, and AI-assisted workflows. As a Staff Product Designer, you will help define that direction and lead the work from early problem framing through shipped product. You will partner closely with product and engineering, bring a high level of interaction and visual craft to complex workflows, and help raise the quality of design across Modern Compute. At Datadog, we place value in our office culture, the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to help our Datadogs find a work-life rhythm that works for them. What you’ll do Lead end-to-end product design for Modern Compute, initially focused on our Autoscaling product. Help define the product direction for an area that is still evolving, from early framing and exploration through detailed design and delivery. Design clear, trustwort
From $80K/yr
Datadog Sales Engineers help qualify and close opportunities with customers and partners. You will provide technical expertise through sales presentations, product demonstrations, and supporting technical evaluations (POCs). Sales Engineers have a voice with the product team to help prioritize features based on input from customers, competitors, and partners. Additionally, you will work with various teams to resolve customer concerns, escalate bug issues, and serve as an ambassador for our brand. If you want to join a friendly, passionate team with limitless potential, we’d love to meet you! At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Partner with the sales team to articulate the overall Datadog value proposition, vision and strategy to customers Continually learn new technology to build competitive knowledge, technical skill, and credibility Deliver product and technical presentations with potential clients Have a direct line of communication with the product team to collaborate on feature requests Help clients onboard the product and assist when they run into roadblocks. Think creatively about a wide variety of technical challenges during the pre-sales life cycle Who You Are: Enjoys a hybrid office culture and spending time in-person at our Boston office Has at least one year of experience in a technical sales and/or post-sales role Has a working knowledge of modern web development, DevOps monitoring, and cloud architecture/technologies Enjoys collaborating and teaming with others Demonstrates strong written and oral communication skills Comfortable and confident in delivering technical presentations/demos Can multi-task, manage time well, and be responsive to others Curious, eager to learn, and persistent in solving problems Able to si
About the Team At OpenAI, our User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We translate real-world user and operational signals into timely decisions, practical interventions, and improvements to our products and systems. This role will take on new, ambiguous, or underdeveloped operational risks and help mature them into scalable capabilities. We work across USRO and partner closely with Product, Engineering, Data Science, Product Policy, Legal, Safety, Support, and external vendors or partnership stakeholders. About the Role We are seeking a Senior Operations Analyst to take on complex, ambiguous safety and risk problems and turn them into practical operational solutions that can scale. This is a senior individual-contributor role for a versatile operator who is comfortable moving between queues, investigation, analysis, workflow design, hands-on execution, and cross-functional leadership. Depending on team needs, the role may focus on emerging-risk incubation, cloud deployment partnerships, or other new operational areas. You will be expected to move quickly, work hands-on, and create structure without waiting for perfect requirements or a large support team. The work starts with the problem, not a prescribed process. You may investigate unstructured user signals, stand up a lightweight workflow, build an AI-assisted tool, improve an existing operation, or help a new launch become operationally ready. The goal is to produce durable systems that other people can run, not simply complete a series of individual tasks. The portfolio will change with company priorities and may span established harm areas, emerging-risk incubation, cloud deployments and partnerships, device safety, or new product launches. Some hires may focus primarily on cloud deployment operations, including launch readiness, partner coordination, safety workflows, and operational monitoring. You will ty
About the Team OpenAI’s Governance, Risk, and Compliance team helps ensure security and privacy are grounded in how our products and systems actually operate. Assurance Operations partners with Security, Engineering, Infrastructure, Product, Privacy, and Legal to make controls provable, risk decisions explicit, and audit readiness a result of well-designed systems. About the Role We are hiring a technical, product-minded GRC builder who can own consequential audits while improving the control and evidence systems behind them. You will build a reusable common control framework, use Codex to automate assurance work, validate changing system scope, and turn repeated audit friction into measurable improvements. We are looking for someone who questions inherited assumptions, solves novel problems creatively, works closely with engineers, and makes the next audit easier by improving the underlying system. You’ll be responsible for: Lead external, internal, customer, and certification audit work from scoping through evidence review, fieldwork, remediation, and closeout. Build a common control framework linking risk, control intent, implementation, owner, system, environment, evidence, and applicable frameworks. Validate actual scope and ownership instead of assuming last year's controls, product boundaries, or evidence remain accurate. Use Codex to build and test evidence checks, control mappings, request triage, owner workflows, monitoring, and remediation reporting. Partner with engineers on cloud architecture, identity, logging, data flows, software changes, vulnerabilities, and control effectiveness. Design maintainable, permission-aware tools that preserve source provenance, human review, and evidence integrity. Reduce repeated requests and operational burden for control owners through measurable workflow improvements. Define roadmaps, decision rights, milestones, success metrics, and clear cross-functional escalations. We’re looking for someone with: Direct ownership
About the Team OpenAI’s Cyber team works to make frontier AI safe, trusted, and transformative for developers and enterprises. This team is building the security foundation for Codex: the native controls that govern what Codex can access and do, and the interfaces that allow customers and security partners to inspect, constrain, approve, and respond to Codex activity. Our goal is to make Codex secure by default, governable by enterprises, and interoperable with the security products customers already trust . This extends the existing product direction around tenant-scoped tools, guarded actions, approval systems, and scalable partner interfaces. About the Role We are looking for a deeply technical Product Manager to help build Codex security controls and the partner ecosystem around them. This role focuses on securing Codex itself : how identity, permissions, tools, MCP servers, repositories, secrets, networks, and high-impact actions are governed across Codex products. You will also help define standard interfaces through which authorized customer and partner systems can provide security context, inspect activity, return policy decisions, receive telemetry, and initiate bounded responses. You will work closely with Codex product and engineering, OpenAI Security and Safety, enterprise customers, and partners across application security, identity, cloud security, data security, infrastructure, and security operations. In this Role you Will Build native security controls for Codex Partner with engineering, design, security, and safety teams to develop controls for: Identity, roles, permissions, and tenant isolation. Access to repositories, files, tools, MCP servers, secrets, networks, and infrastructure. Read, write, execute, and deployment authority. Human and policy-based approvals. Prompt-injection and untrusted-content defenses. Audit trails, provenance, stop conditions, revocation, and rollback. Help establish a graduated authority model in which local, read-only
About the Team The Private Computing team works across product, engineering, security, and safety to build advanced privacy products and infrastructure at OpenAI. Our mission is to provide world-class security features to users so their private data remains private, even from OpenAI. We use technologies like confidential computing, trusted execution environments, and end-to-end encryption to ship product features across ChatGPT, the API, and our future consumer devices. About the Role We’re looking for software engineers to design, build, and scale novel privacy features and infrastructure across ChatGPT, API, and future consumer devices. In this role, you will: Ship fast while balancing difficult trade-offs in complex domains Build core abstractions for trusted execution environments and end-to-end-encryption Build product features for private inference and storage across ChatGPT, API, and future consumer devices Update build systems to increase trust and verifiability Integrate with safety and integrity infrastructure Operate systems at scale with high reliability, including an on-call rotation Collaborate with a diverse set of cross-functional teams across product, engineering, security, safety, policy, and legal You might thrive in this role if you: Care deeply about user privacy and security Have 5+ years of experience in professional software engineering Have experience building and scaling confidential computing or encryption technologies in production environments Have experience with Kubernetes and cloud orchestration systems Take pride in building and operating scalable, reliable, secure systems Can collaborate well and drive alignment in the face of difficult trade-offs Are comfortable with ambiguity and rapid change Workplace & Location This role is based in San Francisco, CA. We follow a hybrid model with 4 days a week in the office and offer relocation assistance to new employees. About OpenAI OpenAI is an AI research and deployment company dedicat
About the team The AI Deployment Engineering team is responsible for helping developers and enterprises safely and effectively deploy OpenAI technologies in production. We act as trusted technical advisors and thought partners for customers, working side by side with their teams to identify high-value use cases, design practical architectures, and move from prototype to durable deployment. Cybersecurity is one of the most urgent domains where AI can help. Security teams are under pressure to reason across code, logs, infrastructure, tickets, alerts, and vulnerability data faster than ever. As frontier models become more capable, organizations need deep technical guidance on how to evaluate, validate, and safely deploy AI systems in security-critical workflows. About the role We are looking for a Cyber AI Deployment Engineer to partner with customers and help them apply OpenAI models, APIs, Codex, and agentic workflows to real cybersecurity use cases. You will work with CISOs, security executives, application security leaders, SOC teams, security engineering teams, and hands-on practitioners to identify where AI can create measurable security outcomes. This is a customer-facing technical role for someone who can move fluidly between executive strategy, practitioner-level cyber depth, and hands-on solution design. You will help customers evaluate and deploy workflows such as secure code review, vulnerability triage, threat modeling, remediation, SOC and incident response workflows, detection engineering, cloud security, GRC automation, and security validation. You will collaborate closely with Sales, Solutions Engineering, Product, Engineering, Research, and Security to turn customer needs into safe deployment patterns, reusable field assets, and product feedback. This role is based in our San Francisco HQ. We offer relocation support to new employees. In this role, you will: Deeply embed with strategic customers as the technical lead for AI-enabled cybersecurity work
Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role We’re building the observability product for OpenAI—from scalable infrastructure to a rich, AI-powered UI. Our systems ingest over petabytes of logs and billions of time series metrics across our fleet. We're now layering intelligence on top—think agents that summarize SEVs, auto-generate dashboards, or help engineers debug through notebook-like UIs. We’re hiring software engineers across the stack—infra, backend, and product. You’ll join a small, gritty team building both foundational infra and novel internal tools to make OpenAI's production systems reliable, performant, and observable. What You’ll Do Own core observability infrastructure, including distributed logging, time series, and trace storage Build AI-native tools that help engineers detect, understand, and resolve issues autonomously. Contribute to UI experiences like dashboards, notebooking, or interactive debugging Collaborate closely with engineers, researchers, user ops, and other teams across the company to build the next generation observability product You Might Be a Fit If You: Have operated large-scale distributed systems in production. ( especially logging systems or some other time series databases) Thrive in ambiguous environments and roll up your sleeves to solve unscoped problems. Have full-stack chops or product sensibilities—you're excited to build real tools people use. Have strong fundamentals in systems, networking, and cloud infra (Kubernetes, AWS, etc). Bonus : built or contributed to observability systems (e.g. Prometheus, OpenTelemetry, etc). Why This Team We’re b
About the Team The Coding team is reimagining how software is built in the AI era. We build tools and workflows that help software engineers work faster, tackle more ambitious projects, and spend less time on repetitive tasks. AI has already transformed how code is written, but software engineering extends far beyond coding. Our mission is to apply AI across the entire software development lifecycle (SDLC) — from design and implementation to code review, testing, debugging, issue remediation, maintenance, documentation, and user support. The team is also responsible for developer-facing Codex experiences including the Codex IDE Extension and the terminal interface, which are used daily by developers ranging from individual open-source contributors to some of the world’s largest engineering organizations. The team also works closely with the open-source software community, building tools that help maintainers and contributors manage increasingly complex projects. We believe AI can make open-source development more sustainable by reducing the operational burden of reviewing contributions, triaging issues, maintaining quality, and supporting growing communities. By building the future of software development, we're helping advance OpenAI's mission of ensuring that the benefits of AI reach people around the world. About the Role We’re hiring a Full Stack Software Engineer to help invent the next generation of AI-powered software development workflows. “Full stack” in this role means much more than traditional frontend and backend development. You'll own complete product experiences, spanning user interfaces, workflow orchestration, agent and prompt design, backend systems, and cloud infrastructure. This is a highly product-oriented role. You'll work directly on the workflows developers use every day, identifying bottlenecks and rethinking how software gets built in a world where AI agents are active participants in the development process. The features you ship will inf
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
$293K – $405K/yr
About the team Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security that could scale to an extreme level of severity. Our work involves: Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework , and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. About the role As AI agents become more capable at software engineering, and automate more of our internal work, they could become a dangerous cyber threat. People in this role will help OpenAI prepare for security threats from advanced AI agent insiders. In this role, you will: Identify paths by which capable future internal AI agents could compromise OpenAI. Design security controls - focusing on measures with long lead times that benefit from advanced preparation. Stress-test defenses with AI agent evaluations and penetration tests You might thrive in this role if you: Are deeply technical across security and modern infrastructure, and are comfortable digging into the details of operating systems, cloud, containers, CI/CD, or distributed systems. Have strong software engineering skills and enjoy building prototypes yourself. Are interested in engaging with stakeholders and can do so effectively. Bonus: have experience securing cloud infrastructure, and are deeply familiar with core components of the AI stack. Compensation Range: $293K - $405K USD About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy
$216K – $240K/yr
About the Team OpenAI Finance is responsible for ensuring the organization is set up for success in pursuit of its mission. The Technical Accounting team plays a crucial role in helping OpenAI navigate complex, judgmental, and rapidly evolving accounting matters with rigor and clarity. We aim to bring both technical excellence and strong business partnership to some of the most novel accounting questions in the industry. About the Role As Senior Manager, Technical Accounting, Compute Infrastructure, you will lead the evaluation, documentation, and operationalization of complex accounting matters related to OpenAI's compute infrastructure, strategic investments, and other non-routine business activities.. This role sits at the intersection of U.S. GAAP technical accounting, infrastructure strategy, financial reporting, controls, and cross-functional execution. Key areas may include cloud compute arrangements, data center and colocation arrangements, lease accounting under ASC 842, power purchase agreements, strategic investments, consolidation evaluations under ASC 810, financial instruments, and other emerging or non-standard arrangements. This role is based in San Francisco, CA or remote. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead technical accounting analysis for complex, judgmental, and non-routine transactions under U.S. GAAP. Evaluate accounting implications for compute infrastructure arrangements, including cloud compute, data center, colocation, lease, PPA, infrastructure procurement, and related commercial arrangements. Partner with Controllership, Tax, Legal, FP&A, Procurement, Infrastructure, and other cross-functional teams to assess the accounting implications of new products, commercial arrangements, strategic transactions, and business initiatives. Prepare and review technical accounting memoranda, position papers, and other auditor-ready documentation.
Other cities to consider
More places hiring for this role
Get new cloud platform administrator sign on bonus potential jobs in United States by email
Daily job updates · Unsubscribe anytime