About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're hiring a Compute Strategy and Operations lead to own how Modal plans for and acquires GPU and CPU capacity. You'll size our infrastructure needs ahead of demand, source supply across hyperscalers, neoclouds, and datacenter operators, and negotiate and close the contracts to secure it. The compute you secure directly determines what Modal can sell and build. In this role, you will: Own end-to-end procurement of GPU and CPU capacity across hyperscalers, neoclouds, and datacenter operators Build and maintain a strong pipeline of supplier relationships Evaluate supply options on price, availability, hardware specs, networking capabilities, and SLA terms Negotiate and close contracts: reserved capacity agreements, spot arrangements, MSAs, DPAs, and order forms Work closely with our engineering teams to translate technical requirements into procurement specs Track
Jobs in United States
Lead Cloud Operations Engineer in New York
273 active opportunities · Updated October 2026
Showing
15 jobs
Explore current lead cloud operations engineer jobs in New York. Filter by work mode, employment type, experience, department, date posted and distance.
$145K – $170K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. CLEAR is seeking a Senior Security Operations Analyst III to join our SOC team to help strengthen our ability to detect, investigate, and respond to evolving security threats. In this role, you’ll lead complex investigations, improve CLEAR’s threat detection and response capabilities, and serve as a trusted security partner while helping develop the analysts and program around you. What you'll do: Lead complex investigations of security events across corporate networks, endpoints, data centers, cloud environments, and other critical systems, driving incidents from initial analysis through escalation and remediation Develop, tune, and optimize threat detection logic across SIEM, EDR, and other security platforms, proactively identifying coverage gaps, reducing false positives, and improving the fidelity of security alerts Partner with Engineering, Infrastructure, and other teams to investigate threats, identify root causes, communicate risk, and drive timely remediation and improvements to CLEAR’s security posture Apply threat intelligence, data, automation, and AI-enabled tools to identify emerging attack patterns, accelerate investigations, improve detection workflows, and strengthen decision-making while applying sound security judgment Serve as a subject matter expert and escalation point for other analysts, mentoring junior team members, sharing knowledge, and helping establish scalable processes, playbooks, and standards for threat detection and analysis Continuously evaluate CLEAR’s detection coverage against the evolving t
$275K – $350K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We are seeking a strategically-minded, technology-focused, and customer-centric Engineering Manager to lead one of our Infrastructure teams here. You will lead a team responsible for building, operating, and scaling the cloud infrastructure and platform systems that underpin CLEAR’s services, ensuring reliability, performance, and security across our environments. A successful candidate brings strong experience in cloud infrastructure, distributed systems, and operational excellence, along with a solid foundation in software engineering. You are an effective communicator who can lead complex infrastructure initiatives from inception through delivery, and thrive in fast-paced environments. This role requires a focus on building resilient, scalable systems, driving automation, and leading and developing high-performing engineering teams. What you'll do: Hire, develop, and grow engineering talent through coaching, mentorship, performance management, and career development planning Set clear goals and expectations, provide regular feedback, and foster accountability across the team Own and execute the roadmap for cloud infrastructure and platform engineering, and reliability initiatives Design, build, and operate a scalable, secure, and highly available cloud platform infrastructure Drive automation across infrastructure provisioning, deployment, and operations to improve efficiency and reduce manual overhead Establish and enforce best practices for system reliability, observability, incident response, and disaster recovery Partner with eng
SUMMARY STATEMENT We are looking for a Solution Architect to design the technical solutions behind our client engagements and give delivery teams a clear, workable path from concept to production. You will work across enterprise data, software applications and GenAI - translating complex business problems into practical architectures that delivery teams can build and scale. This could include architecting an agentic workflow for clinical operations, a conversational analytics product grounded in enterprise data, or an AI-enabled decision platform for commercial teams. You will work directly with clients, define the architecture, test the most important technical decisions yourself and establish the foundations for successful delivery. This is an architecture-first role with meaningful hands-on engineering: you will stay close enough to implementation to prove the architecture works and support it through production delivery, without becoming the primary engineer for every component. You will also help shape the reusable patterns, technical standards and accelerators behind Lynx’s growing AI-native life sciences practice. KEY RESPONSIBILITIES Solution Architecture Own the end-to-end solution architecture for client engagements, including data models, system design, integration patterns and technology choices. Translate business requirements into clear technical designs and implementation paths that delivery teams can build from. Design solutions spanning enterprise data, APIs, applications, cloud platforms and GenAI capabilities. Lead technical discovery with clients: understand requirements, assess existing systems and identify dependencies, constraints and delivery risks. Present architectural options and trade-offs clearly to technical teams, business stakeholders and senior leaders. Make pragmatic decisions across build speed, cost, scalability, security and maintainability. Review key implementation decisions and remain
From $156K/yr
As a Product Manager – IaC Detection, you will define, build, and launch capabilities that proactively detect infrastructure issues in code (e.g. Terraform, Helm) before they can be deployed into production and escalate into production incidents. The Infrastructure Monitoring team has pioneered shift-left detection in the industry with Bits Infrastructure Operations , and we’re looking for a Product Manager to expand this capability to a broader set of use cases Customers (and thus developers) are increasingly standardizing on IaC tools to deploy and maintain ever-growing infrastructure in the cloud. At the same time, SREs and Infra teams struggle with an increasing number of production incidents. By shifting-left and identifying high-impact infra changes before they are deployed, we help reduce production incidents, reduce waste, and free up SRE time to focus on value-added tasks. You will own the roadmap to expand IaC detection to a broader set of use cases, including cost detection, blast radius impact, as well as configuration changes on infrastructure powering applications like nginx, postgres and more. You’ll partner closely with Engineering, Design, and customers to build and iterate on the roadmap, build product market fit, drive customer adoption (including internal usage), and focus on coverage and correctness of the AI system. This is an opportunity to lead an initiative at the intersection of AI, infrastructure operations, and autonomous observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the product roadmap for IaC Detection, enabling customers to proactively detect and catch high-impact infrastructure and configuration changes before they are deployed into production and escalate into incidents. Define the end-to
From $244K/yr
Role Summary: Datadog is seeking a Staff Software Engineer to help shape the future of our Bring Your Own Cloud (BYOC) Logs offering by unifying observability pipelines with log management software that customers deploy and manage in their own infrastructure. This role will focus on building and scaling systems that process, route, and store high-volume observability data within customer-managed infrastructure. You will operate as a hands-on technical leader, driving architecture, cross-team delivery, and product direction across a complex and evolving space. This is a high-impact opportunity to influence product strategy, mentor engineers, and solve deeply technical challenges at scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Make customer-controlled deployments feel like a managed Datadog product: deployment, upgrades, configuration, observability, diagnostics, reliability, and secure operation across diverse customer cloud environments Build and scale high-throughput systems for log processing, routing, and transformation across distributed environments Lead cross-team initiatives, aligning engineers, product managers, and stakeholders to deliver complex, multi-team projects Design and implement software that runs reliably that customers deploy and operate within their own cloud infrastructure. Improve system performance, scalability, and cost efficiency through thoughtful trade-off analysis and capacity planning Contribute hands-on to critical code paths, debugging, and deployment challenges in customer environments Who You Are: You have significant experience building software that is installed, deployed, and operated in customer environments rather than only as a fully managed SaaS service. You have strong expertise in distributed systems,
From $131K/yr
Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're hiring a VP of Finance to build the finance function from the ground up as our first full-time finance hire. This is a high-impact role for someone who thrives at the intersection of strategic thinking and hands-on execution. We are looking for someone who can architect the systems and processes that will scale with Modal, partner closely with the founders and executive team, and grow into the company's CFO. You'll report directly to the CEO and collaborate closely with our BizOps, GTM, and Product teams. In this role, you will: Build and maintain Modal's operating model, tying financial performance to company KPIs and resource allocation Lead all budgeting, forecasting, and long-range planning processes, and develop the reporting infrastructure that gives leadership and the board clear, timely visibility into the health of the business Partner with the found
From $154K/yr
Datadog’s Implementation Services team helps customers implement and deploy Datadog quickly and successfully. Our team of architects leads the discovery, design, build, and launch of the Datadog platform to help customers accelerate time to value and get the most out of their investment. As a Senior Services Architect focused on Security and Cloud SIEM, you will help customers design, implement, and operationalize Datadog’s security capabilities across cloud, infrastructure, application, and log data sources. You will lead structured, outcome-driven professional services engagements delivered through a day-based professional services delivery model, partnering directly with customers through co-development working sessions, architecture workshops, implementation planning, and operational handoff. This role is ideal for someone who combines customer-facing consulting experience with strong cybersecurity knowledge, hands-on Cloud SIEM implementation skills, and an understanding of security control frameworks such as NIST 800-53, the NIST Cybersecurity Framework, CIS Controls, MITRE ATT&CK, SOC 2, PCI, HIPAA, ISO 27001, or similar standards. At Datadog, we place value in our office culture — the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Design and guide execution of Datadog implementations, focusing on Security and Cloud SIEM deployment, including discovery, requirements gathering, technical architecture, deployment planning, and launch. Partner with customers to map security requirements, controls, and monitoring objectives to Datadog capabilities, including frameworks such as NIST 800-53, NIST Cybersecurity Framework, CIS Controls, MITRE ATT&CK, SOC 2, PCI, HIPAA, ISO 27001, or similar standards. Advise customers on security data strategy, including log source prioritization, parsing, normalizat
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. What You’ll Do Regional & Executive Leadership Own the overall Americas Channel Sales strategy, operating model, and results across Enterprise West, Enterprise East, LATAM, and Federal/SLED. Build, lead, and scale the Americas channel team hiring, developing, and retaining 4-5 regional channel ICs plus evolving Americas-wide scope. Establish clear priorities, coverage models, KPIs, and operating rhythms aligned to global Postman objectives. Act as the senior leader and voice for Americas Channel within Postman. Partner Ecosystem Strategy Define and execute the Americas partner strategy across SIs, resellers, distributors, and technology alliances. Build a scalable, services-capable partner ecosystem (WWT, SHI, CDW, Insight, SoftwareONE, Carahsoft, plus FSIs BAH, Deloitte Federal, Accenture Federal Services, Leidos, SAIC, CGI Federal, GDIT, CACI plus LATAM partners including Caylent, Mission Cloud, and regional SIs). Ensure partners are enabled, certified, and accountable for sourcing, selling, and delivering Postman Enterprise. Revenue & Pipeline Ownership Own partner-sourced and partner-influenced pipeline and revenue acro
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for a strong technical lead to guide the engineers designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. You'll lead the team responsible for Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll own the full lifecycle of a machine, from accepting and benchmarking new hardware from a growing set of providers, to network bring-up, kernel and image management, GPU and disk health tracking, and automated remediation of unhealthy hosts. You'll manage a team of 3–8 engineers while staying hands-on across the stack which involves BMCs, firmware, PXE, bootloaders, Linux networking, drivers, and distributed control-plane services, and you'll shape our long-
In this competitive market, the Group Product Manager for Cloud Security will play a mission-critical role in providing product and strategy leadership to grow Datadog’s market share through differentiation, innovation and compelling customer value. This leader will lead a talented and growing team of product managers and work with world class engineers to build and grow multiple Cloud Security products that play an essential role for our customers’ cloud security programs, and growing Datadog into a security industry leader. You Will: Run and grow multiple Cloud Security products to meet revenue and business targets with the goal of building a multi-hundred million dollar annual business. Lead and own product strategy and roadmap for accountable security products, fully aligned to revenue and business goals and with compelling differentiation and customer value. Ensure predictable roadmap execution across direct and partner teams to achieve product and business outcomes required to meet the revenue and business goals. Analyze and develop pricing and packaging strategies to maximize revenue through attaching deep understanding of market dynamics and other strategic leverage points. Drive GTM strategy with GTM partner teams to achieve revenue and business goals Fulfill the role as the product-leader representative across accountable product lines with analyst, press and customer communities Manage and grow individual contributor (IC) Product Managers, ensuring they are motivated, delivering high quality work and finding high fulfillment. Model and contribute to a culture of learning and collaboration across product management and the broader organization Raise the bar of PM leadership through leadership in strategy development, roadmap planning and execution, GTM and overall product-leader responsibilities. You Are: An experienced leader in Product Management with demonstrated business growth and customer adoption success in the CNAPP, data and AI
From $244K/yr
We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—allowing for seamless collaboration and problem-solving among Dev, Ops and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team: As organizations rapidly invest in AI applications and build out AI labs, telemetry volumes are growing exponentially and costs are becoming unpredictable. From LLM interactions to agentic workflows, AI systems generate unpredictable streams of logs, driving up costs and making it harder to maintain efficient observability. These challenges are critical for organizations in regulated industries with strict data residency requirements, where data must remain within controlled environments. Datadog’s Bring Your Own Cloud (BYOC) team is reimagining what observability and security look like at petabyte scale in the AI era. The Opportunity: The Group Product Manager - Bring Your Own Cloud (BYOC) role is responsible for defining and bringing to market the next generation of telemetry analytics and insights capabilities in an AI-first environment. This role is highly technical and creative in nature as you will envision novel ways to enable customers to cost-effectively explore, analyze and report over petabytes of data through a welcoming and easy-to-use interface. You will partner with various teams to take advantage of BitsAI capabilities and surface critical insights on volume usage and retention for popular use cases. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead and grow a team
From $200K/yr
We are looking for a talented engineer to lead evaluation of startup acquisition opportunities in the AI, cloud and security space. You will drive product evaluations, prepare and manage technical architecture discussions with target groups in Product and Engineering and provide roadmap suggestions for M&A and investments for Datadog. You will be a key partner to Datadog’s C-level leadership and highly visible at the most senior levels of Datadog. The role is reporting into the Senior Director of Product Strategy and falls within the Product organization. We are looking for an innovative and strategic thinker who is passionate about the latest tech being developed by startups in the cloud, AI and security space. The ideal candidate enjoys researching and evaluating new technologies, works effectively with cross-functional teams, and communicates opinions concisely to our leadership team. Broad understanding of relevant Cloud Technologies and deep understanding of the full coverage of Datadogs current offerings is necessary. The Corporate Development team is small and values authentic, strong-willed individuals who think creatively and proactively. This role leads technical due diligence from a product and architecture perspective across our acquisition pipeline. You'll scope and stand up proof-of-concept and sandbox environments to stress-test candidate products, then give an honest, unvarnished view of their quality and depth - the kind of assessment that holds up regardless of deal momentum. You'll assess technical architecture, flag the risks and open questions that matter most early, and turn that into a clear post-acquisition integration path. Working closely with engineering, you'll keep the evaluation focused on what's actually decision-relevant, then translate the findings into strategic recommendations for leadership and help carry the integration through by partnering with the right people on the other side. What You’l
From $192K/yr
As Engineering Manager for Threat Detection, you will lead a high-performing team that powers Datadog's detection program. Threat Detection is the organization responsible for keeping Datadog ahead of an evolving threat environment: closing coverage gaps faster, raising the bar on signal quality, and shipping detections that hold up under the scale and complexity of cloud-native infrastructure. Your team will combine direct detection expertise, platform engineering, and applied AI to ship detections at a pace and scale traditional rule-writing alone cannot match. Examples of what your team will work on include detection-authoring agents, the detection platform that powers every rule in production, coverage analysis, alert triage and response automation, and the evaluation infrastructure that holds these systems to a high bar of fidelity. Detection authorship is a shared responsibility across the organization, and your team will contribute both by building the systems that scale our authoring capacity and by writing detections directly when their domain expertise is the right tool. You will partner closely with our Security Incident & Response Team (SIRT), Cyber Threat Intelligence (CTI), AI Engineering teams, and Datadog's broader Security organization. This is a high-impact leadership role: you will grow a team of security and software engineers responsible for building and executing our detection and AI strategy. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the strategy, roadmap, and execution of Datadog Security's shift to AI-accelerated detection and response. Drive development of high-fidelity detections as a shared responsibility across the organization, ensuring your team's systems and direct contributions raise the bar on coverage and
Other cities to consider
More places hiring for this role
Get new lead cloud operations engineer jobs in New York, United States by email
Daily job updates · Unsubscribe anytime