About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Requirements: 5+ years of experience writing high-quality production code Experience building high-performance distributed systems at a large scale (the more battle scars, the better) Strong cloud skills Strong knowledge of low-level operating system foundations (Linux kernel, file systems, containers, etc.) Experience with performance engineering (tell us a story of when you shaved off a few milliseconds!) Ability to work in-person in our NYC or SF office. Prior experience with Rust is nice to have, but not required. Ability to participate in on-call rotation and respond to production incidents.
Jobs in United States
Platform Architecture Manager in New York
386 active opportunities · Updated October 2026
Showing
14 jobs
Explore current platform architecture manager jobs in New York. Filter by work mode, employment type, experience, department, date posted and distance.
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role We're looking for an Engineering Manager to lead a team of highly experienced engineers building the infrastructure that powers Modal's serverless GPU platform. This is a hands-on leadership role — expect to split your time between technical contribution and people management depending on what the team needs. You'll set direction, remove blockers, and build a strong engineering culture as your team tackles hard problems in distributed computing, large-scale data handling, and performance optimization. Who You Are You're an experienced engineering leader who stays close to the work and builds alongside your team when it counts. You earn trust through technical depth, not title. You communicate clearly, help strong engineers move fast without cutting corners, and stay calm and pragmatic under pressure. You care as much about how your team gets to an answer as the answ
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We’re looking for an Infrastructure Security Engineer to design and secure the core systems that power our platform. This role focuses on building security directly into our infrastructure—from container isolation and orchestration to identity and secrets management in a multi-tenant, cloud-native environment. You’ll work closely with engineering teams to define secure primitives and ensure our platform is resilient, scalable, and trustworthy by design. This is a hands-on, deeply technical role focused on real systems, not compliance or policy. What You'll Do: Platform & Runtime Security Design and improve isolation mechanisms for multi-tenant workloads (containers, sandboxing, execution environments) Strengthen boundaries between customers, workloads, and internal systems Identify and mitigate risks in distributed, dynamic compute environments Container &
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for engineers with deep AI/ML and low-level systems experience who want to build the best technical support experience in the world. This isn't a traditional support role — it's an engineering role where you happen to be closest to our customers. You'll split your time roughly 50/50 between working directly with customers and shipping fixes, features, and automation that improve Modal for everyone. When you help a customer debug a training run, you'll also fix the underlying issue in the platform. When you notice ten customers hitting the same friction point, you'll build the tooling or automation that eliminates it entirely. This role is for people who solve problems, not people who answer tickets. The problems you encounter are deeply technical and arise from running some of the most demanding AI workloads in the world. You'll be a member of our eng
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a Detection & Response Engineer to build the systems that help us identify, investigate, and respond to threats across our platform. This is an engineering role focused on automation. You'll build detections, investigation tooling, and response capabilities that scale with our infrastructure, using AI where it meaningfully improves signal, investigation speed, and operational effectiveness. You'll work closely with infrastructure, platform, and security engineers to ensure every incident makes the platform more resilient. What You'll Work On: Detection Engineering Design and build high-fidelity detections for attacks, abuse, and anomalous behavior across our infrastructure and production systems Continuously improve detections based on telemetry, threat intelligence, and lessons learned from incidents Improve visibility across cloud infrastruc
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're building a platform that covers the whole life of an LLM: training it, deploying it, and observing it in production. We already run multi-node training, elastic inference, sandboxes, and distributed volumes, and we control the infrastructure underneath. We’re looking for research depth in post-training to sit alongside our systems and product work. What you'll do: We are looking for research scientists with a strong track record in reinforcement learning, machine learning, and foundation models, including large language and multimodal models, to join our research team. This role is well suited to candidates interested in improving existing methods and developing new techniques for large-scale model training, optimization, and inference, extending models to long-context and long-horizon tasks, and improving inference-time efficiency, reliability, and robustnes
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. About the Team: The Security Governance, Risk, and Compliance (GRC) team is part of Plaid’s security organization, focused on enabling the business by proactively managing information security risks and maintaining effective controls. Our mission is to reduce the likelihood and impact of security risks while operating a robust assurance program that builds trust with our customers, consumers, and data partners. We partner closely across the company to ensure Plaid’s platform remains secure, resilient, and aligned with industry and regulatory expectations. The Security Contracts workstream is a core part of our Security Assurance and Trust Enablement program — ensuring Plaid's contractual security obligations with customers and data partners are defensible, consistent, and never a bottleneck to deal velocity, all while building trust. About the Role You will own Plaid’s Security Contracts workstream end-to-end—the DRI for how security contract reviews get done, how fast they move, and how the program improves over time. You will review security provisions in customer MSAs, DPAs, and security addenda, identify unacceptable clauses, and provide Legal and GTM with the actionable security feedback they n
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely. Release Engineering owns the path from merge to production, including Plaid's zero-touch deployment system, progressive rollouts, metric-gated analysis, and automatic rollback. Our goal is to make safe shipping the default for every product team. As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale Plaid's reliability practices across product engineering. You'll architect our SLO and error-budget programs, drive the adoption of progressive delivery, and ensure new products are production-ready. By partnering across product and platform teams, you'll translate complex production needs into intuitive, self-service tooling. This is a hands-on technical leadership role where you'll shape the future of our deployment systems—ensuring they remain fast and safe even as AI-assisted development increases code velocity. What excites you Lead the expansion of reliability standards across product engineering, converting foundational infrastructure into lasting operational habits and tooling. Architect and manage the SLO and error-budget
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Team: The Security Governance, Risk, and Compliance (GRC) team is part of Plaid’s security organization, focused on enabling the business by proactively managing information security risks and maintaining effective controls. Our mission is to reduce the likelihood and impact of security risks while operating a robust assurance program that builds trust with our customers, consumers, and data partners.We partner closely across the company to ensure Plaid’s platform remains secure, resilient, and aligned with industry and regulatory expectations. Third-party ecosystem risk is a core part of how we keep Plaid safe—we vet the security of both the vendors we rely on and the customers and partners who connect to our platform, so trust runs in both directions. Role: You will run security risk assessments for Plaid’s third parties end-to-end—from intake and questionnaire through risk rating, findings, and tracked exceptions. You will assess the security posture of customers and partners onboarding to the platform with the same rigor we apply to vendors. You will keep the third-party risk lifecycle moving—risk tiering, reassessment cadence, remediation follow-through, and a clean, current risk register. You
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Security Engineering is the engineering function inside the Plaid security org that focuses on developing the industry-leading security systems and infrastructure. Security Engineering owns most of Plaid’s security-related infrastructure: secure data storage, key management systems, internal identity platform, internal authentication systems, internal permission management, and internal authorization service. We develop solutions across data encryption, key management, access control, and data loss prevention to protect sensitive consumer data. We believe in the Zero Trust security model and are always looking for ways to improve our authentication and access control platforms. About the role: You will develop security capabilities to secure Plaid infrastructure and sensitive data access. You will lead the team’s strategic planning in collaboration with the manager and other senior engineers. You will own, maintain, and build Plaid’s security infrastructure and services like IAM Gateway, Key Management System and Network Firewall. You will consult with product engineers to ensure Plaid services meet security standards. You will help educate and support other engineering teams to improve security in
From $81K/yr
We are Datadog's in-house product experts. The Technical Solutions team enables Datadog's worldwide growth by educating potential clients and ensuring that existing customers are happy and successful. Premier Support Engineers (PSEs) are primarily focused on assisting prospects and customers with any technical questions about Datadog. PSEs engage with Datadog’s Premier Customers via standard technical support channels, but are also involved with cadence calls, demos/presentations, conferences, and various side projects. You will work directly with Datadog’s Premier Customer base, and will be immersed in a fast-paced environment where you will be challenged, but will also immediately witness your contributions to Datadog and to our customers. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Respond to client requests (phone / chat / tickets) on our fast paced team while continuing to educate our clients on the use of the platform Develop relationships with our Premier Customers, working hand-in-hand to truly know their distinct environment Reproduce issues and dive into the 600+ integrations that Datadog works with Build out documentation and knowledge based articles for a variety of technology Drive product conversations based on needs and problems learned during client interactions Participate in routine health check meetings with Premier Customers Work from a Datadog office 3 - 5 days per week Who You Are: Experienced in multi-channel technical support at a SaaS company (2+ years of related experience) A tinkerer with some programming experience and a basic knowledge of Linux Self-motivated, detail-attentive, and have a desire for continuous learning A critical thinker who defaults to a client-centric approach A decision
From $131K/yr
At Datadog, People Operations is more than just human resources—it’s a data-driven team dedicated to constantly improving the way we hire, develop, and support our most valuable asset: our people. Our People Operations team are strategic problem solvers who work closely with leadership and employees to ensure that Datadog keeps scaling smoothly and remains a great place to work. Datadog is seeking a HRIS Manager who will be responsible for managing and supporting projects within People Technology. This hands-on technical role demands excellent knowledge of HR business processes and methodologies along with a strong analytical and reporting background. A successful candidate will have a solid understanding of Workday and the ability to focus on one or more of the functional areas in People Operations. This will include the ability to assess systems and business processes, coordinate with peers, project managers, and management on impacts to other Datadog systems or business processes. You will play a critical role in the continued deployment of new functionality, developing solutions and enabling the continued growth of Datadog. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the end-to-end product lifecycle for your pod's systems and processes—from gathering requirements and designing solutions through testing, delivery, and adoption. Serve as the Workday subject matter expert for your domain, advising stakeholders on configuration best practices, optimization opportunities, and governance. Assess when Workday is the right tool and when a specialized third-party platform better serves the business. You'll partner with stakeholders on those decisions and own the requirements that follow. Understand integration concepts
From $232K/yr
Datadog is entering a new chapter in how our product looks, feels and operates. We’re building a small, high-leverage Design Lab team to define that evolution, and we’re looking for a Staff Visual Designer to help author it. Our design system powers everything we ship. Now we’re ready to evolve it. As Datadog expands into AI-driven experiences and more complex product surfaces, the visual language needs to grow with it. We’re looking for someone who can help define that next phase, someone with taste, judgment, and the confidence to shape how the product feels at scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do Define and evolve Datadog’s visual language across product surfaces. Establish the visual quality bar for the platform. Lead exploratory and concept work that pushes our system forward. Translate bold visual ideas into scalable patterns that can live inside a design system. Partner deeply with Brand, Product Design, and Design Systems to unify expression across product and marketing. Influence executive stakeholders through clear, compelling visual storytelling. Prototype, test, and refine new patterns before they become systemized. Help shape how visual direction is integrated into our design process. You will report directly to the Design Director and work as part of a small, focused team defining the future state before it scales across hundreds of designers and engineers. Who You Are You have 10+ years of experience in visual and/or product design You have led a meaningful visual evolution or rebrand within a mature product ecosystem. You have strong taste and are comfortable defending a point of view. You know how to translate brand into scalable product systems. You think in systems, but you’re motivated by craft. Yo
From $232K/yr
Datadog is entering a new chapter in how our product looks, feels and operates. We're building a small, high-leverage Design Lab team to define that evolution, and we're looking for a Senior Staff Visual Designer to author it. Our design system powers everything we ship. Now we're ready to evolve it. As Datadog expands into AI-driven experiences and more complex product surfaces, the visual language needs to grow with it. This is the role that decides what that language is — someone with taste, judgment, and the confidence to set a direction and defend it. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do Define and evolve Datadog’s visual language across product surfaces. Lead exploratory and concept work, creating a range of distinct visual directions for the next generation of the product. Synthesize product context, research, references, and stakeholder input into clear visual principles and a coherent point of view. Establish the quality bar for typography, color, composition, iconography, imagery, motion, and visual expression across the platform. Create high-fidelity product exemplars and prototypes that make an emerging direction tangible. Partner deeply with Brand, Product Design, Motion, and Design Systems to test and refine the direction across different surfaces. Influence executives and senior stakeholders through clear, persuasive visual storytelling. Guide and critique the work of other designers, helping the visual direction remain coherent as it evolves. Help shape how visual exploration, critique, and craft are integrated into Datadog’s design process. You will report directly to the Design Director and work as part of a small, focused team defining the future state before it scales across hundreds of designers and engineers. Who Yo
Other cities to consider
More places hiring for this role
Get new platform architecture manager jobs in New York, United States by email
Daily job updates · Unsubscribe anytime