Jobs in United States

Engineering Maintenance Lead in United States

2,830 active opportunities · Updated October 2026

Explore current engineering maintenance lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

NVIDIA is seeking a strong technology leader to manage our Server Software Technical Program Management (TPM) team. This role is at the cross-section of execution and strategy, leading a team of Senior TPMs who drive the firmware and system software for NVIDIA's next-generation server platforms like DGX, MGX, and HGX. These platforms bring together the full power of NVIDIA GPUs, NVLink, InfiniBand networking, Grace CPUs, and our optimized AI/HPC software stack. This deep technical leadership role focused on the Software Development Processes that brings new server hardware to life. What you'll be doing: Lead a team of TPMs driving the technical software and firmware execution for NVIDIA's NPI (New Product Introduction) and sustaining engineering teams. Drive the end-to-end SDLC for low-level server components, including firmware (BMC, UEFI/BIOS), drivers, and system management software, ensuring alignment with hardware schedules. Collaborate closely with NVIDIA product management and hardware engineering teams to define release plans and program objectives. Build a strong connection and feedback loop between sustaining and NPI engineering teams to improve product quality and development velocity. Lead process improvement initiatives and help propagate SDLC standards across multiple engineering and TPM organizations. You will have the opportunity to interact with diverse technical groups, spanning all organizational levels. What we need to see: Bachelor of Science (or equivalent experience) or Master of Science degree in Computer Science, Electrical Engineering, or related field. 12+ overall years of experience developing and leading complex low-level or system software projects. and 7+ years of experience in a people management role. Deep understanding of system a

Artificial IntelligenceAI
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel

PythonKubernetesLinuxArtificial Intelligence
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

NVIDIA's DGX Cloud (DGXC) powers AI for strategic research and product workloads. The company seeks a Senior Technical Program Manager (TPM) to lead complex, cross-functional programs powering NVIDIA’s next-generation AI software platforms. In this role, you will drive software initiatives across platform services, cloud infrastructure, and system integration. The focus is on enabling scalable, reliable, and supportable software for AI workloads. You will be responsible for managing high-impact engineering programs within a dynamic, fast-paced roadmap, aligning priorities across teams, and ensuring timely, high-quality delivery. This role requires strong technical competence, a proactive approach, and the ability to operate effectively across multiple levels of the organization. This is a software-first TPM role. The ideal candidate has extensive experience managing software initiatives. They also understand the full-stack environment, including infrastructure dependencies, system bring-up, integration readiness, and operational needs to support software across stack layers. What You'll Be Doing: Lead end-to-end execution of software platform initiatives, including planning, execution, delivery, and operationalization. Work together with software, infrastructure, product, and operations teams to ensure alignment on goals, deliverables, achievements, and schedules. Lead cross-functional initiatives encompassing cloud-native services, platform software, system integration, and release delivery. Help connect software roadmap execution to full-stack readiness, including dependencies across infrastructure, bring-up, validation, and downstream operational support. Identify cross-functional dependencies, mitigate risks, and drive resolution of complex technical and programmatic issues. Establish clear success metrics and reporting mechanis

KubernetesLinuxMachine LearningArtificial Intelligence
A
📍 San Francisco, CA, United States
✓ Quality checkedCompany trend -100%

The Opportunity The world of design is changing rapidly, and the Pro Design team is leading that transformation. We are the Adobe organization behind Illustrator, InDesign, and emerging experiences that connect creativity, collaboration, and AI. Our teams are reimagining what professional design looks like for the next decade - building intelligent, connected tools that empower creators and teams to move faster without sacrificing craft. We are looking for a Senior Business Data Scientist who is creative, analytical, and unafraid to question the status quo and shape the decisions that move key business metrics at scale. Join us and build Adobe’s future products! What you'll Do Map the user funnel and build the metrics, cohorts, and dashboards that Product and Growth rely on to see how users move across free, trial, and paid tiers—and pinpoint where they drop off. Dig into the hard questions (what drives activation, which behaviors predict retention and expansion) and build propensity models for conversion, upgrade, churn, and expansion that feed real-time targeting and in-product nudges. Find and size growth bets and work with Product to ship them. Set north-star, driver, and guardrail metrics with your partners, and stand up multivariate experiments across onboarding, paywalls, in-product prompts, and pricing. What you need to succeed Minimum Requirements: Bachelor's degree in a quantitative field (Statistics, Mathematics, Computer Science, Economics, Engineering, or similar) or equivalent practical experience. 5+ years of experience in data science, product analytics, or a similar quantitative role. Proficiency in SQL and Python (or R) for data manipulation, analysis, and modeling. Hands-on experience designing and analyzing A/B tests and interpre

PythonSQLAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -86.4%

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Infrastructure Quality team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, general contractors, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. Our work spans from vendor qualification through commissioning, ensuring operational readiness across our global portfolio. About the Role We are seeking an experienced Manufacturing Quality Engineer (MQE) to establish, implement, and manage a manufacturing-focused quality program for datacenter infrastructure. This role will be responsible for vendor oversight, quality assurance, process improvement, and issue resolution for all critical systems. You will lead vendor audits, monitor performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced risks, and operational reliability. By partnering with vendors, construction teams, and internal stakeholders, you will help ensure OpenAI’s datacenters are delivered on time and built to the highest operational standards. Travel Domestic and international travel as needed (estimated 40–60%) to manufacturing sites, datacenter locations, and partner facilities. Key Responsibilities Vendor Oversight & Performance Management Conduct manufacturing evaluation, audits, and improve vendor performance across production, inspection, testing, and delivery phases. Develop and track quality metrics to assess manufacturing performance and identify trends. Partner with vendors to refine processes, training, and quality controls to mitigate risks before shipment. Program Development & Execution Develop and maintain a datacenter-focused m

AWSRestAIGo
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for GPU firmware and GPU system software, working directly with engineering teams of key CSP / hyperscale customers to ensure they can reliably manage, update, and operate NVIDIA GPU firmware at fleet scale. You will drive work streams with engineering teams of key CSPs/hyperscale customers to build shared understanding of GPU firmware and system software integration, incorporate their feedback into NVIDIA's feature roadmap and delivery plan, and ensure customer-side automation and recovery procedures are ready before each firmware release. Your cross-CSP visibility enables you to identify patterns in GPU firmware operational challenges that drive systemic improvements no single customer engagement could surface alone. What you'll be doing: Drive GPU firmware & siftware work streams with CSP engineering teams — ensuring they understand GPU firmware architecture (VBIOS, InfoROM, microcontroller firmware), update sequencing, recovery procedures, and GPU power management Gather and synthesize CSP feedback on GPU firmware/software — covering manageability, observability, security requirements (e.g., multi-tenancy isolation, secure boot, attestation), and performance — and champion those priorities into NVIDIA's GPU firmware/software feature roadmap and delivery plan Drive GPU firmware update orchestration for large-scale deployments — multi-GPU update sequencing, rollback strategy, failure handling, and validation across hundreds of GPUs per rack Serve as the technical focal point between NVIDIA and CSP firmware/software engineering — ensuring GPU behaviors (error recovery flows, thermal protection, power state transitions) are well-documented and accessible for customer integration Identify cross-CSP GPU SW/FW issue patterns — common update failu

Artificial IntelligenceAI
C
📍 Maine, United States· Full-time
✓ Quality checked

# Best Email Phishing Simulation Services Email remains one of the most common ways cybercriminals target organizations. Attackers use convincing messages to trick employees into clicking malicious links, opening harmful attachments, sharing credentials, or transferring sensitive information. Even with firewalls, antivirus solutions, and security monitoring in place, one successful phishing email can create a serious security incident. **Email Phishing Simulation Services** help organizations measure employee awareness and strengthen their ability to recognize and report suspicious messages. ## What Is Email Phishing Simulation? Email phishing simulation is a controlled security awareness exercise that recreates realistic phishing scenarios without exposing employees to actual malicious activity. Authorized security professionals design simulated phishing emails based on common attack techniques and send them to selected users according to an approved testing plan. The objective is not to blame employees. Instead, it is to understand how users respond to realistic threats and identify areas where additional security awareness training may be required. A well-designed simulation can test responses to credential-harvesting emails, fake password-reset notifications, suspicious invoices, delivery alerts, account warnings, and other social engineering scenarios. ## Why Phishing Simulations Are Important Technology alone cannot completely prevent phishing attacks. Cybercriminals continually improve their techniques, making fraudulent messages increasingly difficult to distinguish from legitimate communications. Regular phishing simulations can help organizations: * Measure employee phishing awareness. * Identify users who may require additional training. * Test reporting and response procedures. * Improve recognition of suspicious emails. * Reduce the likelihood of credential theft. * Strengthen the organization's security culture. * Track awareness improvements

N
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88.6%

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. This internship will take place from January 25 - April 16 and you will need to be able to work out of our SF office during this time. What You'll Achieve: Conduct data analyses to gain insights about Notion and use these insights to uncover opportunities for improvements in our product and business. Communicate these insights with actionable recommendations to cross-functional teams (insights are useful, impact is even better!). Work with cross-functional partners across the product and business to learn about their functions and use data to advance their respective areas. Create metrics and build dashboards to monitor the growth and health of Notion. Communicate insights and recommendations effectively to leadership and have an impact on strategic decision-making. Qualifications: Pursuing a bachelor's or master's in a quantitative field such as Economics, Statistics, Applied Math, Engineering, Computer Science, or Natural Sciences. Must graduate before December 2027. This internship will take place from January 25 - April 16 and you will need to be able to work out of our SF office during this time. Previous research or internship

PythonSQLRestAI
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We’re seeking an experienced Engineering Program Manager to join Cohere’s customer-facing Engineering Program Management team. We need someone with curiosity, drive, independence, and leadership, who has hands-on experience working directly with customers and managing projects for enterprise-grade software or enterprise-focused machine learning solutions. In return, you’ll have the unique opportunity to shape Cohere’s operations, collaborate with leading minds in the LLM space, work directly with our Strategic Customers as well as Applied ML (AML) Engineering, Forward Deployed Engineering (FDE), Platform , Product and Go-to-Market teams, and be the “technical” voice of Cohere for the customer. You will get a chance to create extremely high-impact contributions to our fast-growing company, product and culture. As an Engineering Program Manager/ Technical Program Manager, you will: Communicate: Provide clear, timely, and objective communication across the tech organization, Cohere teams, leadership, and most importantly - our strategic customers and external partners. Optimize: Break down complex issues into strateg

GitMachine LearningAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -89.8%

From $184K/yr

Quick readStrong listing-quality and freshness signals

TPMs at Datadog see the problems hiding between teams, engineer away the work that shouldn’t require humans, and drive the company’s most technically complex and consequential bets to completion. Technical Program Management at Datadog operates at the intersection of engineering depth and organizational reach by driving high priority, cross-functional programs that are too complex and consequential for any single team to own. We partner with engineering on solving deeply technical problems at scale by connecting the people, decisions, and context to move Datadog's most important work forward. We build the systems and automation that make entire classes of program work self-executing. We are in the architecture conversation early, earning trust through technical judgment. We use AI to surface risks earlier, accelerate program execution plans, and find cross-team patterns that would otherwise stay hidden. The faster teams move, the more essential it is to have someone who can operate across them. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What We Expect: These are the expectations we hold for every TPM at Datadog. Technical depth, product domain expertise, and AI systems literacy; knowing how AI solutions work, where they fail, and the scope and impact of those failures. AI brings more complexity into the picture - the technical bar is higher, not lower. Build the systems that reduce the need for coordination Identify what matters before anyone asks, and automate the rest Engineer program lifecycles end-to-end See what no single team can see and own the solution Drive the company's most technically complex and consequential bets through cross-functional agreement, organizational visibility, and influence Build AI powered automation tools and

RestAIGoRust
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -86.4%

About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. Our product teams work across research, engineering, product, data, and design to bring OpenAI’s technology to people and businesses around the world. Design plays a critical role in making powerful AI intuitive, useful, and trustworthy. We hold a high bar for quality and build experiences that earn users’ confidence, while moving quickly and learning from real-world feedback. We favor early validation and continuous refinement over waiting until every detail is perfect. About the Role We’re looking for a Product Design Leader to shape how startups, small businesses, and growing teams discover, adopt, and unlock lasting value from Codex and ChatGPT. You’ll lead design across the B2B growth journey, creating experiences that turn the potential of AI into meaningful, everyday impact for this user base. This role is a mix of team leadership and hands-on design execution. You’ll lead and develop a small team of designers, contribute directly to high-impact projects, and help set the standard for impactful, simple, user-centric experiences. You balance exceptional craft with momentum, creating a culture where the team launches, learns, and improves quickly. You already use tools such as Codex to extend what you can build. As a leader, you help designers strengthen their judgment, confidence, and influence through clear and actionable feedback. You’re also a highly effective cross-functional partner who brings people together, navigates ambiguity, and builds alignment through clear communication and collaborative problem-solving. Codex is on an extraordinary trajectory, and this role offers a rare opportunity to work across both sides of the product: the consumer experiences through which people first discover and adopt our technology, and the business experiences that help teams use it together. You’ll bring a high bar for craft, strong product judgment, and comfor

AWSRestAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.2%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Forward Deployed Engineer at Baseten, you will partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten’s platform. You’ll own the journey with customers from initial exploration to production deployment, translating ambiguous business goals into reliable, observable services with clear quality, latency, and cost outcomes. This role is a great fit for entrepreneurial engineers who want a front-row view into how modern companies adopt AI at scale and who enjoy working across product, software development, performance engineering, and customer-facing implementations. To be clear, this is an engineering role with hands-on coding and software development that also includes aspects of product management, technical customer success, and pre-sales solution engineering mixed in. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Develop and maintain software systems and product features using one or more general-purpose programming languages in a production-level environment, with a preference for Python due to its relevance in ML projects. Drive customer impact by designing, implementin

PythonDockerMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.2%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Forward Deployed Engineer at Baseten, you will partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten’s platform. You’ll own the journey with customers from initial exploration to production deployment, translating ambiguous business goals into reliable, observable services with clear quality, latency, and cost outcomes. This role is a great fit for entrepreneurial engineers who want a front-row view into how modern companies adopt AI at scale and who enjoy working across product, software development, performance engineering, and customer-facing implementations. To be clear, this is an engineering role with hands-on coding and software development that also includes aspects of product management, technical customer success, and pre-sales solution engineering mixed in. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Develop and maintain software systems and product features using one or more general-purpose programming languages in a production-level environment, with a preference for Python due to its relevance in ML projects. Drive customer impact by designing, implementin

PythonDockerMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.2%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto

PythonCI/CDGitRest
M
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -100%

What we're building Mutiny is the self-improving AI infrastructure for GTM teams to execute faster and close more revenue. Our ambition is to do for revenue velocity what Cursor and Claude Code did for engineering velocity. With Mutiny, everyone in sales and marketing gets a bench of GTM athletes that handle any work across their revenue motion and learn from what's actually moved their deals. In April we re-launched the product as an agent-first platform. Anthropic showcased us as a leader in AI GTM. MRR is growing more than 70% month-over-month, with customers like Uber, Rippling, and Snowflake. We're backed by Sequoia, YC, and Insight, and we're building a generational company. The opportunity We're looking for a Head of GTM to bring Mutiny to millions of users. You'll sit shoulder-to-shoulder with the founders, own revenue end-to-end from first touch to retention, and help shape the decisions that define the company’s brand and culture. You’ll be building the AI infrastructure layer that reshapes how every GTM team operates. This role is in person in New York City, five days a week. What you'll own Revenue. PLG self-serve signups and conversion, sales-led motion including pipeline and ARR, retention and expansion. Marketing. Build the engine from zero to inevitable, including hiring the team and building the channels. GTM engineering. Build the agent-powered growth stack (including Mutiny!) that lets 5 people out-ship a team of 50. Sales and CX. You'll own pricing, evaluation, and expansion strategy to help us close deals and drive customer adoption. Depending on your skills and interest, these functions could fully report to you or you can own the strategy and ops portions. The team. Hire the best GTM people in the world and build a culture of winners. The story. You'll shape how the market understands category-defining AI products for GTM. Who you are A nose for distribution. You're a savant at breaking through noise. You know which channel will work before th

SQLRestAIGo
🔔

Get new engineering maintenance lead jobs in United States by email

Daily job updates · Unsubscribe anytime