The AI platform is responsible for all AI infrastructure across Datadog. Our mission is to provide tools and platforms that enable data scientists and engineers to conduct large-scale training and inference with ease. We support products such as Bits AI , LLMObs and all our AI research . As an engineering manager for the Evaluation & Annotation team, you’ll join a new and fast growing team and organization. You will support building and scaling the team, define our technical vision and help shape the roadmap. Your team will lead the charge on multiple critical technical challenges: AI model evaluation both offline and online, designing tooling and processes around human annotation, and establishing the standard around synthetics and AI generated datasets. You’ll work closely with sister teams in the AI platform organization ensuring a seamless AI development cycle. You’ll also partner with the Applied AI org and with Datadog infrastructure & tooling teams to build out systems from the ground up. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Manage and grow the Evaluation & Annotation team, directly managing 4-6 engineers Define our technical roadmap in alignment with AI platform goals and the Applied AI team roadmap. Work with our core platform teams to tailor Datadog's storage and data pipelines to our needs Create a strong team culture aligned with our engineering standards and our customer focus Participate in hands-on work: Code reviews, design reviews and some coding Who You Are: A Software Engineer at heart with a previous experience leading software engineering teams, as a tech lead or people manager Excellent leader with strong interpersonal skills, and the
Jobiba hiring network
Lead Infrastructure Software Engineer Jobs
6,876 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current lead infrastructure software engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The AI platform is responsible for all AI infrastructure across Datadog. Our mission is to provide tools and platforms that enable data scientists and engineers to conduct large-scale training and inference with ease. We support products such as Bits AI , LLMObs and all our AI research . As an engineering manager for the Training & Serving team, you’ll join a new and fast growing team and organization. You will support building and scaling the team, define our technical vision and help shape the roadmap. Your team will lead the charge on multiple critical technical challenges: distributed training of foundation models, serving at scale, designing the user experience. You’ll work closely with sister teams in the AI platform organization ensuring a seamless AI development cycle. You’ll also partner with the Applied AI org and with Datadog infrastructure & tooling teams to build out systems from the ground up. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Manage and grow the Training & Serving team, directly managing 10+ engineers Define our technical roadmap in alignment with AI platform goals and the Applied AI team roadmap. Work with our core platform teams to tailor Datadog's storage, infrastructure and data pipelines to our needs Create a strong team culture aligned with our engineering standards and our customer focus Participate in hands-on work: Code reviews, design reviews and some coding Who You Are: Previous experience (1+ years) leading software engineering teams, as a tech lead or people manager Strong technician with a mix of backend, data engineer and infrastructure experience who is interested in remaining a hands-on leader Excellent leader with strong
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team As a platform company powering businesses all over the world, Stripe processes payments, runs marketplaces, detects fraud, helps entrepreneurs start internet businesses from anywhere in the world, builds world-class developer-friendly APIs, and more. The engineering teams at Stripe work on the business logic for all of that. As a software engineering manager at Stripe, you'll lead teams that build and expand APIs, services, and experiences. You'll also work with partners to launch new markets and vertical capabilities. Stripe Terminal helps our users extend their online presence to the physical world. The mission of the Terminal team is to make it as easy for businesses to accept in-person payments as the Stripe API has done for online payments—building for Unified Commerce. With Terminal, businesses can unify their in-person and online experiences, unlocking payments use cases that are right for their business model—whether it's creating a modern retail experience, extending their website to a pop-up store, or enabling a mobile point-of-sale at their next event. Within Terminal, the Commerce team focuses on the following to help make Unified Commerce a reality for businesses across the globe. Enterprise features—on-reader forms, tipping, gift cards Payment features—DCC, surcharge, multi-capture Point-of-sale integrations—enterprise-focused integrations Low-code and no-code—solutions that will enable our users to onboard quickly What you'll
About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API. You’ll work cross-functionally with teams building core services, training models, and developing real-time user experiences to push our latency, throughput, and cost-efficiency to the next level. We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack. In this role, you will: Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond. Develop tooling and metrics that provide deep observability into system performance. Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements. Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale. Lead investigations into high-impact performance regressions or scalability issues in production. Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems. You might thrive in this role if you: Have 7+ years of experience in software engineering with a strong tr
A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO • Create and execute automated and manual test cases to ensure optimal system performance according to specifications. • Work closely with software development and support teams to deliver high-quality applications in a timely manner. • Develop and maintain comprehensive test plans, including manual and automated tests for functional, regression, and integration testing. • Build consensus between stakeholders and developers to define clear and testable acceptance criteria. • Document quality assurance and process flows, both existing and proposed. • Advocate for best quality assurance practices and testing techniques. • Ensure any new software changes meet business, legal, compliance, and technical requirements. • Develop and maintain test automation frameworks for various applications in trading domains. • Maintain and track quality assurance capacity, velocity, statuses, and deliverables with the team lead. WHAT’S REQUIRED • 5+ years of experience in quality assurance with test automation, software engineering, or business analysis roles. • Comprehensive expertise in processing trades, managing trading and lifecycle events, handling corporate actions, and managing cash flows. • Experience supporting fixed income products, Equities and Treasuries. • Solid understanding of position management for listed and OTC products. • Understanding of SDLC, Test lifecycle, and testing methodologies. • Experience creating and writing SQL que
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives, spanning AI research specialists, silicon designers, software engineers and systems architects. Job Summary We are looking for an experienced Principal Engineer to join our System Management team and help lead the development of critical interfaces used by internal and external customers to manage system state. You will provide technical leadership within assigned areas of System Management, guide architecture and implementation choices, mentor engineers and translate broader technical direction into effective execution. This is a hands-on engineering role for someone who can lead complex technical work, improve reliability and operational readiness, and collaborate effectively across multiple engineering disciplines. The Team The System Management team sits within the Software Platform group and helps build Graphcore products into large-scale AI solutions for our customers. The team is responsible for developing the interfaces between hardware, AI software and frameworks, as well as providing interfaces for public and private cloud environments. This includes system management capabilities that abstract complex hardware administration and enable reliable deployment and operation at scale. As one of the first teams to work with new hardware and software, we regularly solve complex system-level problems
About us Graphcore is one of the world’s leading innovators in artificial intelligence compute. We are developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and support the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a family of companies responsible for some of the world’s most transformative technologies. Together, we share a bold vision to enable advanced artificial intelligence and ensure its benefits are accessible to everyone. Graphcore brings together AI researchers, silicon designers, software engineers and systems architects to solve complex technical challenges and deliver innovative computing solutions. Job Summary The Principal Electrical Engineer will be a technical authority within Data Center Engineering, leading the architecture and delivery of safe, resilient and scalable electrical infrastructure for high-density AI computing environments. Working with internal teams, data center developers, utilities, consultants and equipment partners, this role will guide projects from early technical studies through design, construction, commissioning, operation and lifecycle improvement. The successful candidate must reside in, or be willing to relocate to, Austin, Texas. Approximately 10% travel may be required. The Team The Data Center Engineering team is responsible for defining and enabling the infrastructure needed to deploy and operate Graphcore’s computing systems at scale. The team works across electrical, mechanical, thermal, controls, systems and operational disciplines, collaborating with external engineering and construction partners to deliver reliable, efficient and maintainable data center environments. Responsibilities and Duties Act as the technical authority for electrical engineering across data center infrastructure projects, from the utility or on-site power source through to the IT rack. Lead electrical archit
Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu
Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu
Citibank, N.A. seeks an Applications Development Tech Lead Analyst for its Tampa, Florida location. Duties: Responsible for source code review to ensure it satisfies good coding practice. Coordinate all communication between tech team and Ab Initio vendor, including scheduling and chairing weekly meetings. Responsible for reconfiguration and reengineering to optimize performance. Design, develop and modify software systems, using scientific analysis and mathematical models to predict and measure outcomes and consequences of design. Prepare reports or correspondence concerning project specifications, activities, or status. Confer with systems analysts, engineers, programmers and others to design systems and to obtain information on project limitations and capabilities, performance requirements and interfaces. Manage the hardware infrastructure and monitor the performance metrices E2E. Identify areas of growth need and proposal scale up activities. Execute on scale up activities by submitting procurement requests, manage regular calls with the deployment team and plan go live/switch activity. Develop shell scripting code base to support the job executions, interactions between systems. A telecommuting/hybrid work schedule may be permitted within a commutable distance from the worksite, in accordance with Citi policies and protocols. Requirements: Requires a Bachelor’s degree, or foreign equivalent in Information Technology, Engineering (any) or related field and 6 years of progressively responsible, post-baccalaureate experience as a Software Engineer, Associate Director – Data Engineering, Senior Consultant, Data Specialist, Assistant Systems Engineer or related position involving gathering new business requirements, architecture, analysis, estimation, design, implementation, and leading a team of software developers and software development. 6 years of experience must include: Experience with AbInitio Data Processing (SME), Data Modeling
As an Engineering Manager on Coder’s Agentic Engineering team, you’ll lead engineers building and evolving the systems behind our agentic development experience. You’ll help make agents more capable, reliable, and useful across real development environments. You’ll guide technical direction, grow the team, and keep execution sharp. You’ll work closely with Engineering, Product, and Design across the agent harness, integrations, and developer workflows. What you’ll do here Lead and grow a team within our Agentic Engineering organization. Set technical direction across the agent harness, integrations, and workflows. Stay close to the code and contribute to architecture and implementation decisions. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve reliability, performance, and operability across agentic systems. Coach engineers, raise the technical bar, and create clarity around priorities and tradeoffs. What we’re looking for Experience managing and growing software engineering teams. Strong hands-on engineering experience with Go. Experience with React and TypeScript. Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS. Strong technical judgment and comfort working through ambiguity. A track record of helping engineers grow while maintaining a high execution bar. Our tech stack Backend: Go, Postgres Frontend: TypeScript, React Infrastructure: AWS, Kubernetes Observability: Prometheus, Grafana CI/CD: GitHub Actions Bonus tacos if you have (Tacos? If you need an ice-breaker, ask how we say thanks by giving tacos!) Experience building coding agents, developer tools, or cloud development enviro
As Engineering Manager for Threat Detection, you will lead a high-performing team that powers Datadog's detection program. Threat Detection is the organization responsible for keeping Datadog ahead of an evolving threat environment: closing coverage gaps faster, raising the bar on signal quality, and shipping detections that hold up under the scale and complexity of cloud-native infrastructure. Your team will combine direct detection expertise, platform engineering, and applied AI to ship detections at a pace and scale traditional rule-writing alone cannot match. Examples of what your team will work on include detection-authoring agents, the detection platform that powers every rule in production, coverage analysis, alert triage and response automation, and the evaluation infrastructure that holds these systems to a high bar of fidelity. Detection authorship is a shared responsibility across the organization, and your team will contribute both by building the systems that scale our authoring capacity and by writing detections directly when their domain expertise is the right tool. You will partner closely with our Security Incident & Response Team (SIRT), Cyber Threat Intelligence (CTI), AI Engineering teams, and Datadog's broader Security organization. This is a high-impact leadership role: you will grow a team of security and software engineers responsible for building and executing our detection and AI strategy. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the strategy, roadmap, and execution of Datadog Security's shift to AI-accelerated detection and response. Drive development of high-fidelity detections as a shared responsibility across the organization, ensuring your team's systems and direct contributions raise the bar on coverage and
Director of Engineering, Physical AI Role Overview The Director of Engineering will report to the General Manager of Physical AI, and will be responsible for leading a multi-disciplinary engineering organization. In this senior leadership role, you will own the execution of the Physical AI Data Engine — the platform powering the next generation of Physical AI/Embodied AI. You will collaborate closely with Operations and GTM to guide product direction and help solve the data bottleneck that stands between today's robotics research and real-world deployment. This role requires significant ownership in a fast-paced environment and you will motivate internal teams to set the pace for business growth. Travel will come into play. Key Responsibilities: Set and drive the technical vision across data collection infrastructure, teleoperation systems, ML training pipelines, model evaluation frameworks, annotation tooling, and research Lead a multidisciplinary engineering organization—spanning engineering managers, software engineers, ML engineers, and ML research scientists—while designing the organizational structure, talent strategy, and culture required to scale rapidly without compromising on quality or strategic alignment Maintain exceptional technical and operational excellence by deeply understanding team deliverables, asking incisive questions, identifying slipping standards early, and knowing precisely when to step in Drive cross-functional alignment across Engineering, Operations, and GTM on platform architecture, release processes, and shared priorities Collaborate with researchers and clients to architect and deliver scalable, production-grade data infrastructure tailored for complex robotics workloads Required Qualifications: Bachelor's degree in Engineering, Robotics, Computer Science, or a related technical field 8+ years of engineering experience in fast-paced environments, including 4+ years direct people management demonstrated history of recruiting, mentorin
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng
As a Staff Engineer on Datadog's Compute – Disruption and Workload Placement team, you'll help define how our Kubernetes fleet scales to meet the demands of rapidly growing AI and cloud-native workloads. You'll work on the systems that ensure engineering teams have the right compute capacity, in the right region, at the right time across AWS, Google Cloud, and Azure. This is a highly technical, high-impact role where you'll shape the future of capacity orchestration, influence platform architecture, and solve infrastructure challenges that directly support Datadog's continued growth. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead the technical direction of capacity management and workload placement for Datadog's Kubernetes platform spanning 100,000+ virtual machines across multiple cloud providers. Design and build systems that optimize how engineering workloads are scheduled and deployed across regions while balancing capacity constraints, reliability, and performance. Partner across infrastructure teams to evolve multi-region and multi-cloud capacity orchestration as Datadog continues to scale. Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency. Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies. Who You Are: You have significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems. You are an experienced software engineer with strong programming skills, ideally in Go or a comparable systems programming language. You have hands-on experience with at least one major cloud provider (AWS, Google Cloud, or Azure) and understand distributed cloud infrastructure. Yo
Get new lead infrastructure software engineer jobs by email
Daily job updates · Unsubscribe anytime