A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Substrate is the team responsible for Palantir’s core production infrastructure — 100s of K8s clusters — from on-prem to the major cloud hyperscalers, whether they are internet-connected or air-gapped, small hardware footprint or large. As a Senior Software Engineer on Substrate, you will design and build Palantir’s managed Kubernetes product offerings across all these environments. You and your team will be responsible for bootstrapping and operating the entire fleet of K8s clusters with zero manual steps by building industry leading tooling and contributing to core CNCF components. You will also be responsible for ensuring scale, stability and security across a matrix of compliance regimes and hosting infrastructure types. Your team culture emphasizes engineering rigor and operational excellence at scale. This means issues in production should be pre-empted and deeply root-caused, and investments in automation and self-healing systems are key. If you’re excited about infrastructure at scale and working with Kubernetes, this is the right role for you.
Jobiba hiring network
Senior Infrastructure Architect Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior infrastructure architect jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Our Global Sustaining Engineering team sits at the intersection of software engineering and infrastructure, ensuring the services our customers depend on are fast, resilient, and always available. As a Senior Site Reliability Engineer, you'll take direct ownership of production services — from initial design through day-to-day operation — while partnering with product, engineering, and security teams to build and maintain business-critical systems. In this role, you will deepen your technical expertise and grow your leadership presence by mentoring the next generation of SREs. You will also gain hands-on experience with intelligent tooling in real-world workflows. What you'll get to do... Design, implement, and operate scalable, highly available production services while diagnosing and resolving complex infrastructure, network, and application issues Build and maintain alerting pipelines, dashboards, and SLO-driven monitoring strategies using Icinga, Prometheus, and Grafana Lead incident response end-to-end — performing root-cause analysis, authoring blameless post-mortems, and driving corrective actions to closure Develop and extend Infrastructure as Code coverage and build internal tooling that eliminates manual, repetitive operational work Mentor SRE I and SRE II engineers through code reviews, debugging sessions, and knowledge-sharing talks Apply LLM-driven log analysis, anomaly detection, and generative AI tools to accelerate incident response and runbook creation — validating all outputs before use Your experien
ROLE DESCRIPTION: We’re looking for a Senior Platform Backend Developer who can help us support the development organization to deliver value to customers in a reliable, efficient, and safe manner. You’ll be working in a focused team that owns one or more pieces of the production application environment and the developer experience, you will own and deliver in service of quarterly goals on the team. ABOUT THE TEAM: This role is within our Backend Platform team. The team primarily uses Go, Scala, and PHP and has expertise in technologies such as Kafka, various AWS services, and some infrastructure-as-code tools. Your primary focus will be on developing services and tools for our product development teams as well as modernizing our existing platform. Based out of British Columbia, you will report to the Senior Manager, Software Development, DevOps. WHAT YOU’LL DO: Design and build software - tools, libraries, automation, services, and glue scripts Responsible for the reliability, security, and integrity of our large, cloud-based platform Participate in a flexible on-call rotation Lead by owning project milestones, epics or features Practice continuous improvement, contributing to culture, process, and direction in your team and across our department Develop processes and automation to eliminate repetitive tasks Design and build our infrastructure platform Identify and implement new platform features Research and evaluate new technologies Refactor, rewrite or retire existing platform features Operate our developer experience and production application environments Diagnose and repair our distributed systems Perform maintenance, upgrades, and migrations Control or eliminate repetitive tasks, alert noise, and business-as-usual work Enable development teams Provide executable interfaces to our infrastructure platform Provide tools and best practices to support the entire software development lifecycle Collaborate with others across the orga
We’re looking for an Senior Software Developer, Backend who can help us support the development organization to deliver value to customers in a reliable, efficient, and safe manner. You’ll be working in a focused team that owns one piece of the production application environment and the developer experience, you will execute on defined projects to achieve team-level goals. In line with Hootsuite's distributed workforce strategy, our flexible work arrangement allows for a hybrid model. This role is open to applicants within commutable distance to Luxembourg. WHAT YOU’LL DO: Write software - tools, libraries, automation, services Design and build our infrastructure platform Identify and implement new platform features Research and evaluate new technologies Refactor, rewrite or retire existing platform features Operate our developer experience and production application environments Diagnose and repair our distributed systems Perform maintenance, upgrades and migrations Control or eliminate repetitive tasks, alert noise, and business-as-usual work Enable development teams Provide executable interfaces to our infrastructure platform Provide tools and best practices to support the entire software development lifecycle Participate in a flexible on-call rotation Communicate by writing documentation, participating in meetings, and showing off your work at demos WHAT YOU’LL NEED: A degree in Computer Science or Engineering or equivalent experience working in a software engineering role An ability to write software and working knowledge of software engineering practice (Java programming language and strong working knowledge of object-oriented programming concepts) Proven experience creating stable, reliable, performing and maintainable code Familiarity with data modeling and schema design Knowledge of data structures and algorithms Open Communication: clearly conveys thoughts, both written and verbally, listening attentively and asking questions for clarific
About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About the Team The Revenue Strategy & Operations team at Mixpanel partners with Regional and Global Business Leaders to set and execute revenue strategy across the customer lifecycle. We build the strategy, operational processes, reporting infrastructure, and decision-making frameworks that make our GTM teams successful. About the Role As Senior Customer Strategy & Operations Manager, you’ll be the strategic advisor and operating partner to our VP of Global Customer Success. You’ll own how we understand, retain, and grow our customer base - diagnosing what drives upsell, what predicts churn, and what we need to build to scale the post-sales motion. This isn’t a role where you inherit a clean process and tune it at the margins. You’ll get your hands dirty in customer-level data, design the systems that turn signals into action, and build AI-powered tooling. You’ll work directly with CS leadership day-to-day and feed field-level insight back into the broader Revenue Strategy team and GTM leadership. You bring the structured thinking of a consultant and the bias for action of an operator. You’re equally comfortable in a strategy session with the VP and three layers deep in a SQL query. Responsibilities Customer analysis at the account level. Get hands-on with the data to understand what drives expansion, what predicts churn, and where the highest-leverage interventions live. You’ll connect product usage signals, customer health, and commercial outcomes and translate the findings into decisions the CS org acts on. Forecasting and pipeline rigor for the post-sales motion. Build and evolve how we
About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About the Team The Revenue Strategy & Operations team at Mixpanel partners with Regional and Global Business Leaders to set and execute revenue strategy across the customer lifecycle. We build the strategy, operational processes, reporting infrastructure, and decision-making frameworks that make our GTM teams successful. About the Role As Senior Customer Strategy & Operations Manager, you’ll be the strategic advisor and operating partner to our VP of Global Customer Success. You’ll own how we understand, retain, and grow our customer base - diagnosing what drives upsell, what predicts churn, and what we need to build to scale the post-sales motion. This isn’t a role where you inherit a clean process and tune it at the margins. You’ll get your hands dirty in customer-level data, design the systems that turn signals into action, and build AI-powered tooling. You’ll work directly with CS leadership day-to-day and feed field-level insight back into the broader Revenue Strategy team and GTM leadership. You bring the structured thinking of a consultant and the bias for action of an operator. You’re equally comfortable in a strategy session with the VP and three layers deep in a SQL query. Responsibilities Customer analysis at the account level. Get hands-on with the data to understand what drives expansion, what predicts churn, and where the highest-leverage interventions live. You’ll connect product usage signals, customer health, and commercial outcomes and translate the findings into decisions the CS org acts on. Forecasting and pipeline rigor for the post-sales motion. Build and evolve how we
About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About the Team The Revenue Strategy & Operations team at Mixpanel partners with Regional and Global Business Leaders to set and execute revenue strategy across the customer lifecycle. We build the strategy, operational processes, reporting infrastructure, and decision-making frameworks that make our GTM teams successful. About the Role As Senior Customer Strategy & Operations Manager, you’ll be the strategic advisor and operating partner to our VP of Global Customer Success. You’ll own how we understand, retain, and grow our customer base - diagnosing what drives upsell, what predicts churn, and what we need to build to scale the post-sales motion. This isn’t a role where you inherit a clean process and tune it at the margins. You’ll get your hands dirty in customer-level data, design the systems that turn signals into action, and build AI-powered tooling. You’ll work directly with CS leadership day-to-day and feed field-level insight back into the broader Revenue Strategy team and GTM leadership. You bring the structured thinking of a consultant and the bias for action of an operator. You’re equally comfortable in a strategy session with the VP and three layers deep in a SQL query. Responsibilities Customer analysis at the account level. Get hands-on with the data to understand what drives expansion, what predicts churn, and where the highest-leverage interventions live. You’ll connect product usage signals, customer health, and commercial outcomes and translate the findings into decisions the CS org acts on. Forecasting and pipeline rigor for the post-sales motion. Build and evolve how we
#Team Nextdoor Nextdoor (NYSE: NXDR) is the essential neighborhood network. Neighbors, public agencies, and businesses use Nextdoor to connect around local information that matters in more than 350,000 neighborhoods across 11 countries. Nextdoor builds innovative technology to foster local community, share important news, and create neighborhood connections at scale. Download the app and join the neighborhood at nextdoor.com . Meet Your Future Neighbors As an Android Software Engineer at Nextdoor, you’ll join a fast moving team of developers, product managers, and designers who are passionate about using technology to cultivate a kinder world where everyone has a neighbor they can rely on. The Nextdoor Android team works on features and infrastructure to deliver our values to our members. We care about making an incredible Android app that respects platform conventions and is delightful to use. We’re always trying to move faster and more safely, by adopting the latest practices, such as Kotlin, Jetpack Compose, and GraphQL. At Nextdoor, we operate in an AI-first environment and expect every team member to actively use AI tools as part of their workflow. We aren't looking for prompt engineers; we’re looking for people who use tools like Claude, Gemini, ChatGPT, and Glean to challenge their own thinking and take full ownership of AI-assisted outputs. We also offer a warm and inclusive work environment that embraces a hybrid employment model, blending an in office presence and work from home experience for our valued employees. The hiring team will go over these expectations with you if you are being considered for a role near one of our offices in San Francisco, Los Angeles, Chicago, Dallas, New York, and London. The Impact You’ll Make We believe in empowering our teams to own all aspects of bringing Nextdoor to life. As such, you’ll get the opportunity to make key contributions across our Android stack - this includes developing and improving our networki
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Job Description/ Responsibilities: Designing, developing and maintaining stable and reliable AI/ML Ops platforms / pipelines Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production. Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow. Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data Technology Eva
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Infrastructure Compute Site Reliability Engineering mission is to own and manage the successful operation of our underlying cell infrastructure system, along with elements of service discovery, secrets management and related software layers. We’re looking for a skilled Senior Site Reliability Engineer with strong programming skills to help us build Roblox's private cloud, productionize our growing Kubernetes-based infrastructure, and institute reliability best practices across the Roblox Compute team. You will: Design and Develop systems & libraries that promote fault-tolerance and resilience, automate much of the management and lifecycle of our clusters, and ensure systems are observable. Promote and Institute reliability best practices across the Infra Compute group, drive common reliability initiatives. Provides collaborative technical reviews and operational guidance to strengthen system reliability. Build, Automate and Standardize process automation to create a "golden path" of tooling and platform support that powers the fundamental Roblox ecosystem. Create Tooling that provides production guardrails, by evaluating release candidate capacity with load testing tooling before de
The Data Engineering team’s mission is to ensure high-quality data to enable data-informed decision-making across Asana. You will build data artifacts that are leveraged by Product and Business Data Science teams to optimize our user adoption, growth, and experience. In this role, you will partner with the Infrastructure team to build a self-service analytics platform for the company. This role is based in our Vancouver office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday; most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Design, implement, and scale end-to-end data products that support growing data processing and analytical needs Transform raw data into actionable insights to drive product strategy and power in-depth analyses and reporting Leverage AI to build self-serve tools and accelerate Data/GTM workflows Partner with data scientists, domain experts, and engineering teams to develop a roadmap that aligns with our business goals Implement systems that guarantee data quality, governance, and availability About you 5+ years of experience in Data Engineering or Software Engineering Experience in data modeling and building scalable data pipelines involving complex transformations Proficiency in data processing and storage technologies like Databricks, AWS/S3, Python/Scala/Java, SQL, Spark, and Airflow Proactive and innovative in identifying and addressing performance bottlenecks in existing workflows Motivated to work closely with cross-functional partners to evolve our analytical data model Demonstrates curiosity about AI tools and emerging technologies, with a willingness to learn and leverage them to enhance productivity, collaboration, or decision-making At Asana, we're com
At Asana, security is foundational to our mission of helping teams work together effortlessly. Our Security organization protects Asana’s employees, users, and customers by proactively addressing threats, enabling secure product development, and embedding security into the fabric of how we operate. We partner closely with Engineering, Product, IT, Legal, and Compliance to build and maintain trust at scale. We are seeking an Engineering Manager, Security to lead and grow our Security Engineering organization as Asana continues to grow globally. This role is pivotal in translating high-level security strategy into technical reality, ensuring our infrastructure is resilient by design and our internal security tooling is world-class. This is a leadership role with accountability for people management, technical mentorship, and the delivery of high-impact security initiatives across a complex, high-growth SaaS environment. This role is based in our San Francisco office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Lead and mentor a multi-disciplinary engineering team, fostering a culture of technical excellence, psychological safety, and high accountability. Own the security engineering roadmap, ensuring high-quality delivery and measurable risk reduction through effective prioritization. Drive recruiting efforts to attract top-tier talent and establish clear, ambitious career paths for your team. Champion a high-performance environment where security rigor acts as a catalyst for innovation rather than a bottleneck to velocity. Collaborate with Product and Engineering teams to embed security requirements directly into system des
We are looking for a Program Manager to join our People Team, reporting to the Chief of Staff to the Chief People Officer. In this role, you will drive large-scale, high-impact, and cross-functional programs that shape and evolve our People products and infrastructure. You will operate at both a strategic and executional level, partnering across the People Team, and with business and functional stakeholders across Finance, Legal, IT, and business leadership to deliver transformative initiatives that support Datadog’s growth. This is a highly visible role requiring strong program management, systems thinking, and the ability to connect initiatives across the organization while improving operational effectiveness and employee experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead end-to-end delivery of complex, cross-functional People programs, from ideation through execution and iteration Partner with senior stakeholders across the business to define program goals, success metrics, and roadmaps Deliver People products through scalable processes and frameworks that improve employee experience and organizational effectiveness Identify dependencies, risks, and tradeoffs across initiatives, proactively driving alignment and decision-making Translate strategic priorities into actionable plans, ensuring clear communication and accountability across teams Establish program governance, reporting, and operating rhythms to track progress and outcomes Synthesize insights across multiple workstreams to inform executive-level updates and recommendations Continuously improve program management practices within the People Team Who You Are: 5+ years of experience in program management, operations, consulting, or a related field, preferably
About Datadog: Datadog is the world-class monitoring and security platform for cloud applications. We’re dedicated to creating, developing, and supporting our product and customers, allowing for seamless collaboration and problem-solving among Dev, Ops and Security teams globally. Built by engineers, for engineers, our SaaS product is used by organizations of all sizes across a wide range of industries to enable digital transformation, cloud migration, and infrastructure monitoring of our customers’ entire technology stack. Given the resilience of cloud technologies and importance placed today in digital operations and agility, Datadog continues to innovate and is well positioned for the long term. The Team: Datadog's Finance team collaborates with teams across the organization, providing commercial, operational, and analytical support to ensure that Datadog's business continues to grow as rapidly and efficiently as possible. The Opportunity: We are seeking a Senior Revenue Accountant to join our growing Finance team at Datadog. As a member of the finance team, the Senior Revenue Accountant will be a key member of the team in developing more efficient revenue / billing data flow and close processes and be on the front lines of supporting the rapid growth of the Company. You Will: Work closely with the broader finance team to support the finance operations, accounting and compliance function for our business. Complete month-end close responsibilities including preparing journal entries, balance sheet reconciliations, and supporting schedules. Prepare memos and analyses surrounding revenue recognition under ASC 606 for more complicated billing arrangements Assist in calculation of key internal revenue metrics and invoicing during month-end close. Generate and issue prepaid and monthly invoices issued in arrears. Help customers understand their bills, which vary with usage and plan type. Effectively detect, communicate and work to resolve customer bi
Here at Datadog, we think about offensive security a little bit differently. We embrace automation and AI to run adversary simulations continuously across a massive cloud-native environment, and we expect our offensive engineers to build the tooling that makes that possible. We're looking for a Senior Security Engineer who can execute sophisticated red team operations, write the code that scales them, and take an AI-first approach to offensive security engineering. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Plan and execute red team engagements end-to-end, simulating real-world threat actors across cloud infrastructure (AWS, GCP), Kubernetes, CI/CD pipelines, and corporate environments Build and maintain custom offensive tooling, automation frameworks, and engagement infrastructure, treating offensive operations as a software engineering problem Develop custom payloads and evasion capabilities tailored to Datadog's environment and modern defensive controls (EDR, SIEM, network monitoring) Improve the efficiency of offensive operations through thoughtful use of automation and AI, accelerating reconnaissance, vulnerability analysis, and reporting workflows Partner with the Detection & Response team on purple team exercises to validate detection logic, improve alert fidelity, and influence threat models Translate offensive findings into concrete improvements by working directly with defensive security and engineering teams to close gaps Who You Are: You have 5+ years of hands-on experience in offensive security (red teaming, penetration testing, or adversary simulation) with a track record of operating against mature, well-defended environments You write production-quality code (Python, Go, or similar), can build your own tools, and automate your w
Get new senior infrastructure architect jobs by email
Daily job updates · Unsubscribe anytime