Senior Infrastructure Architect — Enterprise Observability and Automation Description - Job Summary Senior individual contributor responsible for the architecture, implementation, and operational ownership of enterprise observability, monitoring, and automation platforms across HP's global IT environment. This role modernizes infrastructure visibility capabilities while ensuring operational stability, security, and compliance. Serves as a technical and operational bridge between infrastructure engineering, cybersecurity, SOX/compliance stakeholders, automation teams, and external technology partners — leading complex initiatives such as platform migrations, enterprise integrations, and governance enablement. Responsibilities Enterprise Observability and Monitoring Application owner and senior technical authority for enterprise monitoring and logging platforms (Datadog, Splunk), including platform governance, roadmap alignment, and operational oversight. Lead enterprise-scale monitoring platform migrations, including architecture design, agent strategy, data ingestion models, vendor coordination, and deployment across 5,000+ servers. Define standards for alerting, dashboards, observability data quality, and integration with ITSM platforms (ServiceNow). Design and manage multi-org Datadog architecture, including org structure, RBAC, SSO/SAML, secrets management, and cybersecurity compliance. Oversee SNMP-based monitoring of storage and network devices, including device profiling, syslog/event integration, and NetFlow collection. SOX Compliance and IT Governance SOX control owner for enterprise monitoring applications — approve monthly reviews, participate in internal/external audits (EY), and maintain ITGC/SOX compliance. Provide audit evidence, walkthrough docu
Jobiba hiring network
Senior Infrastructure Architect Jobs
7,292 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior infrastructure architect jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Location Details: Pune, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a hybrid position. You’ll divide your time between working remotely from your home and an office, so you should live within commuting distance. Hybrid teams may work in-office as much as a few times a week or as little as once a month or quarter, as decided by leadership. The hiring manager can share more about what hybrid work might look like for this team. Join our Team Our team builds and operates the foundational infrastructure platforms that power GoDaddy's engineering organization. We own critical services including secrets management, software distribution, host security controls, and live patching for thousands of Linux systems running on OpenStack. This role sits at the intersection of Linux engineering, platform engineering, reliability engineering, and security. You will help define how core infrastructure services are designed, operated, automated, and scaled across the enterprise! What you'll get to do... Design, build, and operate highly available, scalable, and secure infrastructure platforms supporting large-scale Linux environments, with a focus on reliability, resiliency, and operational efficiency Lead the architecture, implementation, and operation of infrastructure services, including OpenStack, enterprise secrets management, package management, software promotion pipelines, and platform lifecycle management Develop and maintain automation solutions using infrastructure-as-code, Ansible, Python, Go, and self-service capabilities to improve efficiency and reduce operational overhead Build and improve observability and reliability practices through monitoring, logging, alerting, dashboards, managing incidents, analyzing underlying causes, disaster recovery, and service health reporting
We're looking for a Senior Infrastructure Engineer who brings strong software engineering skills and a deep understanding of production systems. This role is a good fit for someone who enjoys building systems that make infrastructure more scalable, reliable, and easy to operate – using code, not runbooks. You'll work with a highly collaborative team to design and build the internal platforms that power all of Asana, from product features to AI systems to offline analytics. Our tech stack includes: AWS, Kubernetes (EKS), MySQL (RDS), OpenSearch, DynamoDB, Redis, Terraform, Datadog, TypeScript, Scala, Go, and Python. We’re especially interested in people who think like backend engineers but care deeply about systems – things like failure modes, operational cost, debuggability, and performance. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you’ll achieve: Design and build frameworks, tools, and services that improve the reliability, observability, and scalability of Asana’s infrastructure. Lead end-to-end projects, from scoping and design through to rollout, across multiple systems and teams. Improve the operability of stateful infrastructure like MySQL, OpenSearch, and DynamoDB – and help drive Asana’s long-term vision for storage reliability. Debug production issues across the stack. Yes, there’s an on-call rotation – but this isn’t a pager monkey role. You’re here to fix things properly and make sure they don’t break again. Partner with product teams to shape a service-oriented architecture that enables fast, reliable development. Share knowledge through code reviews, design discussions, and mentorship. Abou
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Infrastructure Engineer Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning
About the Team OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions. Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency. About the Role We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments. This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment. The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries. Key Responsibilities Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure. Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios. Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize
Senior Consultant-We are seeking a seasoned ITES Solution Architect to lead the design and execution of enterprise-grade IT infrastructure projects. -"Key Responsibilities: Infrastructure Architecture & Design • Design and implement end-to-end IT infrastructure solutions including compute, storage, network, and security. • Architect high-availability systems with failover clustering, load balancing, and redundancy. • Define and implement disaster recovery strategies with clear RTO/RPO objectives. Data Center & DR Strategy • Design and manage primary and secondary data center environments. • Plan and execute DR site configurations, ensuring seamless failover and recovery. • Implement real-time or scheduled data synchronization between DC and DR using replication technologies (e.g., SAN replication, DFS-R, Veeam, Zerto). Backup & Recovery Planning • Develop and maintain enterprise backup strategies using tools like Commvault, Veeam, or NetBackup. • Ensure secure, encrypted backups with retentio
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the Team: Bitstamp Exchange Platform Bitstamp made history in 2011 as the world’s first regulated crypto exchange. As a key part of the Robinhood family, the Exchange Platform team owns the full service lifecycle. We are the architects of a modern, high-velocity ecosystem that enables our global expansion, ensuring the world’s longest-running exchange remains unshakeable. The Role As a Senior Backend Engineer integrated into the Exchange Platform team, you will be a key driver of our modernization strategy. You will be an essential part of building the new generation infrastructure for low-latency services in the cloud, adopting cutting-edge technologies to drive our high-velocity ecosystem. This role is based in our London office(s), with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. Requires participation in an on-call rotation to support business needs. What You’ll Do Build Next-Gen Infrastructure: Architect and implement high-availability, low-latency cloud infrastructure that ensures performance and portability across our ecosystem. Modernize and Scale: Lead the transition of our services to modern, containerized environments, optimizing deployment and scaling workflows. Robinhood Ecosystem Integration: Manage and execute high-priority integration projects that align Bitstamp's backend with Robinhood's global infrastructure. Evolve the Stack: Identify and implement back
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Cloud Security Engineer, you will define and implement the security strategy and controls across our hybrid and multi-cloud environment. Embedded within the Platform Security team, you will operate with a high degree of autonomy, partnering closely with Infosec and Infrastructure teams to secure our cloud infrastructure. Your architectural decisions will directly impact the security of a global platform. You Will: Create Paved Roads: Build innovative services and tooling that make the secure path the easiest path, enabling developers to deploy faster and safer. Engineer Self-Healing Infrastructure: Architect and scale systems that monitor our cloud posture and automatically enforce a self-healing security baseline. Design Secure-by-Default Blueprints: Partner deeply with infrastructure teams to bake threat modeling and secure-by-default patterns into the core DNA of our cloud environments. Implement Effective Guardrails: Deploy organizational security controls that protect our developers without slowing them down. Advance Detections: Write and optimize detections tailored specifically to our footprint. You Have: 4+ years: of relevant professional experience. Experience writing a
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Software Engineer, Backend on the Growth Foundations team within the Consumer & Business group, you'll build and scale the core infrastructure that powers growth across Coinbase. This team owns the foundational systems behind powering efforts as: notifications at scale, incentives and referral platforms, user targeting and segmentation, lifecycle communications, and ML ranking that power more than 50% of Coinbase web and app surfaces. You'll design and ship backend systems that enable high-impact growth initiatives for Retail, Institutional, Base, and international teams while maintaining the reliability, scalability, and performance our customers depend on. What you'll do: Own the design and delivery of scalable distributed backend services that handle high QPS, low latency, and high reliability requirements for growth infrastructure Architect resilient service-oriented systems using modern cloud technologies, contributing across backend, data, and ML layers of the stack Partner with engineers, designers, product managers, and senior leadership to translate product and technical vision into quarterly roadmaps Lead investments in developer tooling, shared platform capabilities, and cross-team relationships that increase organizational velocity Build high-quality, well-tested production code for systems that are secure, reliable, and maintainable at sc
Senior Cloud Security Engineer At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale's security needs are growing as we operate more production and cloud infrastructure for larger and more demanding customers. We're looking for a Senior Cloud Security Engineer to own the security of that infrastructure. This is a hands-on, high-ownership role: you will own how our production and cloud environments are hardened, isolated, and monitored. You will set and drive the direction for infrastructure and production security, reporting to the Head of Security and partnering closely with the wider engineering organization. This role is based in the San Francisco, Bay Area. In your first year, success looks like hardened and well-segmented production environments, strong runtime security coverage across our container footprint, and a clear, defensible story for how we secure the infrastructure our customers rely on. What You'll Do Own the security posture of Anyscale's production and cloud infrastructure across AWS and Azure, including hardening, network segmentation, and tenant isolation. Own runtime security coverage across our Kubernetes environments, from deployment through detection of anomalous activity. Partner with engineering on secure infrastructure architectur
About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible. Responsibilities and Duties Develop, own, and maintain tools and services to support the software build and release process Deploy and maintain services with Kub
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox’s daily active users growing at a record pace, we are seeking a senior data infrastructure engineer to join our new Data Insights team. Our team owns the data tooling that empowers Roblox builders to independently make informed and timely data-driven decisions. As an engineer on the team, you’ll work on the platforms behind tools like Superset, Hex, and Python notebooks, which provide critical insights into the health of our business to users at every level of the company. We tackle diverse challenges in data engineering, infrastructure, and analytics, to deliver the insights our customers need. You will collaborate closely with engineers across our data ecosystem to shape the future of product analytics at Roblox. This role offers the chance to be a founding team member and help define both the technical direction and the long-term shape of the product area from the ground up. This role is a great fit for you if you are proficient in designing and scale robust data infrastructure and applications and have a zeal for developing inspiring, easily maintainable, and reusable code. Join our team and make a significant impact at Roblox. You Will: Architect and deliver a high-pe
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Controls Network Engineer, you will design, validate, and scale the controls and OT network architectures that support high-density AI data centers. You will work across controls systems, OT infrastructure, telemetry, commissioning, deployment, and operations, partnering with mechanical, electrical, IT/networking, security, and external delivery teams. About the Role We are seeking a mid to senior OT Network Engineer with a strong controls systems background to lead the design and operation of resilient, secure, and scalable OT network architectures for high-density AI data centers. This role translates compute, power, cooling, and operational requirements into practical OT network designs, evaluates vendor solutions, and drives technical decisions across controls infrastructure, telemetry, commissioning, and operations. The ideal candidate has strong hands-on experience in mission-critical OT environments, including industrial networking, virtualized infrastructure, and OT network operations, with expertise in routing, switching, segmentation, firewall policy, time synchronization, monitoring, and network lifecycle support. Key Responsibilities Define controls, automation, and OT network requirements for AI data center campuses. Develop reference architectures, engineering standards, and reusable design templates. Review and develop basis-of-design and functional design documents, including OT network diagrams, IP/VLAN schemes, telemetry architectures, data flow diagrams, and commissioning requirements. Design OT and infrastructure network architectures, including physical topology, logical topology, IP addressing, subnetting, VLA
Become a part of our caring community *(Selected candidate will be required to live within 60 mins of one of the following metro locations OR be willing to relocate to within 12 months of hire date: Louisville KY, NYC Metro, Dallas Metro, Charlotte NC Metro, Tampa, Miami, Washington DC metro, Chicago, Boston, Atlanta, Nashville) The Senior Security Architect for AI works with EIP Department leaders and Humana enterprise stakeholders to identify, define, and develop security architecture requirements and secure designs for AI technology solutions across Humana's business, information technology, and security domains. The Security Architect leads the development of technical architecture & designs, develop security requirements, perform threat modeling and ensures alignment of security & risk imperatives with business priorities. The Security Architect is responsible for the high-level design and patterns of security program infrastructures to enable the protection of Humana tools, data, systems, and networks. Working with EIP Leaders, the security architect drives alignment between the EIP security strategy, security architecture and infrastructure, and Humana's overall business and technology strategic priorities. In this capacity, the role is responsible for planning, designing, and proposing architectural patterns or security enhancements for Humana's information security and technology infrastructures. The role works with EIP and Humana leaders to review and reconcile Humana business priorities with EIP security requirements. The role engages with relevant EIP stakeholders to identify current and emerging security threats and works to design security architecture elements to mitigate threats as they emerge. Additionally, the role actively works to identify gaps within existing EIP reference architecture and designs updates to impacted sec
Get new senior infrastructure architect jobs by email
Daily job updates · Unsubscribe anytime