Jobiba hiring network

Engineering Architect Jobs

8,135 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current engineering architect jobs. Use filters to narrow by work mode, employment type, experience and date posted.

G
23 days ago

Senior -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to deb

pythonlinuxai
View job →
G
23 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software, and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers, and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Validation leadership team, the Senior Execution and Quality Validation Engineer will be responsible for executing validation plans, developing automation solutions, and improving product quality across Graphcore silicon and platform technologies. Working closely with architecture, design, verification, firmware, software, systems engineering, and validation teams, the successful candidate will contribute to scalable validation methodologies, improve test coverage and execution efficiency, and help ensure products meet high standards of functionality, reliability, and performance before customer deployment. The role requires strong technical skills, attention to detail, and a passion for improving validation quality through effective execution, automation, and continuous improvement. The Team The Validation Execution and Quality team sits within the Validation organisation and is responsible for improving validation effectiveness, test execution efficiency, product quality, and release readiness across Graphcore silicon and platform products. The team develops validation methodologies, automation

pythonci/cdai
View job →
C-
CLEAR - Corporate
📍 New York• Full-time• From $130K/yr
23 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Solutions Engineer, you will be responsible for ensuring our customers successfully integrate with the CLEAR1 platform, with a focus on our Healthcare vertical. You will work across the Growth, Product, and Engineering teams to develop a deep understanding of the problems our customers are facing and influence how we’re building our platform for scale. We’re looking for a candidate with an excellent technical background who is looking to have a direct impact on customer and developer experiences. What you'll do: Lead presales customer engagements, including technical discovery, product demos, use case demos, RFPs, solution and integration scoping, and proof-of-concepts Drive customer integration for enterprise-level customers including building reference applications, prototypes, and architecture diagrams Develop code samples, integrations, and applications to demonstrate best practices and minimize customer development time, across a variety of languages and frameworks Support customers throughout the presales, proof of concept, and integration processes, assisting with common coding issues and patterns or troubleshooting Creating and maintaining public and internal documentation to help scale the team Act as the internal and external expert in the integration and usage of the CLEAR identity platform Collaborate across Sales, Product and Engineering to ensure that CLEAR strategically incorporates customer feedback and new product features into our roadmap Stay up-to-date with the latest trends and developments in the Health

gitrestai
View job →
CH
Cohere Health
📍 Hyderabad• Full-time
23 days ago

Opportunity Overview: We are seeking a Technical, Hands-on Manager to lead a team in building AI-driven healthcare enterprise applications . In this role, you will combine strong people leadership with deep technical expertise to guide the development of scalable, data-intensive solutions that drive meaningful business impact. As a leader who is still technically involved , you will mentor and manage data scientists and analysts while having the ability to guide the team through evaluating, selecting, and implementing the right models , paired with an understanding of end to end workflow and ensure the delivery of scalable AI solutions . This position requires excellent communication and collaboration skills as you partner closely with internal stakeholders and cross‑functional engineering, product, and clinical teams . In our fast-paced environment, adaptability is key—your ability to reprioritize quickly and lead your team through evolving business needs will ensure maximum impact What you’ll do: Lead, mentor, and develop a high‑performing team of Data Scientists and ML Engineers, ensuring strong execution and continuous skill advancement. Contribute to event‑driven architecture design and implementation, enabling asynchronous processing and large‑scale system integration. Work seamlessly across functions—partnering with Data Scientists on model tuning, experimentation, and prompt design; collaborating with Product and Software Engineering to embed AI/ML into user-facing applications; engaging with DevOps/Platform Engineering on environment setup, CI/CD, monitoring, and reliability; and working with Data Engineering on pipeline design and ingestion strategies. Provide technical leadership in the effective use of AWS services such as Lambda, EC2, EMR, S3, Athena, Batch, Textract, Comprehend, Bedrock. Drive the implementation of project scope definition, effort estimation, and planning in close coordination with cross-functional teams. Conduct code reviews, pr

pythonsqlaws
View job →
DU
DoorDash USA
📍 San Francisco• Full-time• From $1.6M/yr
23 days ago

About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus

pythonjavasql
View job →
DU
DoorDash USA
📍 San Francisco• Full-time• From $1.6M/yr
23 days ago

About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that will run DoorDash's data, analytics, and ML compute for the next five years and beyond. You will own the deep, cross-cutting problems that span the runtime, the shuffle service, the scheduler, and the overall service reliability — making the architectural calls that compound across the platform's lifetime. You will partner with the Engineering Manager on technical roadmap, hiring, and team shape, and act as the senior technical voice in cross-team partnerships with Data Engineering, ML Platform, and product engineering teams that depend on the platform. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Set the multi-year technical direction for an in-house Spark-on-Kubernetes platform — runtime, shuffle, scheduler, reliability — and make the architectural calls that compound for years. Own the deepest distributed-systems problems on the team: shuffle architecture, multi-tenant scheduling, runtime performance, and the failure modes that only show up at scale. Partner with the Engineering Manager on technical roadmap, hiring, inte

pythonjavasql
View job →
EA
23 days ago

Research Engineer, Applied AI Location: Bangalore (or throughout India remote-friendly with travel) About EnCharge AI: EnCharge AI is building the next generation AI platform. Our novel in-memory-computing architecture delivers a 10x step-function improvement in compute energy efficiency and performance for AI inference workloads. As the demands of artificial intelligence move beyond today's models, we believe fundamental underlying infrastructure must evolve. We are an experienced team of AI researchers, silicon & systems engineers, and architects backed by leading investors, poised to become the essential platform for the next wave of AI innovation. The Opportunity: Modern AI workloads—from large language models to diffusion-based generators to multimodal systems—represent some of the most compute-intensive frontiers in AI, and some of the most promising applications for our hardware’s energy efficiency advantages. We’re building a vertically integrated AI stack that will showcase the transformative potential of our silicon while delivering real value to customers today. We are seeking a Research Engineer to push the boundaries of AI model capability, quality, and efficiency. You’ll build fine-tuning and post training pipelines, develop rigorous benchmarking frameworks, and work at the intersection of ML research and hardware-aware optimization—ensuring our models run beautifully on our silicon. This is a role for someone who thrives at the boundary between research and engineering. You’ll read papers, implement techniques, and ship production-quality code—all in service of making AI inference faster, cheaper, and better. Key Responsibilities: Algorithmic Acceleration: Research and implement state-of-the-art techniques to accelerate AI inference—quantization, sparsity,

pythonaigo
View job →
O
OpenAI
📍 San Francisco• Full-time
24 days ago

About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We are looking for an experienced Mechanical Engineer with 7+ years of experience in design of IT hardware from chip/package to system levels. You’ll work alongside experts in thermal, mechanical, electrical, software, and systems engineering to support the design, analysis, and validation of mechanical and thermal systems that ensure the reliability, efficiency, and longevity of mission-critical hardware. This position requires strong analytical skills, hands-on testing experience, and the ability to work in a fast-paced, cross-disciplinary environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead mechanical design for AI supercomputer product in the data center application Collaborate with the cross functional team to design and optimize thermal solutions for data center hardware, including chips, power modules, and system-level cooling architectures Collaborate with cross-functional teams to integrate thermal management strategies into hardware design, from concept to mass production Design and validate mechanical systems, including chassis, enclosures, cooling systems, and high-power connections, ensuring alignment with performance and reliability standards. Perform 3D modeling, FEA, tolerance analysis, and prototyping, ensuring manufacturability and a

awsrestai
View job →
O
Okta
📍 Bengaluru• Full-time
24 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta’s TDI Network Engineering team is responsible for the global corporate network, building and supporting a high-performing, reliable network at scale. As a member of this team, you will have a direct impact on network design, deployment, and reliability, enabling our employees to work effectively from any location globally. Your role ensures the overall security and integrity of our corporate network by leveraging network security best practices, innovative products, and rigorous security validation. Reporting to the Network Engineering Manager, this operations-focused role is distinct from core Network Engineering and Network Security, centering primarily on operational execution—including responding to alerts, maintaining service availability, and ensuring system health across our global enterprise network. You will drive the strategic reduction of systemic toil and technical debt across multiple teams, applying a systems-level perspective and leveraging deep expertise in Distributed Systems, Networking fundamentals, Infrastructure as Code, and observability to architect scalable platforms and lead technical efforts to ensure an "Always Secure. Always On." environment. You will own multi-quarter objectives and establish long-term strategies for network reliability. What you'll be doing : Design and Own the resilience, health and availability of our entire global corporate network domain, managing operational responsibilities such as responding to aler

pythonawsrest
View job →
G
27 days ago

Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team The Network Security team is responsible for securing GoDaddy's global hybrid infrastructure across data centres, cloud, edge, and remote-access environments. We partner across Security, Infrastructure, Cloud, and Engineering teams to build scalable, resilient, and secure solutions that support the business. As a Principal Network Security Engineer, you'll act as a senior technical leader, helping define network security strategy, influence architecture across teams, and drive security outcomes at enterprise scale through technical expertise, systems thinking, and cross-functional leadership. What you'll get to do... Define and drive network security architecture across hybrid environments, including data centres, cloud, edge, and remote-access technologies Design trust boundaries, segmentation strategies, secure connectivity patterns, and network controls that reduce risk and improve security posture Lead complex technical initiatives, migrations, and architectural decisions while balancing security, reliability, performance, and operational requirements Establish scalable approaches for policy governance, automation, monitoring, telemetry, and security control validation Partner across engineering organizations to drive large-scale initiatives, mentor engineers, and influence technical direction through architecture reviews and technical leadership Your experience should include... 10+ years of experience in Network Security Engineering, Network Architecture, or Security Engineering, including ownership of enterprise-scale secu

awsci/cdai
View job →
B
Baseten
📍 San Francisco• Full-time• Remote
29 days ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui

REMOTEpythonawsazure
View job →
O
1mo ago

About the Team The ChatGPT Library team is building the place where people can save, organize, rediscover, and build on the content they create with ChatGPT. Our goal is to make ChatGPT more useful over time by helping users seamlessly return to important files, images, conversations, and other content across their devices. The team works at the intersection of product engineering, design, and AI research to create intuitive, personalized experiences that make users’ content easy to find and act on. On Android, we are focused on delivering fast, reliable, and deeply native experiences that put a user’s evolving body of work at their fingertips. About the Role We are looking for a senior Android engineer to help build the future of ChatGPT Library on mobile. You will own high-impact product experiences across the Android stack, shaping how millions of people save, organize, discover, and interact with their content in ChatGPT. In this role, you will: Build and ship new Android experiences that help users easily access, organize, and build on the content they create with ChatGPT. Own features end to end—from early product exploration and technical design through implementation, experimentation, launch, and iteration. Develop scalable, maintainable foundations that allow Libraries experiences to evolve quickly as new AI capabilities emerge. Improve the architecture, performance, reliability, and responsiveness of content-rich experiences across a wide range of Android devices. Thoughtfully integrate Android platform capabilities to create experiences that feel intuitive and native to mobile. Partner closely with product, design, research, data science, and engineering teams to translate emerging AI capabilities into useful, polished products. Help establish technical direction and raise the quality bar for Android development across the team. How We Work We care deeply about building products that are intuitive, useful, and trustworthy. We move quickly from ideas to wo

redisawsrest
View job →
B
Baseten
📍 San Francisco• Full-time• Remote
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture. EXAMPLE INITIATIVES The B200 Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's first Blackwell GPU clusters. Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems to move customer workloads seamlessly across regions to optimize cost and latency. Precision GPU Triage: Developing automated Go-based operators to identify, cordon, and repair unhealthy H100 nodes in under an hour. The Supply Chain of Intelligence: Partnering with lead

REMOTEpythonawsazure
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng

pythonsqlpostgresql
View job →
V
Vanta
📍 United States• Full-time• Remote
1mo ago

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. This is not a back-office compliance role and it is not a generic sales-engineering role. You will operate as a named member of deal teams under our pod model: paired with Strategic and Enterprise Account Executives, embedded in their weekly cadences, engaged from first discovery through POC, onsite, close, and expansion. You will be the practitioner in the room that a buyer's CISO or GRC lead trusts — and the internal expert our AEs, SEs, and marketing team build around. What you’ll do as a Subject Matter Expert at Vanta: Serve as the dedicated GRC SME for a book of Strategic/Enterprise Account Executives: join discovery and qualification calls at the earliest deal stages, scope compliance programs against Vanta's platform, and support demos, POCs, workshops, and customer onsites across Compliance, Third-Party Risk Management, Risk Management, and Trust/Questionnaire Automation. Advise prospects on program architecture: multi-framework strategy, shared controls, business-unit and workspace scoping, custom frameworks, and audit sequencing. Answer field questions through our SME channels at customer-forwardable quality — including reviewing and validating AI-agent-generated answers before they reach customers. Our team runs AI-first: you'll use agents daily, act as the quality gate on their output, and (ideally) build tooling of your own. Design and deliver enablement: live sessions for GTM teams, bootcamp scenarios and mock-customer roleplay, async curriculum modules, and review of GRC marketing and SEO content. Own monthly alignment cadences with sales front-line managers; feed structured product feedback to our Product and PM

REMOTErestaigo
View job →
🔔

Get new engineering architect jobs by email

Daily job updates · Unsubscribe anytime