At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions
Jobiba hiring network
Ai Senior Systems Engineer Jobs
10,000 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current ai senior systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an OS / K8s Systems Engineer at Baseten, you’ll build the automation and systems that turn raw GPU hardware into production-ready compute. From provisioning to orchestration, you’ll own the software layer that makes our infrastructure reproducible, scalable, and reliable across data centers. This is a senior, hands-on role focused on building systems not operating them. You’ll work close to the metal designing OS images, building provisioning pipelines, and automating cluster bring-up from scratch. Your work will define how quickly we can turn new capacity into usable compute. EXAMPLE INITIATIVES Zero-to-cluster automation Build workflows that take new hardware from unprovisioned to fully operational cluster. Provisioning systems Design PXE-based or equivalent systems for imaging and lifecycle management. Reproducible infrastructure — Ensure clusters deploy consistently across data centers. RESPONSIBILITIES Own the end-to-end automation of cluster bring-up and lifecycle management. Build and maintain OS images, provisioning systems, and configuration pipelines. Deploy and operate cluster orchestration platforms (Kubernetes, Slurm, or similar). Design systems for reproducibility across sites and hardware generations. Automate upgrades, rollouts, and failure recovery. Optimize system performance, including GPU utilization and networking. Partner with hardware and network teams to validate and improve system b
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role At Sentry, Support is an engineering discipline. Our customers are the greatest technical minds in the world—developers at elite enterprises building the future of software—and they deserve answers that go deeper than a knowledge base link. We're looking for an APAC Technical Support Engineer based in Australia to join our global Support Engineering team. This role is designed to overlap with our San Francisco headquarters; Monday thru Friday 9AM-5PM AEST. We are architecting the Technical Support engine . We’re looking for a veteran engineer to help us redefine the standard of technical support by combining deep human expertise with autonomous agentic systems. You are a debugger of both code and systems. You will treat support volume as a data signal to build automated resolution paths, ensuring our human engineers only touch the most complex, high-impact architectural puzzles. Sentry Support Engineers aren't just clearing queues; they are Orchestrators . You will engage with our users across GitHub, Discord, and our internal systems, while acting as the Technical Lead for our Agentic Ops. You ensure that when a developer asks a complex question, our systems have the right context and a seamless "Human-in-the-Loop" path to you when deep, nuanced expertise is required. In this role you will Master the Sentry Ecosystem & Support Elite Developers Deep-Dive Debugging: Perform root-cause analysis on complex issues and distributed tracing gaps across polyglot environments. Support the Great Minds: Act as a strategic consultant for senior engineers at our largest enterprise customers, solving high-stakes archite
Perforce is a community of collaborative experts, problem solvers, and possibility seekers who believe work should be both challenging and fun. We are proud to inspire creativity, foster belonging, support collaboration, and encourage wellness. At Perforce, you’ll work with and learn from some of the best and brightest in business. Before you know it, you’ll be in the middle of a rewarding career at a company headed in one direction: upward. With a global footprint spanning more than 80 countries and including over 75% of the Fortune 100, Perforce Software, Inc. is trusted by the world’s leading brands to deliver solutions for the toughest challenges. The best run DevOps teams in the world choose Perforce. About the Role The Perforce IP Lifecycle Management (IPLM) suite - used by semiconductor and hardware design teams to manage their intellectual property (IP), releases, and dependencies at scale - includes a Python command-line client (pi), the IPLM API server, the PiCache caching layer, and a Web based interface. As a Senior Software Engineer on the IPLM team, you will own end-to-end features across the CLI and PiCache. You will be expected to work closely with the server team that provides the API and database layers, as well as solutions and support teams to help resolve customer errors. This role blends product feature work, systems engineering, and modern AI-assisted development practices.
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role We are a team of high-output generalists where ML and systems engineering converge to push autonomy performance forward. As a Senior Perception ML Data Infrastructure Engineer, you will own the critical bridge between our autonomous vehicle hardware, our human labeling operations, and our ML models. You will take ownership of our core native perception data platform. This is not a standard web development role; the stack is deeply adjacent to our core robotics infrastructure. You will be dealing with massive, dense 3D point clouds at a scale that pushes the boundaries of industry state-of-the-art, alongside complex sensor parsing, pipeline propagation, and rigid performance constraints. You will operate in highly ambiguous environments, inheriting complex systems, establishing strict API boundaries, and building the "good enough, fast enough" infrastructure that guarantees our ML models learn from high-quality data. About the wo
About the Role: We're hiring Senior and Staff Data Platform Engineers to join the Data Infrastructure teams in Toronto. Together these teams own the infrastructure that processes billions of events per day: Spark-on-Kubernetes, Flink and Kinesis pipelines, a multi-petabyte Delta Lake, a large-scale MemoryDB feature store, Databricks multi-environment operations, and the catalog and lifecycle systems that govern it. The team is small and senior. Each engineer owns major platform components: you design it, build it, and support it in production. This is a hybrid-role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Spark-on-Kubernetes — EKS-based compute platform for Spark workloads: cluster configuration, Pod Identity IAM, job environment setup, Kustomize overlays, and shadow canary validation Event ingestion — Rust services and Flink jobs processing billions of events per day over Kinesis; throughput, reliability, on-call response, and AI-assisted operational tooling to reduce toil Platform infrastructure — Terraform modules for environment provisioning, cross-account AWS IAM, ARC runner infrastructure, and CI/CD for data platform changes Feature store and ML compute — Flink-based real-time feature pipelines feeding a large-scale MemoryDB cluster; GPU capacity governance and Databricks multi-environment operations for ML training workloads Workflow orchestration and CDC — Airflow-based DAG deployment, change data capture pipeline operations, and data quality monitoring Your Background: 3+ years building and operating production data platform infrastructure at the cluster or platform level, across Spark, Flink, Kinesis, Kubernetes, or equivalent Deep experience in at least one of: Spark-on-K8s cluster operations, Rust-based data or systems engineering, Kubernetes platform engineering and IaC, or data catalog and governance tooling Production AWS experience or equivalent: EKS, S3, Kinesis, and mu
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software, and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers, and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Validation leadership team, the Senior Execution and Quality Validation Engineer will be responsible for executing validation plans, developing automation solutions, and improving product quality across Graphcore silicon and platform technologies. Working closely with architecture, design, verification, firmware, software, systems engineering, and validation teams, the successful candidate will contribute to scalable validation methodologies, improve test coverage and execution efficiency, and help ensure products meet high standards of functionality, reliability, and performance before customer deployment. The role requires strong technical skills, attention to detail, and a passion for improving validation quality through effective execution, automation, and continuous improvement. The Team The Validation Execution and Quality team sits within the Validation organisation and is responsible for improving validation effectiveness, test execution efficiency, product quality, and release readiness across Graphcore silicon and platform products. The team develops validation methodologies, automation
NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and automations that power Smartsheet’s core business functions across Finance, Sales/GTM, and People & Culture. Our team owns mission-critical systems and workflows that enable how the company hires, sells, bills, pays, reports, and scales. We operate at the intersection of software engineering, enterprise platforms, and business-critical data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Automation team builds human-to-system and system-to-system automations that reduce manual effort and friction across the business. We combine cloud-native services, agentic AI, and workflow orchestration to enable employees to interact with enterprise systems through intelligent, secure, and auditable automation. As a Senior Software Engineer I (Automation), you will lead the design, build, and operation of systems and workflows that directly support business execution at scale. You will own complex technical initiatives, partner with Product Managers and stakeholders on technical roadmaps, and mentor junior engineers. This full-time position reports to the Sr. Director, Development and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Architect AI Agents: Take a leading role in designing Agentic Workflows using AWS Step Functions and Bedrock Agents that reason
Bloomreach is building the world’s premier agentic platform for personalization .We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey. We're taking autonomous search mainstream, making product discovery more intuitive and conversational for customers, and more profitable for businesses. We’re making conversational shopping a reality, connecting every shopper with tailored guidance and product expertise — available on demand, at every touchpoint in their journey. We're designing the future of autonomous marketing , taking the work out of workflows, and reclaiming the creative, strategic, and customer-first work marketers were always meant to do. And we're building all of that on the intelligence of a single AI engine — Loomi — so that personalization isn't only autonomous…it's also consistent.From retail to financial services, hospitality to gaming, businesses use Bloomreach to drive higher growth and lasting loyalty. We power personalization for more than 1,400 global brands, including American Eagle, Sonepar, and Pandora. About the Team: AI Search The AI Search team owns Bloomreach’s core search platform, serving hundreds of millions of queries per day across enterprise customers. We build and operate highly scalable, low-latency systems that combine traditional information retrieval with modern ML-driven ranking and semantic understanding — all in production at scale. This team sits at the intersection of systems engineering, search relevance, and applied ML , with direct impact on customer revenue and experience. The Role: As a Senior Staff Engineer , you will be a technical leader responsible for shaping the architecture and long-term evolution of Bloomreach’s Search platform. You will lead complex initiatives, set technical direction, and mentor engineers, while remaining deeply hands-on. This role is ideal for someone
Senior: GBP 73,500 - 99,500 Staff: GBP 97,300 - 131,700 Subject to alignment to the responsibilities and duties of the role - we currently have multiple positions available at both Senior and Staff level. About the job Build the Linux distribution foundation that turns upstream software into trusted Graphcore platform releases. You will help create the Linux distribution that powers Graphcore AI systems. The team produces production-ready system images from proven upstream distributions. Your work will shape how releases are built, validated and prepared for deployment. You will strengthen the engineering path from upstream Linux software to dependable platform releases. You will build and improve automated pipelines, run established Linux test suites, and diagnose issues across build and validation flows. As the platform evolves, you will introduce controlled configuration and tuning changes with evidence-led validation. This is hands-on systems engineering with visible impact. You will help define reliable processes for a new team building a critical part of Graphcore’s platform. The team and culture You will join one of Graphcore’s newest engineering teams, helping shape its culture from the start. It is a small, co-located team where ownership matters and progress is visible. Work happens through close technical discussion, practical problem-solving and evidence-led decisions. Ideas are challenged openly, and the best path wins regardless of hierarchy. You will report to a leader who values technical credibility and invests in people’s growth. The team moves with pace, takes responsibility and changes direction when the evidence demands it. What we're looking for Strong practical experience working in Linux environments Experience building or maintaining automated CI/CD pipelines for reliable engineering workflows Proficiency in Python, Bash or similar
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role We’re a team of high-output generalists where ML and systems engineering converge. This is not a "run the models" role. We reason from first principles about why a perception model learns what it learns, close the gaps that cap its performance, and raise the bar on the data and evaluation loop that drives autonomy. Your work will directly impact how autonomous systems understand rare scenarios, adapt to global geographies, and scale safely. About the work You’ll solve autonomy’s hardest data challenges through applied ML and systems rigor: Diagnose why perception models underperform on the long tail, and turn that into targeted data and training priorities. Design eval metrics and regression detection that tell us whether a model is ready. Curate and clean training data for segmentation and occupancy; hunt the data problems that silently cap performance. Run controlled ML experiments and ablations; cleanly sepa
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX About the Role You’ll help define how machine learning models run across Cloudflare’s global network, from frontier open LLMs and real-time voice models to customer-deployed models served on heterogeneous GPUs and next-generation accelerators. You’ll work with systems engineers, product teams, hardware partners, and AI/ML engineers to bring models into production with low latency, strong reliability, and effic
Come join the Server Ingress Security team, where we are rearchitecting MongoDB Server’s ingress networking to make MongoDB clusters even more secure. This new team is building the Atlas Network Protection layer, a set of performant, security-critical services that harden MongoDB's pre-authentication attack surface and provides the ability to respond rapidly to emergent threats. We are looking for talented Senior Engineers to join the team and be founding members, where you will play a crucial role in our multi-year roadmap. Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies security and systems engineering fundamentals to protect a popular database at scale, join us! We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Candidate Profile 5+ years of experience building production-quality systems software Experience with large backend/compiled codebases and performance-sensitive software, preferably in Rust Bonus points for experience working hands-on in security-sensitive or networking-adjacent domains Strong systems fundamentals, including multi-threaded programming and performance profiling. Bonus points for: Understanding of network protocols, TLS, and connection lifecycle management Familiarity with security concepts such as attack surface reduction, input validation, memory safety, and defense-in-depth architectures Excellent verbal and written technical communication skills, with a strong desire to collaborate with colleagues Strong time management skills and the ability to realistically assess project complexity B.Sc. in Computer Science or a related field, or equivalent practical experience, with strong competencies in data structures, algorithms, and software design/architecture. Interest in the theory and practice of high-availability, security-critical systems Position Expectations Design, implement, and operate production
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Security Operations Engineer at Brex, you will focus on preventing, detecting and responding to security threats across Brex's corporate and cloud environments. You will use existing systems and develop tools to improve our security capabilities. Our team is responsible for functions across corporate security, detection & response and infrastructure security domains; and we perform systems engineering and automation to support those functions. Security Operations is part of our wider Trust & IT o
Get new ai senior systems engineer jobs by email
Daily job updates · Unsubscribe anytime