Jobiba hiring network

Senior System Software Safety Engineer Jobs

7,101 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior system software safety engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions. Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency. About the Role We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments. This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment. The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries. Key Responsibilities Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure. Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios. Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize

pythonawsgit
View job →
G
Gitlab
📍 United States• Full-time• Remote• From $115.2K/yr
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Backend Engineer on the Chat Engine team, you'll build the engine behind GitLab Duo Chat, the conversational AI experience for GitLab. You'll work primarily in Python to build and maintain our agentic runtime: the Flow Registry, LangGraph flows, and the Duo Workflow Service. You'll also work in the GitLab Rails monolith, where Chat connects with the product. You'll own scoped parts of the system, ship small features and improvements with minimal guidance, and collaborate with the team on larger projects. You'll work alongside senior and staff engineers who will partner with you on design and support

REMOTEpythongitrest
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an OS / K8s Systems Engineer at Baseten, you’ll build the automation and systems that turn raw GPU hardware into production-ready compute. From provisioning to orchestration, you’ll own the software layer that makes our infrastructure reproducible, scalable, and reliable across data centers. This is a senior, hands-on role focused on building systems not operating them. You’ll work close to the metal designing OS images, building provisioning pipelines, and automating cluster bring-up from scratch. Your work will define how quickly we can turn new capacity into usable compute. EXAMPLE INITIATIVES Zero-to-cluster automation Build workflows that take new hardware from unprovisioned to fully operational cluster. Provisioning systems Design PXE-based or equivalent systems for imaging and lifecycle management. Reproducible infrastructure — Ensure clusters deploy consistently across data centers. RESPONSIBILITIES Own the end-to-end automation of cluster bring-up and lifecycle management. Build and maintain OS images, provisioning systems, and configuration pipelines. Deploy and operate cluster orchestration platforms (Kubernetes, Slurm, or similar). Design systems for reproducibility across sites and hardware generations. Automate upgrades, rollouts, and failure recovery. Optimize system performance, including GPU utilization and networking. Partner with hardware and network teams to validate and improve system b

pythonkuberneteslinux
View job →
G
Gitlab
📍 United Kingdom; Remote, United States• Full-time• Remote
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As an Engineering Manager, Git at GitLab, you’ll guide a deeply technical team focused on building, maintaining, and providing expertise on the Git version control system. The team’s work spans upstream development of Git, support for teams across GitLab, new tooling, scalability improvements, new data formats, and ongoing maintenance of the Git codebase. This role is a good fit for you if you can guide senior and staff engineers through complex, long-horizon technical work while building a culture of technical rigor, clear accountability, and ownership. You’ll help the team balance upstream open source d

REMOTEgitrestai
View job →
C
Cloudflare
📍 Hybrid• Full-time• Hybrid• From $174K/yr
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: San Francisco, CA, Austin, TX or NYC About the role What you'll do As the Marketing Strategy and Operations Lead, you will be the strategic right hand to the Chief Marketing Officer and a senior operator for the entire marketing organization. You will design and own the operating system that allows a global marketing function to move with speed, alignment, and accountability — setting the strategic agenda, running the planning

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Team Lead for Initiator & Protocol Engineering, you will spearhead the critical bridge between our industry-leading FlashArray and the Linux/VMWare ecosystems. You will drive the performance and reliability of our storage protocol stacks—spanning NVMe over Fabrics and Fibre Channel—ensuring Pure Storage remains the gold standard for enterprise connectivity. Collaborating closely with cross-functional hardware and software teams, you’ll mentor a high-caliber engineering squad to solve complex kernel-level challenges and influence the global Linux upstream community. WHAT YOU'LL DO Own the Protocol Lifecycle: Lead the development, maintenance, and optimization of Linux and VMWare initiator stacks (NVMeoF, FC-SCSI, iSCSI) and target drivers to ensure seamless, high-performance integration with Pure FlashArray. Drive System Resilience: Architect enhancements for Fibre Channel and NIC driver stacks that improve RAS (Reliability, Availability, and Serviceability), specifically focusing on multipathing logic and link health monitoring. Technical Leadership & Mentorship: Guide a team of senior and junior engineers through complex project deliveries, conducting deep-dive code reviews and setting the technical bar for C/C++ and Python development within the kernel space. Solve the Impossible: Act as the final escalation point for the most challenging system-level bugs found in the field or internal testing, u

pythonawslinux
View job →
E
Everpure
📍 Prague• Full-time
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Engineering Manager for FlashArray File, you will lead a Prague-based team dedicated to evolving Everpure’s™ industry-leading file services. You’ll drive the mission of delivering enterprise-grade, high-availability file capabilities that empower global innovators to manage mission-critical data. This role differentiates itself by blending deep systems-level engineering with high-impact people leadership, requiring close collaboration with Product Management and global R&D teams to redefine the modern data experience. Your goal is to ensure our customers can seamlessly scale their file environments with the reliability and performance Everpure™ is known for. WHAT YOU’LL DO Drive End-to-End Execution: Lead the full software development lifecycle (SDLC) for core file features, ensuring the team delivers high-quality, scalable code from initial design through production rollout. Shape Technical Strategy: Partner with senior architects and Product Management to define the roadmap for multi-server and directory services, making critical trade-offs that balance innovative feature delivery with long-term system stability. Cultivate Engineering Excellence: Champion rigorous testing strategies and CI/CD health, fostering a "follow-your-feature" culture where engineers take ownership of operational excellence and defect prevention. Empower and Develop Talent: Manage and mentor a team of 5–8 engineers, providing t

awsci/cdrest
View job →
Z
Zscaler
📍 Netherlands• Full-time• Remote
17 days ago

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer to join our Cloud Infrastructure & Operations team. This is a remote role based in the Netherlands, reporting to the Senior Director, Software Engineering. As a Staff SRE, you will leverage your expertise in Linux/UNIX System Administration to build scalable infrastructure and manage platforms like Kubernetes using automation and high security standards. You will troubleshoot complex Linux networking and security issues, manage firewall technologies, and ensure secure access across our global platforms and applications. What you’ll do (Role Expectations) Create and maintain highly scalable solutions based on KVM LINUX, Kubernetes, and Public Cloud Providers Analyze and troubleshoot systems performance and issues across the OS and Applications Maintain platform security and observability using nftables and robust monitoring tools Manage and deploy systems and s

REMOTEpythonawsdocker
View job →
F
Fin
📍 Sydney, Australia• Full-time
1mo ago

Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. The opportunity We are looking for an experienced Account Executive to join our Sydney-based sales team and own new-business growth across Malaysia, Singapore, and Hong Kong.You will manage the full Mid-Market sales cycle, from outbound prospecting through to close, while building trusted relationships with senior customer stakeholders across the region. This role requires a consultative, value-based selling approach and the ability to adapt to different APAC markets and buying environments. What you’ll be doing Own new-business sales across Malaysia, Singapore, and Hong Kong. Build and manage a healthy pipeline across Mid-Market and Enterprise accounts. Generate and progress pipeline through a balanced inbound a

aigorust
View job →
D
Datadog
📍 New York• Full-time• From $99K/yr
1mo ago

Datadog's Finance team collaborates with teams across the organization, providing commercial, operational and analytical support to ensure that Datadog's business continues to scale rapidly and efficiently. The Financial Planning & Analysis (FP&A) team analyzes company financial data (revenue, customers, headcount, expenses, etc.) in order to support the business’ growth and success. As an analyst supporting the team, you will play a key role in delivering insights through the management of essential data infrastructure, including our financial planning tool, Pigment. Your role will be highly cross-functional, leveraging systems and data to unlock analytical capabilities for both FP&A and business leaders. Your role is critical in synthesizing information from across the organization to foster operational alignment and support informed strategic decisions. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the team’s forecasting and reporting software, Pigment, supporting data-driven insights through the development of dashboards and KPIs, both for standard FP&A reports and ad hoc projects Work cross-functionally with FP&A leaders to improve existing datasets and models Ensure data and system best practices in processes across the organization, including during planning and reporting cycles Represent FP&A in the data & analytics community, collaborating with analytics partners across the organization to democratize data and share insights Work on strategic projects and initiatives for senior management, assessing various business opportunities and proposing solutions Support Datadog’s data-based decision making and continued efficient growth Who You Are: 2+ years of professional experience in FP&A, Data Analy

pythonsqlrest
View job →
G
13 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

pythonartificial intelligenceai
View job →
🔔

Get new senior system software safety engineer jobs by email

Daily job updates · Unsubscribe anytime