Staff Embedded SW/FW Engineer (Bringup) Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Bengaluru which will play a central role in Graphcore's work building the future of AI computing. Job Summary We have an exciting opportunity to be part of a collaborative, cross-functional development team developing C code used to validate cutting-edge, high-performance AI chips and platforms. You will play a critical role in supporting new product introductions and post-silicon validation. Working within the Post-Silicon Bringup team, you will be involved with bringing first silicon to life, developing code primarily in C to configure and exercise systems and sub-systems on new silicon devices, and working closely with many other teams to help it become a fully characterised and working product, reporting project status/progress to program management on a regular basis. You will have the opportunity to, and be responsible for, leading, mentoring, and providing technical guidance to other engineering team members. In this role, you can leverage your experience and industry knowledge to architect and drive implementation of continuous improvements to test infrastructure and processes. The Team The Post-Silicon Bringup team sits within the Architecture and Validation team, we are responsible for bringup and validation of new silicon when it returns from manufacture, enabling and supporting the production SW and FW teams to bring up their software and supporting the Silicon Characterisation team. Respons
Jobiba hiring network
Linux System Administrator Jobs
742 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current linux system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team An exciting opportunity to join a new team within the Software Operations group. The Build Engineering team is a new function within Software Infrastructure, which focuses on the overall process of building and integration of the Machine Learn ing S oftware S tack. You will work closely with the QA and development teams to get an understanding of how our ML SW stack is built, helping to ensure good build practices, and proving that the stack works together and is reproducible in secure, sandboxed environments. Responsibilities and Duties Developing our internal t
About us Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Job Summary Working within the Silicon verification team, the silicon verification engineer is responsible for a wide range of tasks within the silicon verification team. This person is responsible for verification activities within Graphcore, helping the silicon team meet the company objectives for quality silicon delivery. The Team The verification team sits within the Silicon design team. We are responsible for ensuring that the RTL created by the logical design team and used by the physical design team matches the architecture specification for Graphcore silicon. Responsibilities and Duties V erification activit i es within the verification team Ensuring good communication between sites Verification planning, specification and closure of functional coverage Providing feedback to architects Test generation and failure diagnosis/triage Contributing to shared verification infrastructure Candidate Profile Essential: verification experience in relevant industry Proven leadership and planning skills Be highly motivated, a self starter, and a team player Ability to work across teams and programming languages to find root causes of deep and complex issues Ability to research along with the knowledge to solve complex problems Presents technical and functional knowledge to design experiments/ projects that contribute to overall
About the Team DoorDash Labs is a team within DoorDash building autonomous delivery robots and other autonomy solutions from the ground up for DoorDash's core delivery platform. If you have a passion for applying robotics solutions to a service loved by millions of people, then we want to talk to you! About the Role As Team Lead, Autonomy Tech Support, you will lead the day-to-day execution and development of the Mexico City Autonomy Tech Support team while maintaining a strong understanding of its technical workflows. You will set a high bar for troubleshooting, escalation quality, documentation, and operational readiness as the autonomous delivery fleet continues to scale. You will report into the Manager, Autonomy Tech Support on our Autonomy Tech Support team in our DoorDash Labs organization. Work model: 100% in-office in Coyoacán, Mexico City, Open for T4 levels, Across all LOBS. You’re excited about this opportunity because you will… Lead the day-to-day execution of the CDMX ATS team, including coverage, workload, priorities, and high-impact fleet issues. Coach and develop ATS Specialists through regular feedback, technical coaching, and hands-on support. Maintain a high bar across troubleshooting, escalation quality, Failure Mode execution, documentation, and cross-functional communication. Serve as the first leadership escalation point in CDMX for complex or high-urgency issues. Be a first line of defense for potential software regressions and emerging fleet-level trends, ensuring abnormal behavior is identified, validated, and escalated quickly. Lead AI adoption, integration, and enablement across the CDMX ATS team, ensuring AI capabilities are effectively incorporated into day-to-day workflows while identifying opportunities to improve troubleshooting, decision-making, and team efficiency. Partner with Operational Intelligence and Operational Reliability to ensure the team has the tools, systems, and processes needed to operate effectively and effic
About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta enables universal, secure access to technology for every user and organization. The Device Identity and Access team ensures that every endpoint interacting with corporate resources is trusted, healthy, and secure. Our pillar spans four domains: Device Identity, Device Authentication, Device Security Posture, and Endpoint AI Security. We are hiring a Senior Product Manager to drive our Zero Trust Authentication group. In this role, you will lead the core execution, performance, and deployment strategy for Okta’s foundational endpoint products. You will own the operational excellence of Okta Verify , FastPass , and Device Assurance , ensuring that our core passwordless authentication flows and device posture engines remain performant, secure, and incredibly easy for global enterprises to deploy at scale. This is a critical, high-impact role designed for a highly autonomous product manager. You will take on the tactical execution and smaller strategic investments for our endpoint authenticators, acting as the team's operational anchor so that our senior product leaders can continuously expand our strategic frontiers. What you will own You will bridge user experience and deep systems engineering to ensure our core authentication products scale reliably for millions of global users. Core Device Assurance Execution: Own the roadmap for our baseline security posture checks (such as OS version validation, disk encryption, and firewall status). You will m
As a Staff Software Engineer on Coder’s Agentic Engineering team, you’ll shape the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on while setting the team's technical direction. You’ll lead complex work, make sound architectural decisions, and help other engineers do their best work. What you’ll do here Set technical direction across Coder’s agent harness, integrations, and workflows. Design and build production systems in Go, with work across React and TypeScript where needed. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Lead complex projects from early ambiguity through production. Raise the engineering bar through design reviews, code reviews, and technical mentorship. Partner with Product and Design on clear, useful agent experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Deep experience building and operating production software systems. Strong hands-on experience with Go. Experience with React and TypeScript. Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS. A track record of setting technical direction without formal authority. Strong architectural judgment and comfort working through ambiguity. Someone who makes the engineers around them better. Our tech stack Backend: Go, Postgres Frontend: TypeScript, React Infrastructure: AWS, Kubernetes Observability: Prometheus, Grafana CI/CD: GitHub Actions Bonus tacos if you have (Tacos? If you need an ice-breaker, ask how we say thanks by giving tacos!) Experience building coding agents, developer tools, or cloud development environm
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We're hiring experienced performance engineers to find and land performance wins across the Supabase stack — from the query optimiser to block device level, and across the network paths and cloud infrastructure that connect our distributed platform. This is a performance delivery role: real, measured improvements to our database, platform, and APIs as we scale toward enterprise and platform-grade workloads. The methodologies you develop and tooling you build exist to multiply that delivery. You'd join a new performance team as one of its first members, working alongside a deeply seasoned performance engineer, with the team growing to roughly six this year. There's no legacy process to inherit — you'll help define how performance engineering works at Supabase. What you'll do Find bottlenecks in live production systems and characterize them precisely enough that the owning team can act on them without you in the room. Partner closely with database teams (e.g. Multigres, OrioleDB) and infra teams to land concrete performance improvements. Build, communicate, and evolve performance methodologies and tooling that turn live production data into actionable insight. Partner with the observability team to capture the right signals and establish a unified view of platform performance (latency, throughput, tail behavior) across products. Help teams self-serve performance analysis and make performance a first-class engineering concern. Problems you might work on Tune the kernel's paging behavior to production workloads Reduce memory overhead in OrioleDB and Multigres Identify and eliminate sources of latency spikes from application to kernel level Identify scalability bottlenecks across the
ROLE DESCRIPTION: We’re looking for a Senior Platform Backend Developer who can help us support the development organization to deliver value to customers in a reliable, efficient, and safe manner. You’ll be working in a focused team that owns one or more pieces of the production application environment and the developer experience, you will own and deliver in service of quarterly goals on the team. ABOUT THE TEAM: This role is within our Backend Platform team. The team primarily uses Go, Scala, and PHP and has expertise in technologies such as Kafka, various AWS services, and some infrastructure-as-code tools. Your primary focus will be on developing services and tools for our product development teams as well as modernizing our existing platform. Based out of British Columbia, you will report to the Senior Manager, Software Development, DevOps. WHAT YOU’LL DO: Design and build software - tools, libraries, automation, services, and glue scripts Responsible for the reliability, security, and integrity of our large, cloud-based platform Participate in a flexible on-call rotation Lead by owning project milestones, epics or features Practice continuous improvement, contributing to culture, process, and direction in your team and across our department Develop processes and automation to eliminate repetitive tasks Design and build our infrastructure platform Identify and implement new platform features Research and evaluate new technologies Refactor, rewrite or retire existing platform features Operate our developer experience and production application environments Diagnose and repair our distributed systems Perform maintenance, upgrades, and migrations Control or eliminate repetitive tasks, alert noise, and business-as-usual work Enable development teams Provide executable interfaces to our infrastructure platform Provide tools and best practices to support the entire software development lifecycle Collaborate with others across the orga
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity Are you passionate about building foundational technology that fuels the world’s digital innovation? At New Relic, we provide the leading unified data platform for all things observability – helping engineers, developers, and operators make decisions using data at every stage of the software lifecycle. Our New Relic Control team is at the heart of that mission, creating groundbreaking capabilities that orchestrate, manage, and optimize observability agents and telemetry pipelines at scale. We’re looking for a Senior Product Manager to lead a new strategic initiative within our Pipeline Control team. In this role, you’ll be responsible for defining vision, strategy, and roadmap for an enterprise-grade solution that orchestrates observability pipelines to streamline telemetry data in flight. You'll collaborate closely with engineering and product design to solve complex challenges around instrumentation, configuration, and data management with new systems and experiences. If you love turning ambitious ideas into impactful enterprise products that delight customers, let's talk! What You'll Do Define and drive the product strategy for our next-generation observability pipelines solution and align your vision with broader company goals Lead cross-functional collaboration with Engineering, Product Design, Sales, and Marketing to translate customer insights and technical opportunities into impactful product capabilities Evangelize the product vision internally
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, Atlanta Software Engineer, Network Firewall Role Summary As a Software Engineer on our team, you will work across a wide range of technologies and systems to deliver new features, improve performance, and increase the scalability of our Network Services products. You will help evolve our Network Firewall and related security capabilities across both north-south and east-west network architectures, with demand
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX About the Role Cloudflare’s Emerging Technologies & Incubation team builds and launches bold new products on Cloudflare’s edge network. This role is on the Durable Objects team, which powers stateful serverless applications for Cloudflare customers. You’ll help evolve the runtime and the low-level routing and storage systems that support real-time chat, multiplayer games, AI agents, and other stateful workloads
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Senior Network Engineer. About the job This position is needed to help build, operate, and maintain the global corporate network infrastructure, VPN, ZTNA, and connectivity to our cloud-based systems. As part of the Identity and Network Trust team, you will monitor network performance to proactively respond to incidents. You will need to possess the technical knowledge to troubleshoot and perform root cause analysis for escalated connectivity issues. You also will collaborate with the security team and other engineering teams to develop a secured global network infrastructure. Responsibilities Build, operate and support the global Twilio IT network infrastructure. Keep track of network assets, including network hardware and software. Monitor the network performance, availability, utilization, throughput and latency, providing measurable metrics to the IT leadership team. Work closely with stakeholder teams to determine c
The Team We are Datadog’s in-house product experts. The Technical Solutions team enables Datadog’s worldwide growth by educating potential partners and ensuring that our integration ecosystem is high-performing, secure, and valuable. Partner Technology Solutions Engineers (TSEs) are the technical bridge between Datadog and our third-party developer community. We act as consultants, helping partners build world-class monitoring solutions on the Integration Developer Platform (IDP) . The Opportunity Datadog is looking for a Partner Technology Solutions Engineer to join our fast-paced team. You will be the primary technical contact for our partners, guiding them through the entire integration lifecycle—from initial architectural design to final publication on the Datadog Marketplace. This is a unique role that combines deep technical troubleshooting with high-level consulting and platform advocacy. You will work directly with external developers and see your contributions immediately reflected in the Datadog ecosystem. You Will Act as the technical lead for partners, advising on OAuth flows, log pipelines, OpenTelemetry, and agent-based vs. API-based configurations Perform architectural assessments and deep-dive code reviews for partner integrations in the integrations-extras and marketplace repositories, ensuring they meet our Quality Rubric Solve complex technical challenges for partners via Zendesk, Slack, and dedicated technical consultations Identify friction points in our Integration Developer Platform (IDP) and partner with our internal Product and Engineering teams to build a better developer experience Maintain public-facing developer documentation and internal tracking systems ( JIRA ) to ensure transparency and scale You Are A technical expert with 3+ years of experience in a technical role (Support Engineering, Solutions Architecture, or Software Development) Proficient in at least one language (Python or Go preferred) An observability enthusiast who unders
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our
Get new linux system administrator jobs by email
Daily job updates · Unsubscribe anytime