Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation of AI computing systems. In this highly visible technical leadership role, you'll drive reliability from architecture through production, partnering across hardware, software, and manufacturing teams to build high-performance AI platforms that set the standard for uptime, durability, and quality. If you're passionate about solving complex engineering challenges and influencing products at scale, you'll have the opportunity to shape technology powering the future of AI. This role is hybrid, based out of Toronto, Canada. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You've spent 8+ years in reliability engineering, ideally in high-performance computing, AI hardware, or data center systems. You're comfortable with the statistical side of the job, HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA are all familiar territory. You can work through a technical problem in a thermal lab and then explain the risks and trade-offs clearly to leadership. You're good at bringing people together, mechanical, electrical, thermal, softw
Jobiba hiring network
Performance And Systems Engineer Jobs
6,348 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current performance and systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Position Location: Austin, TX | New York City, NY | Washington, DC About the Role Join the Workers Deploy & Config team as a Principal Engineer, the most senior technical role on the team behind Cloudflare’s serverless edge developer platform. You’ll build the critical large-scale systems that let developers deploy, configure, and manage Workers globally, powering everything from simple static sites to full-stack applications serving millions
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Position Location: Austin, TX | New York City, NY | Washington, DC About the Role Join the Workers Deploy & Config team, the engine behind Cloudflare’s serverless edge developer platform. You’ll build the large-scale systems that let developers deploy, configure, and manage Workers globally, from simple static sites to full-stack applications serving millions of users. This team powers the foundation behind much of Cloudflare’s developer platf
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . We're hiring a Senior Software Engineer to join the CB Node team within the Platform organization. This team runs all the blockchain nodes that power every asset Coinbase offers to customers, currently spanning 55 unique protocols across Ethereum, Bitcoin, Solana, and beyond. You'll focus on performance optimization and reliability across the blockchain platform stack, diagnosing root causes instead of over-provisioning, reducing infrastructure spend, and building the observability and health-check systems that keep our nodes reliable at scale. What you'll do: Own deep performance analysis and optimization across the blockchain node stack, diagnosing root causes of resource consumption and implementing system-level fixes that reduce infrastructure spend while maintaining reliability SLOs. Build observability, health-check, and automated failover capabilities for blockchain nodes, ensuring the team can detect degradation and respond proactively rather than reactively scaling resources. Drive a right-sizing initiative across the node fleet, profiling workloads, identifying over-provisioned instances, and establishing capacity models that balance cost efficiency with headroom for reliability. Partner with protocol integration engineers, wallet teams, and indexer teams to extend performance improvements across the full blockchain platform stack, not just the node la
Want to be a bswifter? At bswift we’ve been transforming benefits administration since 1996, making it simpler, smarter, and more human. Our state-of-the-art, cloud-based technology and services empower employees to understand, manage, and love their benefits. From downtown Chicago, and remotely across the country, we serve thousands of companies and millions of people nationwide, reducing administrative burdens and freeing HR teams to focus on creating thriving, people-first workplaces. We’re looking for motivated and goal-driven individuals who share our passion for delivering excellence and creating solutions that make a difference. The reward is a fun, flexible and creative environment with ample opportunity for professional and personal growth. If you love the bswift values of pursue excellence, embrace accountability, deliver superior service, and be a great place to work, we want to hear from you! About the Role We are looking for an AI Engineer II to design, build, and deploy cutting-edge generative AI applications , including agentic workflows and intelligent chat experiences , that transform how employees interact with their benefits. This role goes beyond individual contribution—you will own complex AI features end-to-end , influence architecture decisions around LLMs and agent systems , and mentor junior engineers . You will play a key role in bringing secure, scalable, and high-performance AI solutions to production using AWS cloud technologies. You will collaborate closely with software engineers, architects, product managers, and cross-functional teams to deliver seamless and impactful user experiences Key Responsibilities 1. Solution Design & Development Lead the design and development of scalable, secure, and high-performance applications . Architect robust backend systems and integrate them with frontend applications and third-party services. Drive technical decisions and participate in design reviews alongside senio
We're looking for a Senior Platform Reliability Engineer who brings strong software engineering skills and a deep understanding of system behavior under load and stress. This role is a good fit for someone who wants to own reliability as a first-class concern – building the foundational systems that protect Asana's platform, not just responding when things go wrong. You'll build core platform systems like load shedding, rate limiting, circuit breakers, and traffic controls that protect Asana under real-world load. This is deep, cross-cutting work that shapes stability and performance of our entire infrastructure – and you'll partner closely with other platform teams to make reliability something that's built in, not bolted on. Our tech stack includes: AWS, Kubernetes (EKS), CloudFront, Istio, Cilium, MySQL (RDS), OpenSearch, DynamoDB, Redis, Terraform, Datadog, TypeScript, Scala, Go, and Python. (Yeah, we know this sounds like buzzword bingo – but we want this post to actually show up in your searches.) Why this role? Reliability as a first-class feature : You won't be patching things up after the fact. You'll build the systems that make Asana resilient by design. Foundational work : Load shedding, traffic management, ingress/egress – these are the building blocks that protect everything else. You'll own them. Strong collaboration, reasonable hours : You'll work closely with infrastructure teams in Warsaw and Reykjavik, making deep collaboration practical without constant timezone gymnastics. Room to grow : This is a new team, and you'll help shape what Platform Reliability Engineering looks like at Asana – whether that means leading projects, mentoring others, or defining our technical direction. In this role, success means shipping systems that other teams rely on by default – because they make the platform safer, not because they're mandatory. We're especially interested in people who think like backend engineers but obsess over failure modes, capacity plan
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Join Tenstorrent’s AI Models team and work at the layer most ML engineers never see: bringing advanced models to life on custom AI hardware. You’ll own real workloads end‑to‑end including porting, tuning, and validating LLMs and vision models on our accelerator, and chasing down every last millisecond and percentage point of accuracy. This role is for people who love the craft of ML engineering and want their work to matter at silicon scale, not just behind another API. This role is hybrid , based in Cyprus. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Bring up, run, and debug modern ML models (e.g., transformers) using PyTorch or TensorFlow. Analyze model behavior and performance, and identify bottlenecks across the stack. Improve efficiency, correctness, and scalability of model execution in real systems. Work closely with compiler, kernel, and hardware teams to drive performance and system-level improvements. Help translate state-of-the-art model architectures into production-grade, high-performance deployments. What We Need Strong experience building and working with ML models in PyTorch or TensorFlow. Strong understanding of mod
About the Team The Storage teams build and operate online stateful systems and abstractions that are reliable, efficient, secure and easy to use for DoorDash Engineering. The teams are responsible for understanding Product Engineering’s evolving needs and developing platform and infrastructure capabilities to serve them. The team currently supports CockroachDB, Cassandra, Kafka and Redis as well as data abstraction services to reduce the complexity of interacting with storage systems for Product Engineers. About the Role The Storage team is building and operating a high-performance, scalable, and reliable data abstraction layer that optimizes both efficiency and reliability. Our goal is to create a platform that manages itself and fades into the background—empowering engineers to focus on delivering product experiences our customers love. This role is available across two teams within Storage, each solving unique and high-impact challenges: One team is building the orchestration layer for DoorDash’s storage platform—unifying lifecycle management, operations, and self-serve APIs for databases and streaming systems, turning complex, stateful infrastructure into reliable, developer-friendly services used across the company. One team builds and operates the distributed data platform powering DoorDash's largest stateful workloads -- including Cassandra, which backs critical product surfaces across DoorDash, Wolt, and Roo. You'll design high-throughput data abstractions, smart clients, and platform services that make distributed data reliable and easy to work with at multi-petabyte, multi-million-QPS scale, with opportunities to go deep on distributed systems internals and contribute to the open-source Cassandra ecosystem. If you're passionate about distributed systems, developer experience, and building foundational infrastructure at scale, we'd love to hear from you. You must be located in San Francisco, Sunnyvale, Seattle, or the New York Metro Area for this hybrid pos
About the Team Consumer Monetization builds the experiences and systems that power how customers purchase and pay for OpenAI products. Our scope spans purchasing flows, payments, subscriptions, and billing, along with the shared capabilities that support new products, offers, and distribution channels. We own both customer-facing experiences and the underlying platforms that power them. We partner closely with Product, Design, Growth, Data Science, and engineering teams across OpenAI to make purchasing effective and reliable, and new offerings easier to launch and monetize. About the Role We’re looking for experienced Staff+ engineers with deep iOS or Android expertise and demonstrated experience contributing beyond mobile to frontend web or backend development. You’ll shape and build purchasing and subscription experiences across mobile applications, web, and supporting product APIs. The work spans improving conversion, performance, and reliability, building reusable components, and enabling new product launches. You’ll combine product judgment with technical depth to set direction, lead initiatives across teams, and remain hands-on through implementation and delivery. Approximately 50% of the work will initially be native mobile development, with the remainder across web and product APIs. We’re looking for engineers who enjoy working across the stack and have concrete examples of doing so professionally. Depth in either iOS or Android is required; experience in both is not required. This role is based in San Francisco, with three days per week in the office. In this role, you will: Architect and build purchasing and subscription experiences across native mobile, mobile web, and supporting product APIs. Partner with Product, Design, and Growth to identify customer needs, prioritize improvements, and shape technical direction for purchasing experiences. Improve conversion, performance, and reliability through experimentation, product analytics, and customer insights
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building the world’s fastest, most efficient AI compute clusters. TT-Fabric is the high-performance nervous system of this platform: the low-level networking layer that lets thousands of RISC-V and AI processors snap together into a single, massively parallel distributed supercomputer. If you love squeezing nanoseconds out of hot paths, designing protocols that move data at absurd scale, and turning messy hardware constraints into elegant distributed systems, this is an opportunity to shape the fabric that future AI models will run on This role is hybrid based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who We Are Strong systems engineer with deep C or C++ experience and comfort working in low-level or bare-metal environments. Passionate about hardware-software interaction, performance tuning, and eliminating inefficiencies at the protocol level. Curious about networking, synchronization, and communication across large clusters. Comfortable reasoning from first principles and challenging industry conventions. Motivated by building infrastructure that directly impacts large-scale
About the Team OpenAI, in close collaboration with our capital partners, is building the world's most advanced AI infrastructure ecosystem. The Scaling Analytics team serves as the data backbone for this effort, enabling leaders and operators to make informed decisions across infrastructure deployment, hardware operations, supply chain, capacity planning, and site execution. As OpenAI’s Industrial Compute expands across an increasing number of global data center campuses, the complexity of managing infrastructure capacity, hardware health, supply flows, and operational performance continues to grow. Scaling Analytics develops the data models, pipelines, metrics, and reporting systems that transform fragmented operational data into actionable insights, helping OpenAI operate infrastructure at unprecedented scale. About the Role We are seeking a Data Engineer to help build and scale the analytical foundations that power OpenAI's infrastructure organization. This individual will partner closely with Hardware Operations, Capacity Planning, Supply Chain, Infrastructure Delivery, Finance, and Engineering teams to create reliable data products that support critical operational and strategic decisions. Today, much of the team's expertise is concentrated within several highly specialized domains including hardware health, GPU attribution, and supply analytics. As Stargate grows and new sites come online, the demand for analytics support continues to expand across both existing and emerging problem spaces. This role will increase the team's ability to move quickly, reduce operational bottlenecks, and provide additional depth across critical infrastructure analytics functions. The ideal candidate combines strong data engineering fundamentals with an ability to navigate ambiguous operational environments, translating complex infrastructure problems into scalable data solutions that improve visibility, decision-making, and execution. Key Responsibilities Design, build, and maint
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations - Bengaluru, India About the Role The Data Intelligence & Analytics organization builds the core data platform and internal products that power decision-making across the company. We design and operate large-scale data systems, own the company’s data lake, ingestion infrastructure, and platform tooling, and develop end-to-end applications that transform complex datasets into fast, reliable, business-critical products u
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: Millions of people use Notion — and this number is increasing every day. That means millions of people trust us to deliver a fast, reliable, and secure experience, and we value this more than anything. We want to keep earning trust, while also continuing to amaze our users with the tools they can build in Notion. The AI Platform team is responsible for building the shared foundations that let Notion ship AI products quickly and operate them safely at scale. You’ll join a team of talented engineers focused on making speed and quality compatible: reliability and availability through provider changes, quality and correctness systems like evals and release gates, observability that makes failures explainable, and shared primitives for model integrations, context management, long-running actions, and cost/performance tradeoffs. Notion’s AI platform is vital to helping product teams move faster with production-grade guardrails as models, providers, and AI capabilities rapidly evolve. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days)
Senior Container Security Engineer – CVE Remediation & Image Hardening About the Role We are looking for a hands-on Senior Container Security Engineer to lead vulnerability remediation and image hardening across Linux-based container environments. This role focuses on deep operating system and container security engineering rather than simple vulnerability scanning. You will analyze, remediate, rebuild, harden, and continuously optimize container images used in modern cloud-native platforms. You will work closely with platform engineering, DevOps, infrastructure, and security teams to build automated remediation pipelines, reduce the attack surface, and deliver production-ready hardened images. What You’ll Do - Own end-to-end CVE remediation across Linux-based container images. - Analyze vulnerabilities across OS packages, libraries, runtimes, and dependencies. - Patch, rebuild, validate, and maintain hardened container images at scale. - Reduce attack surface by removing unnecessary packages, binaries, services, and dependencies. - Build and scale automated remediation pipelines for continuous image patching. - Improve image security posture while minimizing operational disruption. - Generate, validate, and maintain SBOMs to support supply chain visibility and compliance. - Integrate remediation workflows into CI/CD and GitOps pipelines. - Optimize image size, startup performance, and operational efficiency. - Research emerging Linux, container, Kubernetes, and software supply chain threats. - Troubleshoot complex dependency, package compatibility, and runtime security issues. - Help define internal standards for hardened images and secure software delivery. What You Bring - 5+ years of experience in Linux systems engineering, platform engineering, DevSecOps, security engineering, or SRE. - Deep understanding of Linux distributions (Debian, Ubuntu, Alpine, RHEL). - Strong hands-on experience with Docker, Kubernetes, and
We’re looking for Software Engineering Interns to help build and scale the systems that power Datadog’s observability and security platform. Interns contribute directly to real-world engineering challenges across backend, frontend, infrastructure, data engineering, and developer tooling while working alongside experienced engineers and mentors. You’ll help design, build, and improve systems that process and analyze massive volumes of metrics, logs, and application data in real time. Whether you’re interested in distributed systems, Kubernetes, AI-powered products like Bits AI, or developer platform tooling, you’ll work on meaningful projects that deliver impact to customers at global scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Contribute to production systems that process and analyze large-scale observability and application data in real time Build and improve distributed systems across backend infrastructure, developer platforms, and cloud-native services Help identify and solve performance, reliability, and scalability challenges in critical services supporting Datadog’s growing customer base Own and deliver technical projects from design through deployment with support from experienced engineers and mentors Develop technical expertise through hands-on experience with technologies such as Kubernetes, distributed systems, and cloud-native infrastructure Collaborate with fellow interns, mentors, and engineers while building software that delivers impact at global scale Who You Are: Pursuing a degree in Computer Science, Software Engineering, or a related technical field, or have equivalent practical experience Targeting a 2028 full-time start date Demonstrate strong computer science fundamentals, including data structures,
Get new performance and systems engineer jobs by email
Daily job updates · Unsubscribe anytime