Jobiba hiring network

Workload Porting And Performance Engineer Jobs

751 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current workload porting and performance engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Infrastructure organization builds and evaluates the systems that power advanced AI workloads. We work closely with hardware, modeling, and architecture teams to ensure that new platforms deliver real-world performance aligned with workload needs. Our team focuses on understanding workload behavior across evolving hardware platforms—bridging the gap between theoretical capability and observed system performance. About the Role We are seeking a Workload Porting & Performance Engineer to evaluate new hardware platforms by porting benchmarks and real-world workloads, analyzing performance, and identifying system bottlenecks. In this role, you will bring up workloads on new systems, characterize performance behavior, and adapt workloads to better utilize hardware capabilities. You will play a critical role in validating new platforms and ensuring that performance aligns with expectations across compute, memory, and networking subsystems. This role requires strong hands-on experience with performance analysis, workload optimization, and system-level debugging across hardware and software boundaries. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Port and enable benchmarks and real-world workloads on new hardware platforms. Evaluate system performance across compute, memory, storage, and networking subsystems. Identify and analyze performance bottlenecks and inefficiencies. Adapt and optimize workloads to better utilize hardware capabilities. Develop and run performance experiments and profiling workflows. Compare expected vs. observed performance and provide feedback to: hardware architecture teams performance modeling teams system and software engineers. Debug issues across the stack, including software, runtime, and hardware interactions. Provide actionable insights to guide platform readiness and deployment decisions. Qualifications E

awsrestai
View job →
T
Tenstorrent
📍 Australia• Full-time• $100K – $500K/yr
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a highly technical Senior Staff/Principal Engineer to lead the porting and enablement of critical AI workloads. You will be a primary driver in migrating compute workloads to RISC-V architectures, ensuring our hardware is optimized for real-world application performance. The ideal candidate has a strong background in DevOps, workload porting, or application enablement . While this is an individual contributor role at its core, you will have the opportunity to grow and lead a small, specialized team over time as our workload migration efforts scale. This role is remote, based in the United States or Australia. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A Technical Catalyst: You thrive on the challenge of driving AI hardware porting to RISC-V and making complex software stacks run efficiently on new hardware. Systems Expert: You possess deep knowledge of system software, compilers, or low-level OS internals. You are an expert in ARM or x86 environments and are ready to apply those skills to the RISC-V frontier. A Project Driver: You have the technical authority to lead the implementation of a compute migration to RISC-V through

awsaidevops
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About OpenAI OpenAI is dedicated to ensuring that artificial general intelligence (AGI) benefits all of humanity. Our mission requires building not only world-class AI models, but also the infrastructure that enables those models to be deployed reliably, efficiently, and at global scale. As demand for AI continues to grow, we are expanding the ways OpenAI can bring high-performance inference capacity online across a diverse hardware ecosystem. About the Team The GPT Infrastructure team builds software that turns advanced inference and optimization research into production products. One focus is enabling strategic infrastructure partners and accelerator vendors to qualify and onboard new compute without a bespoke porting and optimization effort for every hardware platform. We build the control planes, APIs, secure partner-side execution environments, evaluation systems, artifact pipelines, and operational tooling that make these workflows repeatable and trustworthy. The work sits at the intersection of distributed systems, AI inference, compilers and runtimes, performance engineering, security, and external partnerships. About the Role We are seeking an experienced systems generalist who can work comfortably across the stack to help build an automated inference optimization platform. Given a workload, target hardware profile, compiler and runtime context, and a trusted verifier, the system runs durable optimization campaigns that generate, compile, execute, grade, and improve candidate kernels, runtime configurations, and serving-stack changes. You will design both the OpenAI-hosted control plane and the partner-side software that evaluates candidates on real accelerator hardware. The product must keep long-running workflows reliable, make performance results reproducible, and maintain clear trust boundaries around sensitive model and hardware information. This is a deeply cross-stack role, combining strong software engineering fundamentals with systems thinking and

pythonawslinux
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing

pythonawskubernetes
View job →
T
7 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our diverse team of technologists have developed a high performance RISC-V-based CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a talented engineer to build and run the performance infrastructure shared by our RISC-V Software and RISC-V Performance teams. This is the plumbing that both teams' performance work stands on — benchmarking automation, data collection, and workload capture across silicon, FPGA/emulation platforms and performance models. You'll work across both teams, enabling performance and software engineers to spend their time on analysis and optimization instead of running experiments by hand. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Great at identifying problems and developing solutions, with a bias toward owning the systems you build. Enjoys building tools and optimizing workflows so other engineers can move faster. Strong Linux systems engineer, c

pythonlinuxai
View job →
T
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Join Tenstorrent’s AI Models team and work at the layer most ML engineers never see: bringing advanced models to life on custom AI hardware. You’ll own real workloads end‑to‑end including porting, tuning, and validating LLMs and vision models on our accelerator, and chasing down every last millisecond and percentage point of accuracy. This role is for people who love the craft of ML engineering and want their work to matter at silicon scale, not just behind another API. This role is hybrid , based in Cyprus. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Bring up, run, and debug modern ML models (e.g., transformers) using PyTorch or TensorFlow. Analyze model behavior and performance, and identify bottlenecks across the stack. Improve efficiency, correctness, and scalability of model execution in real systems. Work closely with compiler, kernel, and hardware teams to drive performance and system-level improvements. Help translate state-of-the-art model architectures into production-grade, high-performance deployments. What We Need Strong experience building and working with ML models in PyTorch or TensorFlow. Strong understanding of mod

awsaic++
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building next-generation processors and systems, bringing together world-class expertise across silicon, systems, and software. We are looking for a CPU Performance Modeling Architect to help evaluate, shape, and optimize the performance of future CPU architectures. In this role, you’ll use performance modeling, workload analysis, and deep understanding of CPU architecture to answer complex questions about how a processor should be designed. You’ll work closely with CPU architects, RTL designers, software and compiler teams, and system engineers to identify performance opportunities, evaluate architectural tradeoffs, and turn modeling insights into actionable design decisions. This role is hybrid, based out of Santa Clara, CA or Austin, TX. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A CPU architect, performance architect, or performance modeling engineer with experience influencing CPU architecture or microarchitecture decisions. You have a strong understanding of modern processor architecture and enjoy digging into why a CPU performs the way it does. You are comfortable combining hardware architecture, software, data, and

pythonawsai
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our Tensix team is building custom AI compute cores, RISC-V CPUs, and chiplet-based architectures for datacenter, edge, and automotive AI. Design Verification Engineers on this team validate compute IP and subsystems and build scalable DV infrastructure to keep verification fast, automated, and production-grade. This role is hybrid, based out of Toronto, ON, Austin, TX or Belgrade, Serbia. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced in modern verification methodologies with strong SystemVerilog skills and exposure to structured testbench development. Comfortable working from block-level to system-level verification and reasoning about microarchitecture behavior from specs and waveforms. Proficient in Linux environments with Python, or Bash scripting for automate builds, parse logs, manage CI pipelines. Skilled in coverage-driven verification and confident debugging across RTL, testbench, and workload scenarios. Motivated by AI hardware and eager to learn verification strategies for new architectures and domains. What We Need Contribute to verification of Tensix IP and subsystems from early planning through tape-out, owning cov

pythonawsci/cd
View job →
T
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a Staff Digital Design Engineer to help define, build, and optimize high-performance IP and SoC architectures for next-gen AI and compute workloads. This role is ideal for engineers who thrive at the intersection of microarchitecture, RTL implementation, and performance-aware design. This role is hybrid, based out of Toronto, Ottawa, Boston, or Austin. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A digital design expert with a deep understanding of computer architecture and IP microarchitecture. Skilled in RTL development (Verilog/VHDL) and familiar with full ASIC flows. Comfortable optimizing for power, performance, and area (PPA) under aggressive design goals. A naturally collaborative and technical engineer — you thrive in spec definition, peer reviews, and team-wide planning. What We Need Architecture and RTL implementation of Tenstorrent’s custom IP blocks and SoC components. Performance-aware design decisions for compute, interconnect, or memory-heavy blocks. Occasional contributions to validation using emulation, FPGA prototyping, or UVM flows. Strong synthesis and timing closure awareness to support backend

awsgitai
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are seeking a GCC Compiler Engineer to design, develop, and optimize compilers for next-generation RISC-V and AI compute architectures. You will work across hardware and software teams to improve performance, programmability, and integration of our custom toolchains into real applications. This role is fully hands-on and central to how developers interact with Tenstorrent hardware across both traditional compute and advanced machine learning workloads. This role is Hybrid, based out of Santa Clara, CA, Austin, TX, or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced compiler engineer with deep knowledge of GCC and LLVM internals, comfortable optimizing for custom hardware targets. Strong C/C++ developer with a solid grasp of algorithms, data structures, and performance analysis. Collaborative and analytical, able to work across hardware and software domains to deliver efficient, high-performance toolchains. Passionate about enabling breakthrough compute architectures through compiler innovation and software-hardware co-design. What We Need Design, develop, and optimize GCC and/or LLVM compilers for Tenstorrent’s cust

awsmachine learningai
View job →
T
Tenstorrent
📍 Toronto• Full-time• $100K – $500K/yr
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As a Software Engineer on the Acceleration Kernel Development team at Tenstorrent, you’ll work at the intersection of software and hardware performance. You’ll be writing low-level code that directly powers high-efficiency machine learning workloads, optimizing every cycle, every memory move, every instruction. If you're motivated by performance, precision, and real impact, this is where your skills will shine. This role is hybrid, based out of Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A developer who loves high performance code, parallel algorithms, wrangling bits, optimizing compute, and making hardware fly. Great in C/C++ and able to build fast, efficient code from the ground up. Obsessed with performance and precision, especially in ML workloads. Motivated by complex problems and thrives in collaborative, fast-moving environments. What We Need Expertise in building and optimizing compute kernels for parallel ML and high-performance workloads. Ability to analyze and tune instruction-level performance across latency, memory, and bandwidth. A collaborative mindset to work closely with ML engineers and integrate opti

awsmachine learningai
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We’re looking for a Staff Forward Deployed Engineer who’s excited to build with the engineers using the AI computers Tenstorrent makes. You will create continuity between customers, engineering, and AI inference service products. This is an engineering role first: you contribute production code, operate deployments, and you can explain a trade-off to customer leadership as clearly as to core engineering teams. This is a high-autonomy role with direct customer impact. This role is remote, based out of North America, with preference near one of our main hubs: Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You understand how accelerator compute, memory, and networking topology constrain AI workloads, and don't treat hardware as a black box. You're an early adopter of AI for your work from coding to building agentic workflows that multiply your impact. You work directly with customers to understand their challenges and provide effective solutions. You are comfortable debugging across the full inference stack: from failing requests, through the serving layer, down to OOMs or kernel dispatch if n

awskubernetesmachine learning
View job →
T
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. At Tenstorrent, we build open, state of the art compute for real workloads and real developers.You will own CPU core‑level verification, shaping how our out‑of‑order RISC‑V CPUs behave in silicon. This role is hybrid, based out of Bangalore, India. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You bring 8+ years in CPU verification or closely related digital design. You know high‑performance out‑of‑order CPU microarchitecture in depth. You work comfortably with RTL, waveforms, logs, and complex debug scenarios. You communicate clearly across design, DV, emulation, and post‑silicon teams. What We Need Plan and drive functional verification for CPU core features and complex microarchitectural scenarios. Develop UVM, assembly, and C/C++ based stimulus, functional models, and coverage for ISA, RISC-V extensions, and un-core components. Debug simulation and emulation regressions using RTL understanding, waveforms, and logs to identify and resolve issues efficiently. Build and enhance coverage models, testbenches, and debug infrastructure to improve verification quality and coverage closure. Collaborate with design, validation, and

awsgitai
View job →
T
Tenstorrent
📍 Toronto• Full-time• $100K – $500K/yr
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. At Tenstorrent, we are building open, scalable compute for real AI workloads. As Director, Customer Hardware Engineering, you own the technical relationship with strategic customers and FAEs, turning their silicon needs into precise requirements for our hardware and software teams. You connect customer architectures to Tenstorrent platforms so their models run efficiently on our silicon. This role is hybrid, based out of Toronto, Austin, TX or Belgrade, Serbia. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced leader of customer-facing technical teams and Field Application Engineering organizations. Strong background across RTL, verification, and physical design for custom silicon programs. Systems thinker who understands how RTL choices affect software, performance, and customer solutions. Clear communicator who aligns customers, FAEs, and internal teams around shared technical goals. What We Need Own technical customer relationships and convert high-level asks into concrete engineering specifications. Coordinate with hardware and software leads on customer-specific NEO silicon configurations. Oversee RTL changes and NEO cus

T
13 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a Field Application Engineer to serve as the technical bridge between Tenstorrent and customers across Southeast Asia. Based in Singapore, you will work closely with customers, partners, sales, and global engineering teams to understand AI and machine learning workloads, guide technical evaluations and deployments, troubleshoot issues across hardware and software, and help customers realize the performance of Tenstorrent’s AI platforms. This is a highly visible, customer-facing role that combines hands-on technical problem solving, solution development, and regional relationship building, with regular travel throughout Southeast Asia. This role is remote, based out of Singapore. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A customer-focused technical professional who can build trust with application developers, engineering teams, and business stakeholders. Comfortable translating complex AI/ML hardware and software concepts into clear recommendations for both technical and non-technical audiences. A proactive and self-directed problem solver who can coordinate internal teams, external service providers, and customer sta

pythonawsmachine learning
View job →
🔔

Get new workload porting and performance engineer jobs by email

Daily job updates · Unsubscribe anytime