Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Engine Networking Team pulls the players together by ensuring the communication of the game state to all. As a Principal Engineer on this team you will help the players experience the game as a nearly synchronous world. The networking and asset loading team plays a key role in a smooth experience for the players. You will work in all areas of the game platform in your quest for real-time communication of every part of Roblox. You Will: Lead engineers with 8+ years of industry experience Be experienced with one of these area: asset loading, rendering, and networking coming from a Game Engine/Studio. Be an amazing systems-level C++ programmer and be fascinated by the actual work the CPU does when you use smart pointers, templates, virtual functions, and blocks of memory, both structured and raw Have a keen to each millisecond of the network exchanges: You know where the time goes and how to reduce the waste Understand what happens on the operating system level when certain code is completed You Have: Worked on the guts of a multi-player game engine, solving problems related to scale, performance, latency, and throughput in client/server environments. Worked on a very large multithreaded d
Jobiba hiring network
Cpu Pre Si Verification Engineer Jobs
264 active opportunities · Updated for October 2026
Fresh results
9 shown
Explore current cpu pre si verification engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's Cache team is building a next-generation caching solution designed to deliver sub-millisecond average latency, horizontal scalability, and high efficiency—all at a drastically lower cost. Our ultimate vision is to shape a caching infrastructure capable of supporting 1 billion Daily Active Users while reducing costs by 90%. We are turning hours of onboarding and capacity expansion into seconds, freeing service owners entirely from managing cluster lifecycles. As a Senior Engineer on the Cache team (part of the Infra Storage org), you will innovate and operate large-scale, in-house distributed systems to solve Roblox's ever-growing caching challenges. You will report directly to the Engineering Manager for the Cache team. (Check out our recent engineering blog post here to learn more about the team's latest work!) You will: Lead the architectural transition to a next-generation, multitenant caching service built on ValKey, ensuring strict data, resource, and failure isolation for all tenants. Drive systemic optimizations to mitigate head-of-line blocking, manage hot keys, and maximize CPU and memory utilization across physical machine clusters. Design and build robust frameworks to a
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. About the Role: AI models reshaping how our community creates, plays, and connects, all run on Compute Platform. As Senior Product Manager, Compute Platform , you'll set the strategy and roadmap for Roblox's next-generation AI infrastructure: the rapidly growing fleet of GPUs and AI accelerators spanning Roblox core and edge data centers, and public cloud that decides how fast we can train, serve, and scale every model on the platform. You'll own the products that turn raw GPU hosts into reliable, production-ready AI compute - driver and firmware management, fleet-wide health and performance, and the abstractions product teams across Roblox build on. You Will: Drive strategy and roadmap for Compute Platform spanning Managed Kubernetes (Roblox Kubernetes Service), Managed Compute Services and other critical distributed systems, and our fleet of GPU and CPU machines managed via unified Fleet APIs - all across on-prem and cloud. Drive the evolution of our Compute infrastructure to support Roblox’s most critical workloads - from AI to Storage to Data Analytics and more - each with their own distinct requirements. Build and scale our GPU infrastructure to support training and inferen
About the Team The Workload Networking team is responsible for the collective communication stack used in our largest training jobs. Using a combination of C++ and CUDA we work on novel collective communication techniques that enable efficient training of our flagship models on our largest custom built supercomputers. The models we train are key ingredients to the AI research progress at OpenAI and the field as a whole, and we continually incorporate learnings from our entire research org into our training platform. About the Role As a Software Engineer, Networking you will design and implement custom networking collectives that are tightly integrated into our training stack. We’re looking for people who have a background in low level performance critical software. Experience with collective communication is a bonus. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Collaborate closely with ML researchers to design and implement efficient collective operations in C++ and CUDA. Ensure that our largest training jobs take full advantage of the different network transports used in our supercomputers. Work on simulations to inform our future supercomputer network designs. You might thrive in this role if you: Have written distributed algorithms using RDMA in the past. Are comfortable writing low level performance sensitive CPU and/or GPU code. Are familiar with network simulation techniques. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voic
Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role As OpenAI continues to grow, we are looking for experienced, problem-solving engineers to ensure our systems scale. Our success depends on our ability to quickly iterate on products while also ensuring that they are performant and reliable. You will work in a deeply iterative, collaborative, fast-paced environment to bring our technology to millions of users around the world, and ensure it’s delivered with safety and reliability in mind. Successful candidates will play a crucial role in ensuring the reliability, scalability, and performance of our systems as we continue to expand. As a reliability expert, you will be at the forefront of maintaining and enhancing the stability, scalability, and performance of our rapidly evolving infrastructure. You will work closely with cross-functional teams, including software engineers, product managers, and data scientists, to build and maintain resilient systems that can handle our growing user base and workload. In this role, you will: Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands. Build and maintain the load, chaos and synthetic testing software leveraged by development teams to make the systems they design and operate more reliable. Build and maintain automation tools to streamline repetitive tasks and improve system reliability. Build and maintain the platform for CPU/storage, GPU, and network lifecycle management to drive efficiency, accountability and support dynamic optimization of our resources. Implement fault-tolerant and resilient
Location: San Francisco, CA (Hybrid: 4 days onsite/week). Relocation assistance available. About the Team: We build foundational platform software that enables reliable, secure, and performant products. The team works across system layers and partners closely with adjacent engineering groups to deliver robust capabilities from concept through launch. About the Role: We’re seeking a System Software Engineer to design, implement, and debug core platform components and the pipelines that build and update system images. You’ll work across operating system layers, focusing on performance, security, and deep system debugging to ship production‑grade systems. In this role, you will: Design, implement, and debug system‑level components and services across kernel and user space. Configure and maintain OS platform services (init, services, networking, security policies) and related tooling. Build and operate image and update pipelines, ensuring reliability, reproducibility, and rollback safety. Instrument and analyze performance using profiling and tracing; optimize CPU, memory, I/O, and power usage. Own platform observability and reliability: logging, crash capture, watchdogs, and diagnostics. Collaborate with cross‑functional teams to define interfaces and deliver end‑to‑end features. Establish strong engineering practices: code review, CI, reproducible builds, and release management. Partner with external suppliers to support builds and deployments. You might thrive in this role if you: Have shipped production systems software on modern operating systems. Are proficient in C/C++ and a scripting language, and comfortable with OS internals (concurrency, memory management, filesystems, networking, power management). Bring strong systems debugging skills using debuggers, tracers, profilers, and logs across kernel/user‑space boundaries. Understand configuration of platform services and interfaces, and can translate requirements into stable, well‑documented APIs. Are fluent in u
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
About the Role As a Director, Compute & Infrastructure FP&A, you will own and drive the monthly forecasting process for the Compute & Infrastructure org by partnering with various stakeholders across Finance, Accounting, Tax and Engineering. You will play a critical role in planning and forecasting the company’s largest and most complex cost center ( Compute & Infrastructure ). You will collaborate cross-functionally to develop long-range infrastructure investment plans, evaluate build vs. buy decisions, and ensure capital is deployed efficiently to support rapid growth. You will also provide strategic financial guidance through scenario modeling, ROI analysis, and performance tracking, enabling leadership to make high-stakes decisions under uncertainty. What You’ll Do Own compute financial planning & Forecasting. Build and manage consolidation models for GPU/CPU capacity, storage, networking, and data center investments. Translate infrastructure roadmaps into short- and long-term financial forecasts (LRP, annual planning) Coordinate closely with Corporate FP&A on timelines and process Present insights on a monthly basis to senior management. Drive infrastructure investment decisions. Evaluate build vs. buy, vendor vs. owned infrastructure, and capacity allocation tradeoffs. Develop frameworks for investment trade-offs to guide executive decision making. Build scalable tooling & reporting. Implement stakeholder-facing dashboards to track compute spend, utilization, and efficiency metrics. Improve visibility into unit economics (e.g., cost per training run, cost per inference, cost per customer). Drive forecasting accuracy & accountability. Lead budget vs. actual analysis for compute and infrastructure spend. Identify key cost drivers (utilization, pricing, efficiency gains) and reduce forecast variance. Support close & financial reporting. Partner with Accounting to ensure accurate classification of infrastructure spend (OpEx vs C
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake’s cloud spend is in billions of dollars per year. Hence, it is critical for us to govern and optimize our cloud spend, both for margins and long-term competitive advantage. Cloud Efficiency team’s charter is to build scalable products that enable governance, monitoring and optimization of cloud spend. Think of this as Observability for cloud costs and efficiency. The team’s vision is to “Transform cloud spend into a competitive advantage by empowering teams to continuously optimize the per-unit cost.” In order to improve the overall cloud efficiency (i.e. cost per unit), it is critical to build monitoring products that collate costs with other factors such as utilization, attribution, hardware performance and architecture. Hence, there is an opportunity to build a unified, self-serve cloud efficiency product across Snowflake, that delivers actionable, real-time efficiency datasets through streamlined user experiences. This will enable thousands of engineers at Snowflake and will elevate cloud efficiency at Snowflake for long-term success. When developing these solutions, we think about the problem end-to-end: how do we collect data from different stacks (e.g. costs from AWS, GCP, Azure and CPU, Memory, Utilization) across Snowflake reliably, how do we store it eff
Get new cpu pre si verification engineer jobs by email
Daily job updates · Unsubscribe anytime