Jobiba hiring network

Nodejs Jobs

492 active opportunities · Updated for October 2026

Fresh results

12 shown

Explore current nodejs jobs. Use filters to narrow by work mode, employment type, experience and date posted.

I
1mo ago

Job Details: Job Description: We are looking for a PERC Runset Development Engineer to develop and Validate PERC (Programmable Electrical Rule Checker) rule decks. The role involves working closely with PDK Team and foundry to implement PERC and ensure design compliance. Key Responsibilities: • Rule Development: Develop, validate, and maintain complex PERC rule decks to identify electrical reliability violations (e.g., ESD paths, current density, voltage stress, and floating gates). • Multi-Tool Integration: Develop and optimize verification flows across multiple EDA environments, ensuring consistency in results between Calibre PERC( preferred), Cadence Pegasus, and Synopsys tools. • Cross-Functional Collaboration: Work closely with Device Physicists, Process Integration Engineers, and Analog Designers to translate reliability requirements into programmable rules. • Debug and Analysis: Provide expert-level support to design teams in debugging PERC violations, distinguishing between true reliability risks and tool-induced false positives. • Automation: Develop scripts (Python, SKILL, Tcl, or Perl) to automate the execution, reporting, and tracking of reliability checks across different design versions. • Documentation: Maintain detailed documentation of reliability check-sets and create guidelines for designers to implement correct-by-construction layouts. Qualifications: Required Skills: • Experience with PERC rule development (preferably using Calibre PERC). • Knowledge of ESD, latch-up and reliability verification. • Familiarity with SVRF, TVF, Python or Tcl scripting. • Strong debugging and problem-solving skills. • Knowledge of Linux and scripting (Python/Shell/Tcl) is desirable. • Experience with advanced technology nodes is an advantage. • Team player and be able to work with globa

pythonlinuxai
View job →

NVIDIA is leading groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU -- our invention -- serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables groundbreaking creativity and discovery, and powers inventions that were once considered science fiction, including artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We build communication libraries like NCCL, NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center platforms and scalable communications software. DL and HPC applications have a huge compute demands and already run at scales of up to tens of thousands of GPUs. GPUs are connected with high-speed interconnects (e.g. NVLink, PCIe) within a node and with high-speed networking (e.g. InfiniBand, Ethernet) across nodes. Efficient and fast communication between GPUs directly impacts end-to-end application performance. This impact continues to grow with the increasing scale of next generation systems. This is an outstanding opportunity to advance the state-of-the-art, break performance barriers, and deliver platforms the world has never seen before. Are you ready to build the new and innovative technologies that will help realize NVIDIA's vision? What you will be doing: Investigate opportunities to improve communication performance by identifying bottlenecks in today's systems. Design and implement new communication technologies to accelerate AI and HPC workloads. Explore innovative solutions in HW and SW for our next generation platforms as part of co-design efforts involving GPU, Networking, and SW architects. Build proofs-of-concept, conduct experiments,

linuxartificial intelligenceai
View job →
N
1mo ago

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! This is an outstanding opportunity for someone with HPC and performance background to advance the state of the art in this space. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters. Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available Triage and root-cause performance issues reported by our customers Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information <li

pythondockerkubernetes
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for a customer-obsessed software engineer to come ship with us. You’ll own features like multi-node training and products like serverless reinforcement learning (RL) from conception to MVP (and from MVP to GA!). You’ll work through the stack, architecting solutions from API and UI down to our infrastructure layer. You’ll fine tune models yourself to develop an understanding of user workflows. You’ll work closely with research engineers leveraging state-of-the-art training techniques to build experiences that accelerate model development and solve for real pain points. If you’re excited to dive deep into the training, let’s talk! THE PRODUCT Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done EXAMPLE INITIATIVES Checkpointing Pipeline: Our checkpointing pipeline starts with automated checkpointing, a feature that ensures that versions of models created during training are automatically backed up to the cloud. Users are able to then deploy checkpoints seamlessly into inference servers, providing point-and-click integrations into inference frameworks like vLLM and Baseten’s Inference Stack. This enables customers to quickly evaluate the performance of their checkpoints with real traffic. Multinode training: Multinode training enables customers to easily run training jobs across multiple compute nodes, enablin

kubernetesrestmachine learning
View job →
G
Godaddy
📍 United States• Full-time• From $154K/yr
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the world, delivering the object, block, and file storage platforms that power GoDaddy's hosting infrastructure, internal services, OpenStack environments, and next-generation AI/HPC workloads. If you're passionate about distributed systems, storage architecture, and solving failure scenarios at massive scale, this is an opportunity to work on infrastructure few engineers will experience in their careers. Ceph is a strategic platform at GoDaddy — not an ancillary service. Our global footprint includes 80+ production clusters, 20,000+ OSDs, 1,830 storage nodes, 300 PB of raw capacity, and 69 billion objects spanning five datacenters across three continents. The platform supports RBD, RGW (S3/Swift), and CephFS workloads through more than 1,550 pools, 574,000 placement groups, and 900+ MDS daemons, creating engineering challenges that demand deep expertise in storage architecture, data durability, performance optimization, automation, and observability. As a Lead Senior Site Reliability Engineer, you'll serve as one of the principal technical leaders for GoDaddy's Ceph platform. You'll design the next generation of storage clusters, lead major platform upgrades, drive capacity and hardware strategy, and establish the standards that govern how the platform scales. You'll be the engineer the team turns to for the most complex s

pythonkubernetesai
View job →
N
Nuro
📍 Mountain View• Full-time• From $160.4K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro is seeking a Software Engineer with expertise in large-scale infrastructure, workload orchestration, and data processing to join our ML Infrastructure team . In this role, you will focus on building and evolving the core platform that provides researchers and engineers with seamless access to compute and data resources. You will be responsible for executing the technical strategy for automated resource provisioning, high-performance workload scheduling, and efficient feature management to accelerate the Nuro Driver™ development lifecycle. About the Work You will build the foundation that powers Nuro’s model development from experimentation to production. Key responsibilities include: Resource Provisioning & IaC: Scaling automated infrastructure-as-code (IaC) pipelines to manage thousands of GPU/CPU nodes across diverse environments. Intelligent Scheduling: Designing and optimizing workload orchestration to maximize

redisawsazure
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro is seeking a Software Engineer with expertise in large-scale infrastructure, workload orchestration, and data processing to join our ML Infrastructure team . In this role, you will focus on building and evolving the core platform that provides researchers and engineers with seamless access to compute and data resources. You will be responsible for executing the technical strategy for automated resource provisioning, high-performance workload scheduling, and efficient feature management to accelerate the Nuro Driver™ development lifecycle. About the Work You will build the foundation that powers Nuro’s model development from experimentation to production. Key responsibilities include: Resource Provisioning & IaC: Scaling automated infrastructure-as-code (IaC) pipelines to manage thousands of GPU/CPU nodes across diverse environments. Intelligent Scheduling: Designing and optimizing workload orchestration to maximize

redisawsazure
View job →
R
Roblox
📍 San Mateo• Full-time• From $345K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. We are seeking a visionary Principal Software Engineer to join our Compute organization and provide technical leadership for our Kubernetes infrastructure. You will drive the evolution of a platform that powers our global scale operations, transforming Kubernetes into a secure, reliable, and invisible foundation for our developers across our on-prem and public cloud fleet. Your mission is to balance cutting-edge innovation with rigorous platform stability, ensuring that internal and external customers have a seamless, high-performance experience at massive scale. You Will: Architectural Leadership: Serve as a technical lead for our Kubernetes ecosystem, setting the long-term architectural strategy for a platform that manages thousands of nodes and supports millions of concurrent requests. Customer Focus: Think deeply about how our internal customers consume and interact with compute, designing intuitive interfaces and tooling that simplify consumption of complex infrastructure services. Deep-Dive Engineering: Leverage deep expertise in Kubernetes internals, including custom controllers, operators, API server architecture, and etcd, to solve complex scaling bottlenecks and optimize our contr

awskubernetesgit
View job →
R
Roblox
📍 San Mateo• Full-time• From $243.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will play a key role in enabling our mission to deliver a reliable, high-performing, and cost-efficient infrastructure that powers the world’s play. In this specialized role, you will be the technical lead for our GPU and AI accelerator ecosystem. You will be responsible for the full lifecycle of GPU hardware, from initial architectural evaluation and firmware qualification to large-scale fleet integration and performance tuning. You will ensure that Roblox’s massive-scale rendering and ML workloads run on the most optimized and stable hardware possible. You Will: Architect & Prototype: Prototype next-generation GPU-accelerated hardware platforms, ensuring seamless integration between high-density compute nodes, high-speed interconnects (NVLink/PCIe Gen5/6), and system firmware. GPU Optimization: Drive the integration, performance testing, and debugging of GPUs in our fleet, focusing specifically on hardware-level optimizations, driver tuning, and thermal/power management. Validation & Certification: Develop and execute rigorous evaluation and stress-testing strategies for GPU-heavy server platforms to ensur

pythonawsgit
View job →
O
1mo ago

About the Team The Platform Systems team at OpenAI operates at the intersection of cutting-edge AI and large-scale distributed systems. We build the engineering and research infrastructure required to train OpenAI’s flagship models on some of the world’s largest, custom-built supercomputers. Our team develops core model training software and works deep in the stack - spanning collective communication, compute efficiency, parallelism strategies, fault tolerance, failure detection, and observability. The systems we build are foundational to OpenAI’s research velocity, enabling reliable, efficient training at frontier scale. We collaborate closely with researchers across the organization, continuously incorporating learnings from across OpenAI into the evolution of our training platform. About the Role As a Software Engineer, Platform Systems, you will design and build distributed systems that provide visibility into large-scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical to operating OpenAI’s training stack and is actively evolving to support new use cases and increasingly complex workloads. This role sits at the core of our training infrastructure, blending systems engineering, performance analysis, and large-scale debugging. In This Role, You Will Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high-throughput distributed systems Collaborate with systems, infrastructure, and research teams to evolve platform

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Platform Systems team at OpenAI operates at the intersection of cutting-edge AI and large-scale distributed systems. We build the engineering and research infrastructure required to train OpenAI’s flagship models on some of the world’s largest, custom-built supercomputers. Our team develops core model training software and works deep in the stack - spanning collective communication, compute efficiency, parallelism strategies, fault tolerance, failure detection, and observability. The systems we build are foundational to OpenAI’s research velocity, enabling reliable, efficient training at frontier scale. We collaborate closely with researchers across the organization, continuously incorporating learnings from across OpenAI into the evolution of our training platform. About the Role As a Software Engineer, Platform Systems, you will design and build distributed systems that provide visibility into large-scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical to operating OpenAI’s training stack and is actively evolving to support new use cases and increasingly complex workloads. This role sits at the core of our training infrastructure, blending systems engineering, performance analysis, and large-scale debugging. In This Role, You Will Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high-throughput distributed systems Collaborate with systems, infrastructure, and research teams to evolve platform

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical. The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly. About the Role We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity. In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity. This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions. This role is based in San Francisco, CA, with travel as needed. Key Responsibilities Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint Drive readiness to convert contracted compute capacity into schedulable production clusters Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives Build integrated schedules spanning procurement, logistics, installation, st

awsazurerest
View job →
🔔

Get new nodejs jobs by email

Daily job updates · Unsubscribe anytime