Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Firmware Test Engineer at Micron Technology, Inc., you will build groundbreaking high-performance controller firmware test framework for volatile and non-volatile memory systems. You will assist in the evaluation, creation, build, bench testing, debugging, and failure analyzes of firmware for new high-performance memory controllers and Solid State Drives (SSD) that will improve performance, while reducing power, latency and SoC (System on Chip) complexity for the target sectors. You can expect to partner multi-disciplinary Engineers seek multi-functional product development issues. You will triage failures, file bug reports, and help the development teams with isolating issues. Experience / Skills: 8 to 12 years of experience in managing the test development team within the storage domain. In depth knowledge and extensive experience with embedded firmware development Expertise in the use of scripting languages, programming tools and environments Extensive experience programming in Python Technical Expertise in the storage industry in SSD, HDD, storage systems, or a related technology Understanding of storage interfaces including ideally PCIe/NVMe, SATA, or SAS Experience with NAND flash and other non-volatile storage Ability to work independently with a minimum of day-to-day supervision Experience with team leadership and/or supervising junior engineers and technicians Ability to work in a multi-functional team and under
Jobiba hiring network
System Power Engineer Jobs
10,000 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current system power engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery, integrating frontier AI models, biological foundation models, and lab capabilities, with the ultimate goal of curing disease. Our technology powers scientists around the world, translating AI capabilities into tools that accelerate research everywhere. The Team The AI Cluster Production Engineering team is part of the AI Compute Platform organization at Biohub, a non-profit research lab committed to open science and open-source AI. We own the design, operation, and reliability of large-scale multi-GPU AI clusters that power frontier AI biology research: protein language models, genomic foundation models, and scientific reasoning systems built to be shared, not monetized. Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the intersection of AI tooling, distributed systems, HPC, and frontier AI, debugging deep AI infrastructure problems and building AI systems critical to the entire AI organization. The Opportunity CZ Biohub's mission is to cure or prevent all human disease. Achieving that requires training frontier-scale AI biology models, and that demands reliable, high-performance compute infrastructure. This is production engineering work at a frontier AI lab, with the twist that the mission is biology and the science is open. You'll keep GPU clusters running at high utilization, debug the toughest distributed systems failures, and build the operational foundations for scaling to multi-thousand GPU hero runs. The technical problems are genuinely hard (e.g., multi-node distributed training, InfiniBand fabrics, large-scale storage, Slurm at scale) inside an organization where the work is aimed at helping peop
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Lyft Business Product Platform team builds the systems and experiences that power Lyft's B2B products — enabling companies, organizations, and their employees to seamlessly access Lyft's transportation network. We sit at the intersection of product and platform, owning both the customer-facing features and the underlying infrastructure that makes them reliable at scale. Our work directly impacts how businesses integrate with Lyft, how admins manage their programs, and how millions of riders get where they need to go. Responsibilities: Drive architecture and technical design for systems that are highly available, scalable, and built to last — not just for today's requirements but for where the product is heading Own features end-to-end: from shaping the technical spec and design through to production rollout and operational health Think critically about how AI capabilities can be incorporated into Lyft Business products to improve the experience for business admins and riders — and bring that perspective into roadmap and architecture conversations Make well-reasoned trade-off decisions and communicate them clearly to peers, leads, and cross-functional partners Write clean, well-tested, maintainable code and hold a high bar for the same in code reviews Partner across engineering, product, and design to align on direction and get buy-in on technical approaches Proactively engage in incident response, contributing both to resolution and to long-term reliability improvements Grow the team's technical culture through design reviews, tech talks, and mentorship Experience: 5+ years of software engineering experience, with a track record of designing and shipping production systems at scale Strong system design instincts — you can reason through distributed systems trade-offs, identify failure modes, and
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Merchant Experience team builds the UI operating system and core experiences that power decades of growth for millions of businesses globally, enabling Stripe to be the business SaaS solution of choice for next-gen companies. We own the foundational platforms and product surfaces that merchants interact with every day—from the Stripe Dashboard to the frameworks and tooling that hundreds of Stripe engineers use to build world-class user interfaces. We're at an inflection point. AI is fundamentally changing how users interact with software, and we believe Stripe is uniquely positioned to reimagine what a merchant experience looks like in an AI-native world. At the same time, we're raising the bar on the quality and performance of everything we ship—because the millions of businesses that depend on Stripe deserve an experience that is fast, reliable, and crafted with care. What you'll do As a Senior Staff Engineer on the Merchant Experience team, you'll be a technical leader responsible for shaping the next generation of how Stripe merchants interact with our products. You'll drive a step-function improvement in the quality and performance of our core surfaces, and help define what AI-powered merchant experiences look like at Stripe. Responsibilities Define technical strategy for merchant-facing experiences across Stripe, with a focus on quality, performance, and the integration of AI into core user workflows Lead the design and delivery of
About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe Terminal helps Stripe users extend their online presence into the physical world. The Terminal team’s mission is to make it as easy for businesses to accept in-person payments as the Stripe API has done for online payments. With Terminal, businesses can unlock in-person payments use cases that are right for their business model—whether it’s creating a flagship retail experience, extending their website to a pop-up store, or enabling a mobile point-of-sale at their next event. What we are looking for: As an Android BSP Engineer, you will be responsible for the kernel and driver level system development, which includes building, troubleshooting, and writing automated tests for the Android system on our embedded payments platforms. This team works closely with partner teams throughout the hardware and software product lifecycle, from hardware manufacturing to Android app teams. We also work with external vendors on part selection and initial hardware bring-up. What you’ll do: Bring up new devices and lead debugging and performance tuning exercises that span multiple hardware/firmware/software teams. Design, implement, and maintain drivers and Android services that operate efficiently in a constrained environment and meet the reliability and security requirements of the industry. Own the definition of one or more work streams focused on hardware bring-up, peripheral drivers and communication, and power and performance management and opti
About the Team Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running. That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets. About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models
About the Team The Release Engineer team is responsible for building and maintaining the systems that power software delivery—from CI/CD pipelines and artifact management to release automation and fleet telemetry. We ensure software across bootloaders, firmware, operating systems, and cloud services is built reproducibly, validated rigorously, and released safely at scale. About the Role As a Release Engineer, you’ll design, build, and operate release infrastructure that enables reliable, secure, and traceable software delivery across complex multi-component systems. You’ll partner closely with embedded, cloud, and QA teams to ensure that every build—from development to OTA deployment—is fast, verifiable, and production-ready. We’re looking for engineers who take pride in automation, build reproducibility, and system reliability—and who enjoy building the connective tissue that allows hardware and software to ship together seamlessly. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and operate CI/CD pipelines for multi-component builds (bootloader, firmware, OS images, backend, companion apps) using hermetic toolchains. Define versioning and branching strategies; automate promotions, changelogs, and artifact retention. Integrate unit, integration, and hardware-in-the-loop (HIL) test results; quarantine flaky tests, auto-bisect failures, and block unsafe promotions. Build A/B OTA update flows with verity and health checks; run staged rollouts and canaries; implement safe rollback and roll-forward strategies. Implement code signing for binaries and firmware, generate SBOMs, run vulnerability scanning, and attach build attestations and provenance. Manage dashboards and alerts for build health, promotion latency, failure rates, and fleet update telemetry. You might thrive in this role if you: Have experience building and operating buil
Job Title: Senior QA Engineer - Performance Testing Paytm is India's leading mobile payments and financial services distribution company. A pioneer of the mobile QR payments revolution in India, Paytm builds technologies that empower small businesses with payments and commerce solutions. Paytm’s mission is to serve half a billion Indians and bring them into the mainstream economy through the power of technology. About the Role: We are seeking a skilled Performance Test Engineer to design, execute, and analyze performance tests to ensure application scalability, stability, and responsiveness under varying load conditions. The ideal candidate will have hands-on experience with industry-standard performance testing tools and a strong understanding of system architecture, monitoring, and troubleshooting. Expectations/ Requirements Develop comprehensive performance test strategies and plans aligned with system requirements, project timelines, and business goals. Understand application architecture and identify critical business transactions for performance validation. Design realistic workload models to simulate real-world usage scenarios. Create, maintain, and execute performance test scripts using tools such as JMeter, LoadRunner, Gatling, or similar. Conduct baseline, load, stress, and scalability testing to evaluate system behavior under different conditions. Monitor system performance using tools like Influx DB, Grafana, JVM monitoring tools, and MAT (Memory Analyzer Tool). Analyze test results to identify performance bottlenecks and system limitations. Collaborate with development and infrastructure teams to troubleshoot and resolve performance issues. Assess system scalability and recommend optimizations to improve performance and reliability. Generate detailed performance test reports, including metrics, findings, and actionable recommendations. Work with stakeholders to gather and validate Non-Functional Requirements (NFRs), SLAs, and KPIs. Perform API an
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. About Our Team Our team builds and enables scalable, cloud-native software platforms that power critical manufacturing and business operations. We use modern full stack engineering practices and AI-driven technologies to create innovative solutions that improve automation, decision making, and operational efficiency. Position Overview We are seeking an AI-Centric Full Stack Software Engineer to design, build, and support modern software solutions with a strong focus on AI-enabled applications and intelligent automation. This role goes beyond simply using AI coding assistants. You will have strong understanding of Prompt Engineering, Vibe Coding, Rework Rate Reduction and leverage Custom Agents, and integrations built around the Model Context Protocol (MCP). You will apply strong software engineering fundamentals to assemble, integrate, and operationalize AI capabilities into real production systems being accountable for Full-Stack solutions. The goal of this role is to significantly shorten development and feedback cycles by automating routine engineering work, augmenting human decision-making, and embedding intelligence directly into our development workflows and the applications we deliver! Responsibilities Design, develop, test, deploy, and maintain scalable full stack applications and platform services. Own software solutions end-to-end, including system design, architecture, implementation, testing, observability, deployment, and lifecycle support. Develop an
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. About Our Team Our team builds and enables scalable, cloud-native software platforms that power critical manufacturing and business operations. We use modern full stack engineering practices and AI-driven technologies to create innovative solutions that improve automation, decision making, and operational efficiency. Position Overview We are seeking an AI-Centric Full Stack Software Engineer to design, build, and support modern software solutions with a strong focus on AI-enabled applications and intelligent automation. This role goes beyond simply using AI coding assistants. You will have strong understanding of Prompt Engineering, Vibe Coding, Rework Rate Reduction and leverage Custom Agents, and integrations built around the Model Context Protocol (MCP). You will apply strong software engineering fundamentals to assemble, integrate, and operationalize AI capabilities into real production systems being accountable for Full-Stack solutions. The goal of this role is to significantly shorten development and feedback cycles by automating routine engineering work, augmenting human decision-making, and embedding intelligence directly into our development workflows and the applications we deliver! Responsibilities Design, develop, test, deploy, and maintain scalable full stack applications and platform services. Own software solutions end-to-end, including system design, architecture, implementation, testing, observability, deployment, and lifecycle support. Deve
About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors, software, and data center systems that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our U.S. engineering teams contribute to the hardware and software platforms that support the next generation of AI systems. The Opportunity We are looking for a computer engineering, electrical engineering, or computer science student or recent graduate to join the BMC Development team as a Firmware Engineering Intern. You will work with experienced engineers on low-level and embedded firmware that supports the operation, control, and manageability of advanced compute systems. This internship provides hands-on experience in firmware development, test automation, engineering experiments, and lab-based system testing in a Linux development environment. You will own clearly defined technical tasks with guidance from the team and contribute to production-quality engineering work. Type: 12-week summer internship Timing: May - August (exact dates to be confirmed) Commitment: Full-time What You Will Do Contribute to the design, implementation, and testing of system and embedded firmware. Develop and maintain firmware and supporting software in C, C++, or Python. Support firmware development and debugging in a Linux-based engineering environment. Create automated tests and scripts that improve firmware validation, test coverage, and engineering efficiency. Contribute to continuous integration and delivery workflows for firmware development and testing. Plan and conduct well-defined engineering experiments, record results accurately, and draw conclusions from test data. Support lab setup, system configuration, hardware bring-up, and firmware testing. Use debugging and diagnostic techn
Who We Are Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. Addepar unifies portfolio, market and client data in a total portfolio view and delivers AI-powered insights within investment and client workflows. More than 1,400 firms in nearly 60 countries use Addepar to manage and advise on nearly $9 trillion in assets. Its open platform integrates with nearly 650 software, data and consulting partners to power end-to-end investment operations across firms of all sizes and complexity. Addepar supports clients worldwide with offices in New York City, Salt Lake City, London, Edinburgh, Pune, Dubai, Geneva, Singapore and São Paulo. The Role We are seeking a Staff Full Stack Software Engineer to join the Advisor Experience team as our Technical Lead. Our team is focused on building tools for financial advisors to grow and sustain their business. We oversee bespoke products for advisors and develop advisor-focused capabilities throughout the Addepar platform. In this role, you will be the primary technical anchor for new capabilities including Secure Message Center — a compliant messaging experience built into Addepar's client portal that allows advisors and their clients to communicate directly within the platform. You will partner directly with Engineering Leadership and Product Management to build a modern, scalable architecture from the ground up. Beyond system design, you will act as a true engineering multiplier: setting technical standards, mentoring junior and mid-level engineers, and working alongside other senior engineers and AI specialists to deliver high-impact advisor tools. Applicants must be legally authorized to work in the United States for any employer without requiring current or future visa sponsorship (for example, employment-based visas such as H-1B, F-1/OPT, or similar), and must be authorized to begin work in the U.S. on their first day of employme
DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You'll Do: Deploy, configure, and maintain Kubernetes clusters for our microservices architecture. Utilize Git and Helm for version control and deployment management. Implement and manage monitoring solutions using Prometheus and Grafana. Work on continuous integration and continuous deployment (CI/CD) pipelines. Containerize applications using Docker and manage orchestration. Manage and optimize AWS services, including but not limited to EC2, S3, RDS, and AWS CDN. Maintain and optimize MySQL databases, Airflow, and Redis instances. Write automation scripts in Bash or Python for system administration tasks. Perform Linux administration tasks and troubleshoot system issues. Utilize Ansible and Terraform for configuration management and infrastructure as code. Demonstrate knowledge of networking and load-balancing principles. Collaborate with development teams to ensure applications meet reliability and performance standards. Who you are: Bachelor’s degree in engineering (CS / IT) or equivalent degree from a well-known Institute / University. 2+ years of experience in a Site Reliability Engineer role or similar. Proven experience with Kubernetes, Git, Helm, Prometheus, Grafana, CI/CD, Docker, and microservices architecture. Strong knowledge of AWS services, MySQL, Airflow, Redis, AWS CDN. Proficient in scripting languages such as Bash or Python. Hands-on experience with Linux administration. Familiarity with Ansible and Terraform fo
Get new system power engineer jobs by email
Daily job updates · Unsubscribe anytime