Jobs in United States

Software Reliability Engineer in United States

2,007 active opportunities · Updated October 2026

Explore current software reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the hardware platforms that support the next generation of AI systems. The Opportunity We are looking for a recent graduate or early-career engineer to join the Mech/Thermal Platform team as a Graduate Mechanical Engineer. You will contribute to the mechanical design, thermal validation, and verification of advanced AI hardware and data center systems. You will work with experienced mechanical, thermal, hardware, and systems engineers throughout the development lifecycle. The role combines 3D computer-aided design, prototype evaluation, lab testing, data analysis, troubleshooting, and clear engineering documentation. What You Will Do Create and update mechanical parts, assemblies, and drawings using Creo or comparable 3D CAD software. Take ownership of defined mechanical design tasks from requirements and concepts through detailed design, review, release, and verification. Apply mechanical design, materials, manufacturing, heat transfer, and thermodynamics principles to engineering decisions. Develop mechanical and thermal test plans and procedures for prototypes and development systems. Set up and operate laboratory equipment, collect accurate data, analyze results, and document conclusions. Evaluate prototypes and identify mechanical, thermal, assembly, or manufacturability issues. Support design verification, troubleshooting, root-cause analysis, and implementation of verified design improvements. Maintain accurate engineering documentation, including design notes, drawings, test procedures, results, and change r

Artificial IntelligenceAI
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world’s most transformative technologies. Our new AI Engineering Campus in Austin will play a central role in building the future of AI computing. The Opportunity As a Mechanical Engineering Intern, you will contribute to the mechanical design and thermal validation of advanced AI hardware. Working alongside experienced mechanical and thermal engineers, you will gain hands-on experience with 3D computer-aided design (CAD), laboratory setup, test development, prototype evaluation, and engineering documentation. Type: 12-week summer internship Timing: May - August (exact dates to be confirmed) Commitment: Full-time What You’ll Do Contribute to mechanical design activities using Creo or comparable 3D CAD software, including part modeling, assemblies, drawings, and design updates under the guidance of experienced engineers. Work with the thermal engineering team to develop test procedures and help set up laboratory capabilities for evaluating AI hardware. Support thermal and mechanical testing using appropriate laboratory equipment; collect, organize, and analyze test data. Assist the mechanical and thermal design teams with prototype evaluation, troubleshooting, design verification, and documentation of findings. Collaborate with cross-functional engineering partners, communicate progress and issues clearly, and follow applicable laboratory and safety procedures. What You’ll Bring Current enrollment in a bachelor's, master's, or doctoral program in mechanical engineering or a related discipline during the internship. Students in other disciplines with relevant mechanical engineering exper

Artificial IntelligenceAI
I
📍 Texas, Austin, United States
✓ High-confidence listingCompany trend +315.4%
Quick readStrong listing-quality and freshness signals

Job Details: Job Description: Intel is shaping the future of technology to help create a better future for the entire world. Our work in pushing forward fields like AI, analytics, and cloud-to-edge technology is at the heart of countless innovations. With a career at Intel, you'll have the opportunity to use technology to power major breakthroughs and create enhancements that improve our everyday quality of life. Join us and help make the future more wonderful for everyone. Want to learn more? Visit our YouTube Channel or the link below. Life at Intel The Role and Impact As a SoC Power Thermal Performance Validation and Optimization Engineer, you will play an essential role in ensuring Intel's products achieve optimal power, thermal, and performance benchmarks at the system-on-chip (SoC) level. In this position, you will develop and execute validation methodologies, optimize hardware/software solutions, and perform SoC-level debugging to address power, thermal, and performance challenges. Your contribution will directly impact Intel's ability to deliver competitive, high-performing products to market. Business Group This role is part of Intel's Data Center Group (DCG), which focuses on designing innovative solutions to power the next generation of computing platforms for data centers. As a member of the PTP PreSi Correlation Solution team, you will contribute to shaping the development and validation processes for cutting-edge technologies, ensuring Intel meets and exceeds industry standards for performance and efficiency. Key Responsibilities: Develop and execute SoC-level power, thermal, and performance validation and

PythonAIRecruitment
C
📍 United States· Remote
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking an accomplished Principal Cloud Storage Engineer to lead the design, engineering, and evolution of our private cloud storage platforms. This role will focus on large-scale storage architecture, data protection, cyber recovery, and resiliency technologies across complex enterprise environments. The ideal candidate will combine deep technical expertise in storage systems with strong leadership, architectural vision, and the ability to influence technical direction across the organization. Key Responsibilities Architect and engineer enterprise storage platforms that ensure data integrity, availability, security, and disaster recovery readiness Design and implement end-to-end storage solutions, including Software Defined Storage, SAN, NAS, and object storage across private cloud and data center environments Drive strategic technology decisions by evaluating emerging products, tools, and standards supporting storage, data protection, cloud, and compute platforms Lead infrastructure initiatives involving storage modernization, data protection, cyber recovery, data migration, and resilience engineering Develop and execute enterprise strategies for backup, recovery, cyber vaulting, and business continuity Create and maintain comprehensive documentation of storage architectures, configurations, policies, and operation

KubernetesProject Management
MT
📍 Boise, ID - ID1, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Job Description At Micron-Boise, ID, we are undergoing a historic $15 billion investment in semiconductor manufacturing; construction began in early 2023, with DRAM production slated for the second half of the decade. As a leader in the semiconductor industry, we build solutions that encourage and transform technology. With plans to invest more than $150 billion globally over the next decade in groundbreaking manufacturing, we are looking for hardworking people to join our Boise expansion team and give to the growth and innovation of the semiconductor industry. ID1 High Volume Manufacturing Equipment Engineer is responsible for monitoring, sustaining, and improving the equipment while working in partnership with process team members and area engineers. In this position you will supervise tool performance, schedule and perform preventative maintenance on assigned tool sets, and use mechanical, electronic and PC/software skills to solve and repair equipment issues. They are also required to assist with equipment installations, engineering tests, build and modify equipment procedures, increase tool uptime through systematic problem solving and quality workmanship, and thoroughly document all maintenance. This is a hands-on equipment focused role with time spent in the fabrication facility. </sp

PythonJavaMachine LearningAI
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron has an opening for a Photolithography Engineer (Level 2 or 3) at our Boise, Idaho facility. This position is part of Micron’s Technology Development organization, where advanced ideas become manufacturable technologies. Our team works at the leading edge of memory innovation, partnering across design, integration, manufacturing, and equipment suppliers to enable the next generation of semiconductor devices. This is where curiosity, experimentation, and impact come together. As a Photolithography Engineer, you will play a key role in developing and integrating lithography processes that directly shape future memory technologies. You will work hands-on with tools, materials, layouts, and data, while leveraging Artificial Intelligence and other software, collaborating with global partners to move innovations from R&D into high-volume manufacturing. This role offers broad technical exposure and the opportunity to influence device performance, yield, and cost at scale. Responsibilities: Develop and integrate photolithography unit processes for advanced memory technologies Perform material evaluations, reticle, and layout development, OPC, overlay optimization, and process window analysis Drive yield improvement, failure analysis, and device performance optimization Partner with equipment vendors to evaluate tools and improve process capability Document processes, communicate results, and support technology transfer to manufacturing Leverage Artifi

Artificial IntelligenceAIRecruitment
F
📍 Mclean, Virginia, United States
✓ High-confidence listingCompany trend -26.7%
Quick readStrong listing-quality and freshness signals

At Freddie Mac, our mission of Making Home Possible is what motivates us, and it’s at the core of everything we do. Since our charter in 1970, we have made home possible for more than 90 million families across the country. Join an organization where your work contributes to a greater purpose. Position Overview: We need a highly innovative Technical Lead! How confident are you that you can build sophisticated analytic systems? If you believe you could contribute to the development of innovative principles and ideas in a matrixed environment, please keep reading as we are seeking an individual contributor who has experience with Java and Python and can lead and nurture an inspiring environment in our Virginia office. Our Impact: The Investments and Capital Markets (I&CM) division is looking for a capable technology lead for its trading and analytics development team. This could be you! To thrive in this division, you must have a comprehensive understanding of system implementation and design, experience working in capital markets, and be enthusiastic about leading development of new paradigms in software system architecture. Your Impact: As a Trading Analytics Development Tech Lead, you will develop and maintain software using Java and Python tech stack that adheres to software engineering best practices. You will influence technical decisions, mentor developers, resolve engineering blockers, and partner with engineering managers to help teams deliver secure, reliable, and maintainable solutions. You will provide hands-on directions for full-stack applications, APIs, microservices, and integration services while reinforcing engineering discipline across design, development, testing, deployment, observability, and production readiness. Partner closely with Product Owners, engineering managers, architecture, business stakeholders, and cross-functional tec

PythonJavaReactAngular
B
📍 Berkeley, United States
✓ High-confidence listingCompany trend +515.8%
Quick readStrong listing-quality and freshness signals

Lead Systems Integrator Company: The Boeing Company Boeing Defense, Space & Security (BDS) is seeking a Lead Systems Integrator (Level 4) to support an Air Dominance Fixed Wing Proprietary Program in St. Louis, MO . Step into a fast-paced, cutting-edge program where your expertise will drive the design, integration, and testing of advanced, cloud-based systems that empower mission planning, debrief, and tactical Command & Control (C2) solutions for fixed-wing platforms. In this role, you will lead the development, analysis, and integration of innovative engineering solutions for critical Ground System capabilities and ensure seamless support for ground systems from initial design through to final delivery. You will partner closely with the team’s test lead as well as the overall chief integrator who oversees all activities related to design, integration, and test. You will work with a high-performing, cross-functional team in an agile environment, driving next-generation capabilities from design through delivery. This role offers the chance to innovate with open-architecture, model-based designs and collaborate across disciplines to support critical defense missions. You champion best practices, processes, and standards for the integration team, mentor systems and electrical engineers, and coach solution leads on execution. If you’re passionate about advancing mission-critical systems and thrive in a collaborative, fast-moving environment, this is your opportunity to make a significant impact. Position Responsibilities Lead team of systems and electrical engineers in behavior modeling and requirements decomposition for system, sub-system, and software design

Recruitment
B
📍 Huntsville, United States
✓ High-confidence listingCompany trend +515.8%
Quick readStrong listing-quality and freshness signals

Experienced Electrical & Test Engineer Company: The Boeing Company Boeing Defense, Space & Security (BDS) seeks an Experienced Electrical and Test Engineer to serve as an Automated Test Equipment Engineer to join our Patriot Advanced Capability-3 (PAC-3) Electrical Engineering Capability Team in Huntsville, AL . We design, sustain, and upgrade next generation electronics concepts and support production and special test equipment design for increasing capacity requirements. Position Responsibilities: Support design of automated test equipment, development of engineering documentation, and hardware fabrication, integration, and validation testing Develop, document, and maintain electronic and electrical system requirements according to customer desires and contract requirements Create and implement overall Test Equipment Design strategies Review and apply real-time test anomaly troubleshooting and resolution of issues encountered during hardware validation and verification testing Support fielded test equipment hardware and software over the entire product lifecycle Research emerging technologies for potential applications to meet projected requirements Perform Root Cause Corrective Action (RCCA) investigations, troubleshooting, and resolutions This position requires the ability to get a U.S. Security Clearance for which the U.S. Government requires U.S. Citizenship. An interim U.S. secret clearance Pre Start and final U.S. secret clearance Post Start is required. Basic Qualifications (Required Skills/Experience): Bachelor of Science degree in Engineering (with a focus in Electrical, Mechanical or Aeronautical), Computer Sc

PythonC#Recruitment
C
📍 New York New York United States, United States
✓ High-confidence listingCompany trend +800%
Quick readStrong listing-quality and freshness signals

About the Role Discover your future at Citi Working at Citi is far more than just a job. A career with us means joining a team of more than 230,000 dedicated people from around the globe. At Citi, you'll have the opportunity to grow your career, give back to your community and make a real impact. Job Overview Citi's Integrated Digital Assets Platform (CIDAP) is at the vanguard of institutional blockchain adoption — and security is its foundation. As digital assets move from innovation to regulated infrastructure, the cryptographic integrity of every transaction, wallet, and key lifecycle operation becomes mission-critical. We are building the security layer that the world's most sophisticated financial institution can trust. We are seeking a Senior Security Engineer (VP) to join our New York-based Digital Assets Platform engineering team. This is a hands-on, Java-focused backend engineering role for a security-minded engineer who understands both the craft of secure software development and the cryptographic primitives that underpin digital asset custody, signing, and key management. You will sit inside the core engineering team — writing production code every day — while being the resident authority on cryptographic design patterns, HSM integration, MPC protocols, and security architecture. Your work will directly protect billions of dollars of digital asset infrastructure used by institutional clients worldwide. Key Responsibilities Design, develop, and maintain security-critical backend services in Java — including cryptographic libraries, key management APIs, signing wor

JavaArtificial IntelligenceAI
A
📍 United States
✓ High-confidence listingCompany trend +365.2%
Quick readStrong listing-quality and freshness signals

Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: Position Overview The AI Platform Engineer builds and operates the machine learning and generative AI platform used by teams across Abbott Cancer Diagnostics. You'll own the full model lifecycle in production — data and feature pipelines, training and experimentation, evaluation and promotion, serving, and monitoring — along with the platform services, compute and tooling underneath it. This is hands-on infrastructure work backed by solid platform engineering practice: making inference fast and cheap, making the path from experiment to production repeatable and auditable, and shipping interfaces other engineers can build on — in support of software that ultimately reaches patients. Essential Duties Include, but are not limited to, the following: Build and maintain data, feature, and training pipelines for ML and LLM workloads — ingestion, transformation, fine-tuning, distributed training, and reproducible experiment execution with lineage tracked from dataset and code to resulting model. Implement automated evaluation and promotion gates — performance benchmarks, regression checks, and validation criteria that determine whether a model advances toward production. Automate the model lifecycle end to end through CI/CD and GitOps: packaging, promotion across environments, progressive rollout, and rollback. Build and operate production model-serving infrastructure for LLMs and predictive models, including inference optimization, autoscaling,

PythonJavaAWSKubernetes
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA is a global leader in high-speed computer vision, artificial intelligence (AI), and deep learning. Our team develops data engineering solutions that empower AI developers in autonomous vehicle (AV) domains to innovate quickly and effectively at scale. Are you ready to take on a senior technical role in building high-performance AI data pipelines? We seek an exceptional individual to design and optimize microservices and data pipelines to process massive volumes of AV data and enable seamless data mining and AI training. The ideal candidate will bring expertise in big data processing and distributed computing to create efficient solutions and overarching architectures for challenges such as video data curation, behavioral search, and AI dataset management. What you'll be doing: Scope and build tools, microservices, workflows, and distributed applications to accelerate data mining and AI training. Design and implement solutions for streaming, resilience, logging, security, authentication, workflow orchestration, and data management. Deploy AI models. Design and develop Retrieval-Augmented Generation (RAG) workflows enabling hybrid and agentic patterns. Analyze and operationalize complex distributed systems for speed-of-light performance. What we need to see: Experience developing high-performance, scalable software systems. MS with 6&#43; years, or BS (or equivalent experience) with 8&#43; years of relevant experience in Computer Science, Computer Engineering, or a related technical field. Strong programming skills in Python or Golang Proficiency in key technologies like Kubernetes, Helm, Hive, Parquet, SQL, vector databases, e.g., Milvus. Strong architectural skills with a proactive, problem-solving mentality. Experience in data mi

PythonSQLKubernetesArtificial Intelligence
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

The Silicon Co-Design Group (SCG) sits at the crossroads of architecture, design, marketing, operations, and productization. Our work spans early architecture through final product delivery across Datacenter, Gaming, Robotics, Automotive, and Embedded markets. We work closely across functions to deliver chips that change what is possible. NVIDIA’s Silicon Co-Design Group is hiring a Chip Lead to serve as the technical lead for one of our most consequential silicon programs! This is not project management, and it is not a senior IC role. You are the person the program partners with on its hardest technical questions — when the right answer is not obvious, and when leadership needs a single technical perspective to align on direction. You are accountable for the technical integrity of the chip end-to-end. You guide co-design feature integration, help resolve the toughest multi-functional bugs, and serve as the project-specific custodian of the qualification playbook. The program runs cleaner because you are on it! What you’ll be doing: You will partner across design, validation, software, and manufacturing to keep the program’s technical narrative clear and on track. Day to day, you will: Serve as the single technical point of contact for multi-functional decisions, issues, and trade-offs. Co-lead program-level feature integration from chip to system, surfacing inter-function dependencies and guiding them to resolution. Help resolve the program’s hardest multi-functional bugs by translating ambiguous, multi-team symptoms into root-cause closure on areas such as HBM, power and thermal, high-speed I/O, and packaging. Steward the qualification playbook. When the playbook does not fit a situation, guide the mitigations and capture the lessons as reusable methodology for other SSG programs. Shape the program’s technical narrative by surfacing key risks, trade

AIProject Management
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology and amazing people. Today, we're harnessing the boundless possibilities of AI to build the next era of computing. An era in which our GPU acts as the brain of computers, robots, and self-driving cars that can understand the world. Accomplishing unprecedented goals calls for imagination, inventiveness, and exceptional talent from around the world. As a NVIDIAN, you'll be immersed in a diverse, encouraging environment where everyone is inspired to do their best work. Join our team and discover how you can build a lasting impact on the world. NVIDIA's Silicon Co-Design Group (SCG) leads the full product development lifecycle, from early architecture definition through silicon bringup to product release. The ArchDev team is the hub for silicon and system-level feature development, driving tradeoff analysis, system integration, and POR alignment across the entire organization. This is where ideas become chips, and chips become products that define the state of the art — and we're building that future with some of the most motivated engineers in the industry. We're looking for a Senior Memory Systems Engineer to own HBM and LPDDR integration in sophisticated SoCs. This role covers the full stack, including silicon, package, embedded software, testing, and product development. The engineer will resolve the toughest system-level memory challenges throughout the process, building solutions that hold up at scale. What you'll be doing: HBM & LPDDR System Integration and Bringup: Drive HBM and LPDDR system integration, bringup, characterization, and debug for next-generation SoCs — taking memory subsystems through the full arc from first silicon to production-ready at scale. Full-Stack Memory Closure: <s

N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA's accelerated computing platforms move data at speeds that push the limits of what silicon and physics allow. Whether a high-speed interface trains reliably, maintains accurate margins, and survives every platform topology it will ever see is a question we answer ourselves. This role does that work. The Silicon Co-Design Group leads the boundary between what was designed and what was built. When a GPU, CPU, or SoC ships with interfaces that work at scale, this team is the reason. Most engineers debug within a layer. You will own the full stack. When an interface fails to train, the link margin is unexpectedly tight, or a customer reports a critical silicon issue, you trace the problem through protocol behavior, signal integrity, firmware, platform topology, and silicon marginalities. You then confirm that the fix works. Your methodology shapes how NVIDIA validates high-speed interfaces across generations. Your decisions affect yield, production ramp, and field quality. This is not a coordination role. The engineers who do it well hold protocol depth and system breadth simultaneously, never lose the thread across hardware, firmware, and software, and have the judgment to know when to go deeper and when to act. They are rare. Should that describe you, read on. What you'll be doing: Own post-silicon bring-up, characterization, validation, and debug of PCIe, NVLink, C2C, and other HSIO interfaces across NVIDIA GPUs, CPUs, and SoCs from first power-on through production readiness. Close the hardest failures. Drive root cause across protocol behavior, signal integrity, firmware and driver interactions, platform topology, and silicon marginalities and own every fix through to confirmation. Define validation strategy. Set test coverage, debug priorities, margining methodology, and stress criteria f

🔔

Get new software reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime