About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world’s most transformative technologies. Our new AI Engineering Campus in Austin will play a central role in building the future of AI computing. The Opportunity We are seeking a system validation engineering intern to help drive server blade and rack validation efforts for next-generation AI infrastructure hardware systems. This role focuses on post-silicon system validation across the full lifecycle of server hardware systems, ensuring functional and performance meets product objectives. You will help drive end-to-end blade and rack validation including development, execution, and debug while collaborating across silicon, firmware, systems, and platform teams. The Blade and Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Type: 12-week summer internship Timing: May - August (exact dates to be confirmed) Commitment: Full-time What You’ll Do Help drive and execute post-silicon validation goals of AI compute blades and racks including testcase planning, development, and automation Help drive validation testcase execution and system debug against program achievements and report validation progress and risks. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness Triage test failures, collect debug data, and collaborate on root cause analysis. Track validation coverage and continuously improve test processes and infrastructure. What You’ll Bring Working towards a Bachelor's
Jobs in United States
Engineering Specialist in United States
2,796 active opportunities · Updated October 2026
Showing
15 jobs
Explore current engineering specialist jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors, software, and data center systems that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our U.S. engineering teams contribute to the hardware and software platforms that support the next generation of AI systems. The Opportunity We are looking for a computer engineering, electrical engineering, or computer science student or recent graduate to join the BMC Development team as a Firmware Engineering Intern. You will work with experienced engineers on low-level and embedded firmware that supports the operation, control, and manageability of advanced compute systems. This internship provides hands-on experience in firmware development, test automation, engineering experiments, and lab-based system testing in a Linux development environment. You will own clearly defined technical tasks with guidance from the team and contribute to production-quality engineering work. Type: 12-week summer internship Timing: May - August (exact dates to be confirmed) Commitment: Full-time What You Will Do Contribute to the design, implementation, and testing of system and embedded firmware. Develop and maintain firmware and supporting software in C, C++, or Python. Support firmware development and debugging in a Linux-based engineering environment. Create automated tests and scripts that improve firmware validation, test coverage, and engineering efficiency. Contribute to continuous integration and delivery workflows for firmware development and testing. Plan and conduct well-defined engineering experiments, record results accurately, and draw conclusions from test data. Support lab setup, system configuration, hardware bring-up, and firmware testing. Use debugging and diagnostic techn
About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world’s most transformative technologies. Our new AI Engineering Campus in Austin will play a central role in building the future of AI computing. The Opportunity As a Mechanical Engineering Intern, you will contribute to the mechanical design and thermal validation of advanced AI hardware. Working alongside experienced mechanical and thermal engineers, you will gain hands-on experience with 3D computer-aided design (CAD), laboratory setup, test development, prototype evaluation, and engineering documentation. Type: 12-week summer internship Timing: May - August (exact dates to be confirmed) Commitment: Full-time What You’ll Do Contribute to mechanical design activities using Creo or comparable 3D CAD software, including part modeling, assemblies, drawings, and design updates under the guidance of experienced engineers. Work with the thermal engineering team to develop test procedures and help set up laboratory capabilities for evaluating AI hardware. Support thermal and mechanical testing using appropriate laboratory equipment; collect, organize, and analyze test data. Assist the mechanical and thermal design teams with prototype evaluation, troubleshooting, design verification, and documentation of findings. Collaborate with cross-functional engineering partners, communicate progress and issues clearly, and follow applicable laboratory and safety procedures. What You’ll Bring Current enrollment in a bachelor's, master's, or doctoral program in mechanical engineering or a related discipline during the internship. Students in other disciplines with relevant mechanical engineering exper
From $110K/yr
Datadog AI Research — Scholars Program with Carnegie Mellon University Datadog AI Research (DAIR) is partnering with Carnegie Mellon University to support a small number of PhD students working on open research problems grounded by ongoing efforts at Datadog/DAIR. You will frame a problem, run your own experiments, and write up what you find, with compute and data at a scale most academic labs cannot provide. You will collaborate with colleagues working on the same questions. The Lab And The Research: DAIR is an industrial research lab motivated by practical challenges in observability and software operation: detecting and diagnosing failures, understanding complex production environments, and helping engineers operate software more effectively. The lab focuses on creating specialized foundation models, post-training and evaluating AI agents, and building frontier-scale machine learning systems. By combining fundamental research with Datadog's large-scale, real-world data and infrastructure, the lab develops new AI capabilities and translates them into practical systems with meaningful impact. Internship projects are shaped with your DAIR mentor and your CMU faculty advisor. You do not need prior experience with observability, monitoring, or infrastructure. What You'll Do: Own a research project end to end: framing the question, running the experiments, writing it up Work directly with a DAIR mentor engaged in the same problem, and stay connected to your advisor and lab Publish, and use the work toward your dissertation See research reach production, when it works Who You Are: Currently enrolled in a PhD program at Carnegie Mellon in machine learning, computer science, statistics, or a related field Depth in at least one area relevant to the research above Comfort running real experiments — training models, working with GPUs, reading and reimplementing recent papers Evidence you can do research: conference or workshop papers, preprin
From $139.1K/yr
Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge OneTrust is building the AI-Ready Governance Platform™ to help organizations build trust, manage risk, and use data and AI responsibly. Our platform brings together privacy automation, consent and preferences, data use governance, AI governance, tech risk and compliance, and third-party management across connected, cloud-based workflows. As a Senior Staff Quality Engineer, SDET, you will define the quality engineering strategy that enables OneTrust to scale this platform with confidence. You will shape how we validate integrated customer journeys, protect trust-critical data and workflows, and release reliable capabilities across a broad product portfolio. This role is for an experienced technical leaderwho can solve significant, unique problems requiring independent judgment, evaluation of intangible factors, and the creation of methods and procedures for specialized, high-stakes initiatives. Your Mission Quality Strategy for Trust-Critical Products Define automation architecture, quality strategy, and release-validation practices acro
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are looking for an AI Agent Security Architect to function as the primary technical authority for Replit’s autonomous and AI agent security blueprint. In this critical role, you will design, implement, and maintain the runtime defense systems, guardrail frameworks, and sandboxing architectures that govern AI agents executing code, invoking tools, and reasoning across our platform. You will be a key technical contributor—leading high-impact AI security initiatives and bridging the gap between non-deterministic AI behavior and rigorous cybersecurity controls for both engineering and executive leadership. What You'll Do AI Agent Security Strategy & Technical Execution AI Agent Security Blueprint: Define the long-term vision and architectural patterns for securing autonomous agent workflows, Model Context Protocol (MCP) integrations, multi-turn reasoning loops, and multi-agent coordination. Runtime Guardrails & Policy Enforcement: Architect and deploy dynamic input/output guardrail systems, semantic firewalls, and real-time intent verification filters to prevent goal hijacking, system prompt leaks, and indirect prompt injections. Agent Execution & Tool Sandboxing: Partner with Infrastructure and AppSec teams to design secure, short-lived, micro-isolated environments (e.g., microVMs, WebAssembly, container sandboxes) where agents can dynamically execute code, run shell commands, and interact with host operating systems safely. Agentic Threat Modeling & Red Teaming: Conduct specialized threat modeling against non-deterministic systems. Lead automated and manual AI red-teaming initiatives to uncover vulnerabilities in RAG context pipelines, vector stores, and tool-calling interfaces. Identity
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. The Mandate: Industry Disruption and 10x Scale: Vision: Product Foundry is Replit's internal engine for innovation . We move Replit beyond being a collaborative development environment to becoming the foundational operating system for the entire next generation of software, where AI Agents are the primary actors. The 10x Goal: Our purpose is to launch high-risk, full-stack 0→1 initiatives that define and establish entirely new, multi-billion-dollar product categories. We are seeking non-linear growth opportunities, fundamentally aiming to 10x Replit's value and addressable market by proving out unprecedented technical primitives and disruptive Go-To-Market strategies. The Audience: We build for the next generation of creators and high-leverage users and enterprises, equipping them with tools that enable them to build anything, anywhere . Candidates that do well here will certainly go on to build their own companies in the future! This is a high visibility role reporting to Execution Model: High-Agency Founding Teams Structure: We operate as a collective of in-house technical founders —not just specialized engineers. Initiatives are run by lean, autonomous squads built for velocity and maximum technical leverage. This model is centered around an Engineer DRI (Directly Responsible Individual) who maintains total ownership over the initiative's technical, product, and launch success, supported by fractional PM and Design resources. Cadence: We enforce rapid iteration and rapid market validation via 3-week sprints per initiative. This cadence forces fast deployment, immediate user feedback, and tight alignment with the internal betting table process, mirroring the intensity and speed of a lean startup. Required skills and
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role As a Product Engineer, you’ll take emerging ideas from concept to polished experiences. You’ll build sophisticated, intuitive user experiences while working across AI models, agentic workflows, and all parts of web infrastructure to ship complete products. The Goal: Build products that can define major new categories for Replit. You will combine the flexibility of software with the accessibility of presentations, spreadsheets, dashboards, and visual editors, enabling anyone to create powerful, interactive artifacts through natural language. The Audience: We build for the next generation of creators, knowledge workers, and high-leverage teams. Our users should be able to turn an idea, dataset, or business problem into a polished, editable, and shareable product just by themselves. Candidates who thrive here often have the product judgment, technical range, and ownership mindset associated with founders. This is a high-visibility role working closely with product leadership. How we work Structure: We operate as a collective of in-house technical founders, not narrowly specialized engineers. Initiatives are run by lean, autonomous teams built for velocity and technical leverage. Engineers maintain meaningful ownership over technical direction, product quality, and launch success, supported by Product and Design partners. Cadence: We work in focused three-week sprints, shipping quickly, observing real user behavior, and incorporating feedback into each iteration. We value decisive execution, thoughtful experimentation, and the ability to make progress without perfectly defined requirements. What you’ll do Own major product experiences from initial concept through design, implementation, launch, measurement,
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Summary The Senior Flex Solution Engineer (SE) is a high-impact, customer-facing technical role supporting AI/Machine Learning (AI/ML) initiatives across our strategic US Major Accounts. This role is designed for a technical professional who possesses a strong foundation in AI/ML concepts, technologies, and solution architecture, positioning them above a generalist but not necessarily requiring deep specialization (Level 200-300 technical depth). The Flex SE will act as a critical, hands-on technical resource, accelerating customer adoption and success. This role translates business challenges into AI/ML- driven solutions through the rapid development of MVPs and prototypes, supporting our regional SE teams, and driving growth in this strategic area. The AE and SE maintain ownership and ultimate approval over the account strategy. The Flex SE role is designed to be supportive, not to supersede their authority. The RVP should provide prescriptive guidance on which accounts the Flex SE should prioritize, as the RVP possesses the most comprehensive regional overview. Key Responsibilities and Scope Technical Leadership & Solutioning Partner with regional Account Executives (AEs) and Solution Engineers (SEs) to identify, qualify, and develop AI/ML opportunities within USMajo
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will play a key role in enabling our mission to deliver a reliable, high-performing, and cost-efficient infrastructure that powers the world’s play. In this specialized role, you will be the technical lead for our GPU and AI accelerator ecosystem. You will be responsible for the full lifecycle of GPU hardware, from initial architectural evaluation and firmware qualification to large-scale fleet integration and performance tuning. You will ensure that Roblox’s massive-scale rendering and ML workloads run on the most optimized and stable hardware possible. You Will: Architect & Prototype: Prototype next-generation GPU-accelerated hardware platforms, ensuring seamless integration between high-density compute nodes, high-speed interconnects (NVLink/PCIe Gen5/6), and system firmware. GPU Optimization: Drive the integration, performance testing, and debugging of GPUs in our fleet, focusing specifically on hardware-level optimizations, driver tuning, and thermal/power management. Validation & Certification: Develop and execute rigorous evaluation and stress-testing strategies for GPU-heavy server platforms to ensur
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Orchestration pod within Engine Productivity, you'll design and run the platform that executes large-scale end-to-end and integration tests, running the real, shipping client on real hardware, across Roblox's data centers, cloud, and our own device labs, so our engineering teams can ship the engine, clients, Studio, and more with speed and confidence. Every Roblox engine, client, and Studio change, along with the experiences built on top of them, should ship with confidence, and the Orchestration team is the layer that makes that possible. We build large-scale distributed services that turn thousands of test suites into a reliable, push-button pipeline: fanning work out across fleets of machines and real devices, moving artifacts to where they're needed, managing single- and multi-client test state, and giving test owners and maintainers a system to validate their own runs. It looks a lot like building a specialized cloud platform, with capacity-aware scheduling, isolation and sandboxing, and smart retry and backoff, plus the classic distributed systems problems (fairness, efficiency, failure handling, and reliability) at Roblox scale. You Will: Design a
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is a rapidly growing Series D funded startup that specializes in providing innovative developer products and solutions for both individual developers and large enterprises. Our mission is to empower developers and organizations to create, test, and manage APIs more efficiently. We are currently seeking an experienced and highly skilled Field CTO to join our team, reporting to the CTO. In this role, you will work closely with customers to understand and guide their API strategy, ensuring Postman's products and services are seamlessly integrated into their architecture. What You’ll Do Act as a trusted technical advisor to Postman customers, helping them understand and navigate their API strategies and identify opportunities for Postman's products and services to add value. Collaborate with the sales and customer success teams to provide technical guidance during the pre-sales, onboarding, and ongoing support processes. Develop and deliver customer-facing presentations, demos, and workshops to demonstrate the value of Postman's solutions and help customers optimize their API management processes. Gather custome
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto
About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world's most transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the future of AI computing. The Opportunity As Technical Services Director, you will lead the teams that operate and evolve Graphcore's engineering labs, high-performance computing (HPC) platforms, and data center environments globally. You will be accountable for reliable, secure, cost-effective infrastructure that supports demanding engineering, AI, silicon-development, and validation workloads. This role combines people leadership, infrastructure strategy, operational excellence, capacity and financial planning, procurement, and program delivery. You will partner with Engineering, Information Technology, Security, Finance, Facilities, Supply Chain, customers, and external suppliers. The position is based onsite in Austin and requires travel to company facilities, data centers, and supplier locations, including international travel. What You'll Do Lead, recruit, mentor, and develop the systems administration, lab operations, and technical services teams responsible for the facility supporting global Engineering and Research and Development. Own the reliability, efficiency, protection, safety, supportability, and continuous improvement of engineering labs, HPC systems, and infrastructure facilities. Establish service levels, operating standards, escalation paths, performance measures, monitoring, observability, automation, ticketing, and configuration-management practices. Translate engineering and customer requirements into infrastructure roadmaps, capacity p
Mechanical and Thermal Laboratory Technician Position Summary Graphcore is a globally recognized leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data center hardware that provide the specialized processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Austin, which will play a central role in Graphcore's work building the future of AI computing. Responsibilities The Mechanical and Thermal Laboratory Technician is a hands-on technical role supporting the development and validation of advanced AI hardware systems for data center environments. Working as part of a cross-functional engineering team, this individual will be responsible for executing mechanical and thermal laboratory testing, supporting product validation activities, prototype fabrication and assisting with troubleshooting and root-cause analysis of complex hardware systems. Requirements Associate degree in Mechanical Engineering Technology or a related technical field preferred. Equivalent combinations of education, training, and relevant experience will be considered, including experienced non-degreed candidates or candidates with degrees in unrelated disciplines. Minimum of 5 years of experience working in mechanical laboratories, machine shops, test labs, or similar technical environments. Experience with server hardware platforms and data center equipment. Knowledge of Direct Liquid Cooling (DLC) systems and their implementation in server environments. Experience operating forklifts, pallet jacks, and other material-handling equipment. Key Responsibilities Execute mechanical and thermal test plans to validate hardware designs a
Other cities to consider
More places hiring for this role
Get new engineering specialist jobs in United States by email
Daily job updates · Unsubscribe anytime