About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
Jobs in United States
Technical Operations Engineer in United States
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current technical operations engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
$960K – $1.4M/yr
Position Overview As SingleStore’s IT Operations Engineer, you will help shape the IT toolset used by our end users. This is an active, hands-on position responsible for the planning, design, development, and Tier 1 support of several key technical areas at the SingleStore IT team, including end-user support, client engineering, executive support, and infrastructure application support. This is an incredible opportunity for someone to build upon their technical strengths and be a part of IT at SingleStore team . Roles and Responsibilities: Administering a wide variety of SaaS applications. Some main applications that need to be supported are OKTA (+ Workflows), Google Workspace, Slack, and Atlassian tools (JIRA + Confluence), MDM administration. Keep up to date with new features and new releases in these applications to identify opportunities for better automation or features that could be useful for our environment. Seize opportunities across the IT Operations team to eliminate manual work through tooling, integrations, and automation of IT workflows. Respond to tickets and execute new hire onboarding and user separation processes. Support members of the team with troubleshooting and resolution of complex issues. Design, architect, implement and maintain systems and solutions for various IT-related topics, including but not limited to staff computer hardware, operating systems, software applications, networking, videoconferencing, and printers. Partner and collaborate with all business units to help them evaluate hardware and software solutions. Able to communicate effectively and concisely with the entire company. Analyze existing processes, suggest and make improvements, and implement business processes where none exists. A desire to learn and expand your horizons; take on new challenges as the business scales Required Skills and Experience: Minimum 2 years of relevant experience Prior experience in implementing and administering Google Workspac
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are looking for an IT Support / Operations Engineer to join Baseten as we continue to scale our IT team. In this role, you will play a critical part in bringing our technical support entirely in-house to provide a seamless, high-touch experience for all Baseten employees. As we continue to scale, you will be the primary point of contact for day-to-day technical issues, allowing you to have a direct impact on our team's productivity and overall office environment. This position is ideal for a hands-on problem solver who enjoys a mix of hardware and software troubleshooting, user lifecycle management, and maintaining the physical IT infrastructure of a modern office. While you will focus heavily on elevating our internal support standards, you will also assist with systems administration and workflow automation as our company evolves. This is a hybrid role based out of our San Francisco or New York office, following our standard policy of three days per week in-person to ensure our physical office and AV systems remain high-performing and reliable. RESPONSIBILITIES Serve as the escalation point for day-to-day technical support, diagnosing and resolving hardware and software issues across our Mac and Windows fleet Manage user lifecycle administration including provisioning, deprovisioning, and access management across all systems and services Own the IT onboarding experience for new employees — from laptop set
$192K – $240K/yr
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Security Operations Engineer at Brex, you will focus on preventing, detecting and responding to security threats across Brex's corporate and cloud environments. You will use existing systems and develop tools to improve our security capabilities. Our team is responsible for functions across corporate security, detection & response and infrastructure security domains; and we perform systems engineering and automation to support those functions. Security Operations is part of our wider Trust & IT o
$192K – $240K/yr
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Security Operations Engineer at Brex, you will focus on preventing, detecting and responding to security threats across Brex's corporate and cloud environments. You will use existing systems and develop tools to improve our security capabilities. Our team is responsible for functions across corporate security, detection & response and infrastructure security domains; and we perform systems engineering and automation to support those functions. Security Operations is part of our wider Trust & IT o
$192K – $240K/yr
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Security Operations Engineer at Brex, you will focus on preventing, detecting and responding to security threats across Brex's corporate and cloud environments. You will use existing systems and develop tools to improve our security capabilities. Our team is responsible for functions across corporate security, detection & response and infrastructure security domains; and we perform systems engineering and automation to support those functions. Security Operations is part of our wider Trust & IT o
About the Team The Technical Support team is responsible for ensuring that developers and enterprises can reliably build mission critical solutions using OpenAI models. We provide technical guidance, resolve complex issues and support customers in maximizing value and adoption from deploying our highly-capable models. We work closely with Technical Success, Product, Engineering and others to deliver the best possible experience to our customers at scale. We think from an automation-first mindset and leverage the latest in AI to scale our support operations. Join the Senior Support Engineering (SSE) team at OpenAI and help shape the future of Technical Support in the age of AI. About the Role We are looking for a Senior Support Engineer to collaborate directly with our strategic enterprise accounts and product teams, helping solve some of the most difficult problems faced by our Customers. You will be part of the best technical troubleshooting team at OpenAI, and our Customers and Engineering teams will look to you for technical guidance in addressing the most technically difficult issues in our environment. As a Senior Support Engineer, you will design and run operational processes to monitor our top strategic customers and a 24x7 response team. You’ll work closely with our Infrastructure and Engineering teams to deliver the best possible experience to customers at scale. Working directly with our most strategic Customers - You will be crucial to the success of the most innovative, disruptive, and high-scale AI solutions being built with the OpenAI API platform. The nature of this role will be low volume, high difficulty. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Be among the foremost technical and troubleshooting experts for our API platform at OpenAI. You are the last line of defense before the core Engineering team. Proactively iden
About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and
About Hexnode Hexnode, the Enterprise software division of Mitsogo Inc., was founded with a mission to simplify the way people work. Operating in over 100 countries, Hexnode UEM empowers organizations in diverse sectors. Fueling the transformation to a seamless ecosystem of connected tools, Hexnode is revolutionizing the enterprise software and cybersecurity landscape. Role Overview We are seeking a AWS Operations Specialist to manage and maintain our cloud infrastructure and device ecosystems. This is a highly operational, execution-focused role—not an architecture position. The ideal candidate has 2 to 4 years of experience executing infrastructure as code, monitoring environments, and following documented playbooks to keep our systems secure and resilient. Because this role handles secure environments, candidates must be US Citizens and capable of passing a comprehensive federal background check. Key Responsibilities Infrastructure Execution: Run, maintain, and execute existing Terraform and Ansible scripts to deploy and update infrastructure. GovCloud Monitoring: Actively monitor our AWS GovCloud dashboards, keeping a close eye on system health, performance metrics, and security baselines. Mobile Device Management: Manage Android Enterprise kiosk configurations, ensuring secure deployments and smooth device operations. Incident Response & Triage: Respond swiftly to operational alerts by strictly following our documented team playbooks. Escalation: Identify anomalies or issues that fall outside established, documented procedures and escalate them accurately to the engineering team. Required Qualifications & Profile Citizenship: Must be a US Citizen (required for GovCloud infrastructure management). Background: Must be able to successfully clear a rigorous federal background investigation. Experience: 2 to 4 years of hands-on experience in a technical operations, DevOps, or SysAdmin role. Technical Familiarity: Comfort executing/running Terraform and
About the Team Governance, Risk, and Compliance (GRC) is foundational to Security delivering mission outcomes at OpenAI. The GRC team provides security assurances and builds compliance for OpenAI’s technology, people, and products. We are technical in what we build but operational in how we do our work, and we partner deeply with Product, Security, Legal, Privacy, GTM, and Field Security to help OpenAI move quickly while maintaining trust with customers, auditors, regulators, and the public. About the Role We are looking for an experienced Product Lifecycle Assurance IC to help scale OpenAI’s GRC function across our product stack to ensure products address customer and regulatory compliance requirements at launch and regressions are detected promptly and corrected. You will partner closely with Product, Security, Legal, and Privacy teams to make sure OpenAI can move quickly while maintaining our security, privacy and compliance claims and giving customers, auditors, and regulators assurance about how OpenAI handles user data. You are responsible for product assurance end-to-end from inception to post-launch (continuous) monitoring. You leverage existing workflows, reviews, and data and enhance, augment and build the components needed to create an end-to-end product assurance program. This role is not about supporting SOC or ISO audits; it's a highly cross-functional and deeply technical operations role to ensure that OAI products meet the compliance bar at launch, regressions are prevented and detected, and our compliance state can be evidenced. This role also helps ensure that our product launch governance program operates effectively across key safety, privacy, legal and security stakeholders and lessons-learned from incidents and regressions are used to improve the program. You or in partnership with engineering teams, build key controls in our infrastructure stack, developer workflows and launch tooling to provide developers with guardrails, perform CI/CD confor
From $86K/yr
Datadog’s Technical Solutions organization includes 1,200+ sales engineers, support engineers, post-sales experts, and solution architects. They run on an ecosystem of enterprise platforms and internal tools that directly shape how we serve customers. Technical Solutions Operations (“TSO”) owns that ecosystem. We manage the full lifecycle of the systems TS depends on: Zendesk, Jira, Confluence, and a growing portfolio of off-the-shelf and purpose-built tools. We do the work to operate, maintain, and evolve the platforms powering daily workflows across TS. When a vendor tool reaches its limits, we extend it through customization, integration, or targeted solution development, tapping internal partners across Datadog as needed. We’re looking for a Systems Engineer who wants to own enterprise platforms end-to-end. You go from understanding the business process, to designing the right solution (whether that’s configuration, integration, or code), to measuring whether it actually moved the needle. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own enterprise systems through their full lifecycle. You’ll be the technical owner for one or more platforms that TS relies on daily. That means understanding how the system is used, where it’s falling short, what’s coming from the vendor roadmap, and what needs to change. You drive improvements from assessment through implementation. Engineer solutions that create leverage. Not every problem is solved by configuration. You’ll build and evolve enterprise systems, integrations, automations, and internal tools that multiply the effectiveness of 1,200+ technical experts. Where AI can make a solution smarter (e.g., intelligent routing, automated triage, agent-assisted workflows), you'll include AI in the initial design, no
From $224K/yr
Datadog's Technical Solutions (TS) organization is one of the largest organizations in the company — spanning Sales Engineers, Technical Account Managers, Enterprise Customer Success Managers, Technical Support Engineers, and Solutions Architects who work with prospects and customers across every stage of their journey with Datadog. Technical Solutions Operations (TSO) exists to make that organization faster, smarter, and more scalable. We build the systems, analytics & programs that give TS teams more leverage — and we measure our success by the business outcomes we drive, not the projects we complete. The programs this team runs touch every function in TS, and the operating model you build will define how that scales. We're looking for a Director of Technical Program Management to lead the TSO Program Management team. This role sits at the intersection of strategy and execution: you'll own the programs that shape how TS operates at scale, lead a team of technical program managers, and serve as a peer to the Directors and VPs who run the teams you support. Your counterparts are leaders overseeing hundreds of customer-facing technical professionals, and your role is to successfully interface with each organization with proactive solutions on how your team can help them be more effective. This is a rare opportunity to lead a function where the output isn't a deliverable, it's organizational capability. If you want to build something that compounds across an entire organization, this is the role. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do Run the PMO as a business impact function. Own the PMO operating model end-to-end, including intake, prioritization, scoping, execution, and impact measurement — with every program directly linked to measurable bu
From $126.8K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Data Center Engineer , you'll help us scale our Core/Edge Data Centers and hardware infrastructure at a time of incredible growth for our business. At Roblox, you'll have boundless opportunities to shape the future of the Imagination Platform™ and demonstrate your passion for delivering thoughtful solutions in front of a global audience. If you know what it takes to build and operate hardware infrastructure that can sustain millions of concurrent players year-round and you take play as seriously as we do, you'll fit right into our highly experienced and ever-expanding engineering team. You will report to the Technical Lead Data Center Engineer. You will: Develop and maintain the Core/Edge Data Center and hardware infrastructure to meet the large scale and real-time requirements of our Imagination Platform™ to ensure our community has an awesome experience anywhere in the world. This includes all aspects of the server, network infrastructure, power, and environmental life cycles. Own efforts to track and mitigate systemic issues preventing hosts from returning to service. Identify and solve critical problems and prevent them from re-occurring via root cause analysis and giving rec
From $116K/yr
We are seeking a motivated and experienced Technical Program Manager II to join Datadog’s Technical Solutions organization. This organization is made up of 1,200+ customer-facing technical experts around the world — including Sales Engineers, Technical Account Managers, Support Engineers, and Solution Architects — who work with prospects and customers throughout their journey with Datadog to deliver outstanding experiences and drive growth through product adoption. As a Technical Program Manager II, you will lead and support a variety of technical programs that help these customer-facing teams work more effectively. Depending on the needs of the organization, this could include programs related to internal tooling, process improvement, knowledge and content systems, or other cross-functional initiatives. You’ll partner closely with team leads, subject matter experts, and other stakeholders to bring structure and clarity to moderately complex, cross-team programs, and you’ll act as a multiplier for the teams you support by driving programs from planning through delivery. Candidates who have previously led projects for customer-facing technical teams are especially encouraged to apply. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do Drive delivery of cross-team programs that support one or more Technical Solutions teams — coordinating stakeholders, dependencies, and timelines to keep work on track from kickoff through delivery. Partner with team leads, subject matter experts, and cross-functional stakeholders to translate program goals into clear, actionable plans, and document scope, timeline, and quality expectations along the way. Establish and maintain program management fundamentals — project trackers, status updates, risk logs — so stakeholders always
Other cities to consider
More places hiring for this role
Get new technical operations engineer jobs in United States by email
Daily job updates · Unsubscribe anytime