Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. Discord's Enterprise Technology org is building a software engineering discipline from the ground up, and this is the founding engineering hire. There is no Enterprise Platform team yet. You will not be maintaining a status quo or joining an existing backlog: you will define the discipline, set the technical foundation, and build the platform that lets the internal systems enabling Discordians scale with the company. Discord's mission is to cultivate belonging, and doing that at scale requires internal systems that move as fast as the product itself. The mission of this role is to turn our enterprise infrastructure into self-service software. Your job is to replace manual consoles and tickets with a platform of reusable primitives, building blocks like "provision a device" or "push an IAM policy or app integration through code," exposed through clean APIs and self-service portals. As a code-forward Staff engineer, you will design abstractions that hide the messy reality of Okta, Jamf, and Terraform (our identity, device management, and infrastructure-as-code tools) behind interfaces people actually want to use, operating with a platform-as-a-product mindset. What You'll Do Manage the device fleet as code: Move device configuration off consoles and into versioned, API-driven infrastructure for endpoint management and hardware provisioning. Automate the identity layer: Manage identity systems and their connections to downstream apps as code, from SSO and SCIM integrations to account workflows, so changes ship through version control and CI/CD instead of an admin console. Own app delivery and Infrastructur
Jobiba hiring network
Terraform Jobs
657 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current terraform jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. The Realtime Infrastructure team is responsible for building and maintaining some of Discord’s highest scale and most critical services. Those systems are at the core of our text chat infrastructure and facilitate the dispatching of every update to our users sessions. This role will have a significant impact on Discord’s overall reliability and performance. It will also help our product teams build new features on top of our infrastructure. This team is small but critical, and its work has a direct impact on Discord's success and ability to scale. This role reports to the Senior Engineering Manager of Realtime Infrastructure. What You'll Be Doing Build and operate large-scale, reliable and performant distributed systems. Collaborate with product teams to create new features. Ensure Discord “just works”. Write code but also manage our infrastructure. Work with a talented team of engineers who have built one of the largest communication platforms in the world. What you should have 2+ years of experience writing and designing backend systems. Experience solving complex distributed system problems. Experience operating and maintaining critical tier 0 services. Knowledge of monitoring and alerting best practices. Familiar with open source software, and not afraid to dig into the source code of a library to find the answer you’re looking for. Bonus Points Experience with Elixir or Rust. Experience working with systems deployed in a cloud environment (GCP, AWS, etc.) Knowledge of devops tools like Salt,Terraform or k8s. You have built or contributed to open source projects. You are a Discord power user and hav
Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. We are seeking a QA/DevOps Engineer to join Discord’s Business Systems team. In this role, you will own automated test suite development and maintenance across Oracle ERP Cloud (Fusion) and Salesforce, with a focus on Tosca-based regression and progression testing. You will embed automated testing into CI/CD pipelines and partner closely with ERP Functional, Technical, and DevOps teams to ensure every release is tested consistently and defects are caught early. What you'll be doing Run Oracle ERP Cloud and Salesforce QA automation suites for all quarterly patch upgrades end-to-end. Review new Oracle and SFDC release features and recommend improvements to existing testing processes. Build and improve regression and progression test automation in the Tosca environment for Oracle Fusion and Salesforce. Create, maintain, and update Tosca scripts to reflect system and process changes across GL, AP, AR, and FAH modules. Lead feature enhancement reviews as part of quarterly patch assessment cycles and new enhancement projects. Integrate automated test suites into CI/CD pipelines; trigger regression suites automatically on every release candidate build. Maintain test infrastructure and environments in collaboration with DevOps and ERP Technical teams. Document test coverage, defect trends, and release readiness in Jira and shared reporting dashboards. Automation Use Cases FAH Process Automation: Auto-test Financial Accounting Hub on each new sales item addition to eliminate manual unit testing. AR Regression Suite: Automated flows for invoicing, receipts, and customer setup across every release.
Who We Are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the most ambitious startups — use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the Team The Core Change Management group is responsible for the systems that let every Stripe engineer ship code, configuration, and infrastructure changes safely and at high velocity. You will be embedded primarily on the Service Deployments team — the owners of Stripe's end-to-end code deployment platform — with regular collaboration with the Resource Automation and Feature Deployments teams. What Makes This Role Compelling You own the foundation of how Stripe ships software. The deployment platform sits in the critical path of every engineer's workflow at Stripe. The decisions you make affect thousands of deploys per day across hundreds of services, directly determining how fast and safely Stripe's product evolves. Technically rich, architecturally active. The team is executing several concurrent platform transformations: containerizing host-based services at scale, adding intelligent multi-service deploy pipelines, extending real-time anomaly detection to earlier stages of traffic shifts, and rebuilding deployment event infrastructure on top of a durable message bus. This is not maintenance work — the architecture is in motion. Broad surface area, real ownership. You will span the full stack from container scheduling and deployment orchestration business logic to the developer-facing internal platform UI. The problems are multi-layered: reliability, developer experience, performance, and safety all at once. Your judgment prevents incidents.
Who we are About Privy Our mission is to make privacy and user ownership the default online. To do so, we build simple, flexible APIs and tools for developers that make it easy to build new products on crypto rails. Privy owns the abstractions and infrastructure layer above wallets, integrating across chains, third-party providers, and Stripe products like Treasury and Link. We get to solve hard technical problems while leveraging Stripe's distribution to reach customers like Ramp, Klarna, Deel, Kraken, Hyperliquid, and Fomo — powering experiences for both mainstream users and crypto natives. Learn more about Privy: Privy and Stripe: Bringing crypto to everyone About the team Engineering at Privy is distinguished by: High urgency: Shipping very small iterations, very fast, to learn very quickly. Product taste: Our customers are developers, and to build effective products for them requires technical knowledge - you will often be "the PM". Security mindset: A great portion of our product is trust. While we have a dedicated security team, every engineer brings security to their designs from the start. In practice, we use boring technology like Node, React, and AWS so we can focus our engineering energy entirely on pushing the boundaries of Privy's core product, e.g. through hardware enclaves, multi-region low latency APIs, and blockchain abstractions that are accessible to mainstream developers. What you’ll do Manage and architect multi-tenant infrastructure that serves billions of requests every month Design our build and deploy systems across the stack, from SDKs to apps to high-traffic APIs to complex secure enclaves, and more – all for a team that ships daily Drive observability across our product and help keep our systems up and performant in our role in the most critical parts of our customers' stacks Advocate for making the right technical investments at the right time in a company with ever-evolving technical needs Who you are Minimum requirements 8
About the Team Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. We operate some of the largest Spark compute fleets in production; design, and build data lakes and metadata systems on Iceberg and Delta with a vision toward exabyte-scale architecture; run high throughput streaming platforms on Kafka and Flink; provide orchestration with Airflow; and support ML feature engineering tooling such as Chronon. Our mission is to deliver reliable, secure, and efficient data access at scale and accelerate intelligent, AI assisted data workflows. Join us to build and operate these core platforms that underpin OpenAI products, research, and analytics. We’re not just scaling infrastructure – we’re redefining how people interact with data. Our vision includes intelligent interfaces and AI-assisted workflows that make working with data faster, more reliable, and more intuitive. About the Role This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You’ll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You’ve supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You’re well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per wee
About the Team The Applied AI team safely brings OpenAI's technology to the world. We released ChatGPT, Plugins, DALL·E, and the APIs for GPT-4, GPT-3, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. We serve end-users directly through ChatGPT, and serve developers through our APIs, which power product features that were never before possible. About the Role The Engineering Acceleration team designs, builds and maintains the foundational systems that engineers use to build ChatGPT and the API. This is a fast-growing team and you will get a chance to own and define the strategy, vision, and plan for how to increase developer productivity. In this role, you will: Drive the design, development, and implementation of tools, systems, and processes that accelerate engineering velocity, reduce manual effort, and increase the quality of output. Use our latest AI tools to re-think how we can be the most productive team in the industry. Work closely with various teams within OpenAI to understand their workflows, challenges, and needs, and ensure the tools and systems built by the Engineering Acceleration team address these requirements. Bring new features and research capabilities to the world by partnering with product engineers to lay the necessary technical foundations. Guide and advise product engineering teams on best practices for ensuring observable, scalable systems. Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years of experience in engineering, including 3+ years of experience in infrastructure building tooling for developers. Have experi
About the Team The Infrastructure Engineering function sits within IT and is responsible for reliably building, deploying, and operating critical on prem and hybrid environments that power internal services and critical R&D environments. This is an early, high-leverage technical role focused on applying strong Site Reliability Engineering discipline to environments where uptime, safety, recoverability, and security are non-negotiable. This person helps replace bespoke, one-off infrastructure with standardized infrastructure-as-code building blocks that compound reliability and operational leverage as OpenAI scales. About the Role We are looking for an experienced Site Reliability Engineer working on security infrastructure to design, build, and operate reliable, secure, and scalable infrastructure that underpins identity, access, endpoint, and shared platform services across the company. In this role, you will be a senior technical owner for infrastructure and identity systems end to end, from architecture and implementation through policy enforcement, upgrades, recovery, and day-two operations. You will build durable, production-grade platforms that remove operational friction, enforce security by default, and enable teams to move faster with confidence. This role is well suited for a hands-on senior engineer who thrives in ambiguity, enjoys owning complex systems end to end, and raises the reliability and security bar by replacing fragile implementations with standardized, repeatable infrastructure. This role is based in our San Francisco HQ and requires in-office presence. In this role, you will: Design, build, and operate reliable infrastructure across on-prem, hybrid, shared, and product adjacent environments. Establish standardized infrastructure patterns that replace bespoke implementations with repeatable, auditable, secure-by-default systems. Own the lifecycle of critical infrastructure platforms, including provisioning, deployment, upgrades, patching,
About the Team The Applied Foundations team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale. The Applied Foundations team is at the front lines of defending against financial abuse, scaled attacks, and other forms of misuse that could undermine the user experience or harm our operational stability. Integrity Foundations provides the core building blocks and infrastructure for this work. About the Role At OpenAI, our mission is to advance AI in a way that is safe, reliable, and aligned with broad societal values. The applied foundations role is crucial for maintaining the trustworthiness of our platforms. You will be pivotal in developing robust defenses against a spectrum of adversarial behaviors that threaten our ecosystem. In this role, you'll work with our entire engineering team to design and implement systems that detect and prevent abuse, promote user safety, and reduce risk across our platform. You'll be at the forefront of our efforts to ensure that the immense potential of AI is harnessed in a responsible and sustainable manner. In this role, you will: Develop and enhance systems to detect and prevent various forms of abuse including financial fraud, botting, and scripting. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. Assist with response to active incidents on the platform and build new tooling and infrastructure that address the fundamental problems. You might thrive in this role if you: Have at least 3 years of professional software engineering experience. Have experience setting up and maintaining production backend services and data pipelines. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Are self-directed
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We are seeking a Software Engineer, Security Observability to join our Security team. In this role, you will be responsible for building secure, scalable systems that enhance our security observability infrastructure. Leveraging your strong engineering skills, you will collaborate with cross-functional teams to develop, deploy, and maintain robust software solutions that support our security and detection capabilities. This role is open to remote employees, or relocation assistance is available to one of our OpenAI offices in San Francisco, Seattle, or New York City. Due to requirements associated with work this role may support, applicants for this position must be U.S. citizens. In this role, you will: Design and develop scalable software systems that facilitate security observability across our infrastructure. Build and maintain data pipelines that centralize and store security-relevant data from diverse sources. Proactively improve the resilience and reliability of data systems to ensure high platform availability Collaborate closely with Detection & Response (D&R) and other security teams to reduce the company’s security risk. Contribute to data engineering in support of forensic investigations and compliance efforts. You might thrive in this role if you have: Strong software engineering experience, with proficiency in programming languages such as Python, Golang, or similar. A background in infrastructure as code, with exp
About the Team The Agent Infrastructure team at OpenAI is responsible for building systems that enable training and deployment of highly useful AI agents, both internally and for the world. We work hand-in-hand with researchers to design and scale the environment in which agentic models are trained – providing a workspace for AI models to execute code, debug issues, and develop software just as human SWEs do. Our training environment for agentic models operates at an extremely high scale and has the flexibility to emulate any environment in which an agent might work. At the same time, our team builds and maintains OpenAI’s core platform for the deployment and execution of agents in production. Our systems power products such as Codex, Operator, tool use in ChatGPT, and future agentic products. Some of the most challenging technical problems in scaling the capabilities and utility of agents and agentic models lie in the infrastructure layer – and our team is focused on building the research and production systems that enable OpenAI to train the most capable models in the world, and maximize the utility of our agentic products for users around the world. About the Role As a Software Engineer on the Agent Infrastructure team, you will have the opportunity to work closely with both research and product at OpenAI - building and scaling systems to train highly capable agentic models, and building the platform and integrations to launch new agents to hundreds of millions of users worldwide. Your work will consist of both building new capabilities - standing up the infrastructure and integrations needed to train more complex agentic models - and rapidly scaling these new capabilities to some of the largest compute clusters in the world. At the same time, you’ll be instrumental to the launch of agentic products at OpenAI - building, maintaining, and scaling the production platform on which all agents run. We’re looking for people with deep experience building AI infrastructu
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role Our technologies support some of the most important and impactful work in the world, including our strategic and high-impact customers in the public sector. As a Forward Deployed Security Engineer (FDSecE) you will be responsible for securing these novel applications of OpenAI’s technology. We’re looking for motivated, tenacious, and curious people who will work closely with engineering teams to ensure our infrastructure deployments are highly secure against our adversaries. As an FDSecE, you will embed directly throughout the lifecycle, working on-site and being hands-on to ensure the overall security of these deployments from design to production and through ongoing operations. This role is preferred to be based in Washington DC but may consider remote work. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Travel to and working from customer sites is required for this role. In this role, you will: Deeply embed with our most strategic public sector customers to implement and maintain robust security controls. Be a design and technical thought partner by leveraging security expertise on protective controls including access controls, authentication, encryption, network, and system security. Collaborate closely with teammates, cross-functional teams, customers, and service providers to achieve security and compliance goals. Ensure continuity of critical security and monitoring c
About the Team The Storage Infrastructure team builds and operates the storage foundation behind OpenAI’s most demanding workloads. We work directly with research to design storage systems for rapidly evolving experiments, while also powering production at scale. We own the platform end to end: backend systems, user-facing services and APIs, and the control planes that manage how data is placed, moved, and retained over time. Our stack spans cloud and in-house object stores across very different workload profiles, from GPU-attached systems to dedicated storage hardware. We also build the federation layer that unifies these backends behind a simple interface and routes each workload to the right storage solution. About the Role You will help build the storage platform that powers OpenAI’s research and production systems. This is a hands-on infrastructure role for engineers who want to work on deeply technical systems at scale and own them in production. You’ll work across object storage, cross-region data movement, lifecycle management, and the federation layer that provides a unified interface across multiple backends. Much of our stack runs on Kubernetes, and we primarily build services in Rust. In this role, you will: Build and operate storage services that underpin OpenAI’s research infrastructure Develop object storage systems across cloud and in-house environments Build systems for cross-region data movement, replication, and recovery Design lifecycle management capabilities that keep data durable, available, and cost-effective Evolve the federation layer that unifies multiple backend systems behind a simple interface Improve performance, reliability, and operational excellence across the platform Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads You might thrive in this role if you: Have experience building or operating distributed systems in production Have worked on storage infrastructure, object stores, dist
Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role As OpenAI continues to grow, we are looking for experienced, problem-solving engineers to ensure our systems scale. Our success depends on our ability to quickly iterate on products while also ensuring that they are performant and reliable. You will work in a deeply iterative, collaborative, fast-paced environment to bring our technology to millions of users around the world, and ensure it’s delivered with safety and reliability in mind. Successful candidates will play a crucial role in ensuring the reliability, scalability, and performance of our systems as we continue to expand. As a reliability expert, you will be at the forefront of maintaining and enhancing the stability, scalability, and performance of our rapidly evolving infrastructure. You will work closely with cross-functional teams, including software engineers, product managers, and data scientists, to build and maintain resilient systems that can handle our growing user base and workload. In this role, you will: Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands. Build and maintain the load, chaos and synthetic testing software leveraged by development teams to make the systems they design and operate more reliable. Build and maintain automation tools to streamline repetitive tasks and improve system reliability. Build and maintain the platform for CPU/storage, GPU, and network lifecycle management to drive efficiency, accountability and support dynamic optimization of our resources. Implement fault-tolerant and resilient
About the Role The Engineering Acceleration Delivery / Continuous Deployment team builds and operates the systems that safely ship OpenAI’s infrastructure and product code to production. We own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across OpenAI to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale infrastructure orchestration. In This Role, You Will Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks. Build systems that support agent-assisted or autonomous deployment workflows using modern AI tooling. Technologies commonly used in this environment include: Kubernetes for large-scale container orchestration and runtime infrastructure Python and FastAPI for internal services Terraform for infrastructure as code GitOps-based deployment workflows (e.g., ArgoCD, Flux, or similar systems) Buildkite for CI orchestration You may be a strong fit if you: Have worked with Kubernetes-based deployment systems at scale Have experience building or operating continuous deployment platforms Are familiar with GitOps tooling such as
Get new terraform jobs by email
Daily job updates · Unsubscribe anytime