As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental research problems and collaborating with Datadog's product and engineering teams to translate research advances into products. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Conduct research in generative AI and machine learning, building specialized foundation models and trained agents for observability Train multimodal models on large-scale, diverse telemetry data (metrics, logs, traces, topology, events) using distributed training infrastructure Design and build simulated environments and RL training loops for on-policy agent training and evaluation Collaborate with cross-functional teams (Product, Engineering) to integrate capabilities like multimodal world modeling and autonomous agents into Datadog's products Stay at the forefront of foundation models, world models, and RL-based agent research Contribute to r
Jobiba hiring network
Security Incident Response Engineer Jobs
3,397 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current security incident response engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental research problems and collaborating with Datadog's product and engineering teams to translate research advances into products. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Conduct research in generative AI and machine learning, building specialized foundation models and trained agents for observability Train multimodal models on large-scale, diverse telemetry data (metrics, logs, traces, topology, events) using distributed training infrastructure Design and build simulated environments and RL training loops for on-policy agent training and evaluation Collaborate with cross-functional teams (Product, Engineering) to integrate capabilities like multimodal world modeling and autonomous agents into Datadog's products Stay at the forefront of foundation models, world models, and RL-based agent research Contribute to r
NVIDIA pioneers computer graphics, gaming, AI, and accelerated computing. We are looking for a Technical Platform Operations Lead to join our team and play an important role in scaling Sales AI applications and platforms. This position offers the opportunity to shape how these solutions operate after launch and help ensure they remain reliable, secure, well governed, widely adopted, and continuously improved. You will collaborate with Sales, Product, Engineering, Data, Security, and IT teams to strengthen platform health, improve the user experience, and increase business impact. What you’ll be doing: Lead end-to-end post-launch operations for Sales AI applications, including availability, performance, support readiness, releases, upgrades, and lifecycle planning. Develop effective processes for incident response, problem management, changes, and issue resolution. Coordinate timely recovery and lasting improvements. Analyze service-level indicators and objectives, adoption metrics, dashboards, alerts, and user feedback to identify risks, performance degradation, and usage gaps. Collaborate with partner teams to translate operational signals and user needs into prioritized improvements and roadmap inputs. Improve adoption and business value through usage analytics, enablement, feedback loops, and user experience enhancements. Establish governance practices for security, access controls, compliance, documentation, and platform support. Develop automation, observability, and self-service capabilities that simplify operations and reduce repetitive work and recurring incidents. Prepare new AI capabilities and releases for production with runbooks, monitoring, rollback plans, support models, and partner enablement. What we need to see: 8+ years of experience in technical operations, pl
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: The role is responsible for leading and overseeing end-to-end cloud operations, ensuring the availability, reliability, security, performance, and resilience of cloud platforms and services. It manages service monitoring, incident response, and incident resolution for production applications and cloud infrastructure, while ensuring operations teams are skilled and enabled to execute cloud-related requests with speed and diligence. The role also provides governance over a large third-party managed services organisation, ensuring delivery against agreed KPIs, SLAs, and operational commitments through effective service reviews, metrics reporting, performance management, and continuous service improvement. What you'll be doing: Responsible for overseeing cloud operations, driving operational excellence, and improving the operational landscape through automation and AI-driven solutions delivered by internal resources and third-party partnerships. Product: Work with product and engineering teams to define operational support patterns for each cloud product. Collaborate with business, archi
NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services. What you’ll be doing: Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale. Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation. Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data. Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems. Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation. What we need to see: BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services. Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges. <l
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Head of IT at Baseten, you will build, scale, and secure our internal technology function to support our rapid growth. Reporting to our Chief Information Security Officer, you will lead and mentor a team of 5+ IT engineers, leading the charge to transition Baseten from startup-era IT to a highly automated, enterprise-ready IT organization. You will take full ownership of corporate IT infrastructure, Helpdesk operations, corporate identity management, device lifecycles, and vendor procurement. As we scale to support the world’s most dynamic AI companies, you will ensure our internal systems scale seamlessly with our headcount, providing a secure, frictionless, and world-class technology experience for all Baseten employees. RESPONSIBILITIES Team Leadership: Manage, mentor, and grow a team of IT engineers, fostering a high-performance culture focused on technical excellence and end-user satisfaction. Helpdesk Operational Excellence: Build a fast-response support function by establishing clear response SLAs, tracking employee satisfaction metrics, and formalizing on-call and incident response processes. Zero-Touch Automation: Architect and implement automated employee onboarding, offboarding, and role-based access changes through deep integrations across HRIS, MDM, and IAM systems. SaaS Management & Procurement: Establish comprehensive SaaS management processes to eliminate shadow IT, automate access w
About the team OpenAI’s mission is to build safe artificial general intelligence (AGI) which benefits all of humanity. This long-term undertaking brings the world’s best scientists, engineers, and business professionals into one lab together to accomplish this. In pursuit of this mission, our Go To Market (GTM) team is responsible for helping customers learn how to leverage and deploy our highly capable AI products across their business. The Cybersecurity specialist sales team partners with Account Directors, Technical Success, Marketing, and Partnerships to drive cybersecurity adoption that help bring AI to as many users as possible. About the role Our Sales team has a unique mission to help cybersecurity customers understand the deep impact that highly capable AI models can bring to their businesses, operations, employees, and customers. This role is a mixture of technical understanding, industry expertise, vision, partnership, and value-driven strategy. As an Account Director focused on Cybersecurity, you will own executive-level relationships with leading cybersecurity firms and help them safely and effectively deploy OpenAI’s technology across their organizations. You’ll work with customers to identify and scale high-impact use cases across areas such as security operations, threat intelligence, risk management, incident response, vulnerability management, workforce enablement, customer support, and enterprise knowledge management. You’ll be a key driver of opportunities through the entire sales cycle, from pipeline generation to closure and successful deployment. You’ll work with researchers, engineers, and solution strategists to help customers transform their operations and evolve the cybersecurity industry with AI. This role is based in San Francisco, Seattle or New York. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. We are open to US-based remote candidates. In this role, you’ll: Support Accou
AI/ML – Investment Services A Career with Point72's AI/ML – Investment Services Team The AI/ML – Investment Services team at Point72 spearheads the development of cutting-edge AI solutions that seek to transform our business processes and enhance enterprise intelligence. The team aims to bridge the gap between business challenges and technological innovation, collaborating with stakeholders across the firm and leveraging expertise in generative AI, data engineering, and machine learning. WHAT YOU'LL DO Build and scale core backend services and platforms that power generative AI applications and data infrastructure used across the firm’s investment workflows Design and implement high-throughput, low-latency data pipelines to ingest, normalize, and serve both structured and unstructured data Develop robust APIs and microservices to support model inference, feature serving, and downstream applications Integrate generative AI tools and model-serving workflows into production, including embedding stores, retrieval components, and fine-tuning pipelines Optimize system performance, cost, and reliability through profiling, capacity planning, and architectural improvements Implement automated testing, continuous delivery pipelines, monitoring, and incident response practices to maintain production health Partner with data scientists, AI engineers, product owners, and operations to translate models and prototypes into scalable, production-grade solutions Mentor engineers, lead code reviews, and establish engineering best practices for maintainability, security, and observability Own end-to-end delivery, operational runbooks, and metrics-driven measurement of feature impact and system reliability WHAT'S REQUIRED Bachelor’s degree in computer science, software engineering, or a related technical field Minimum 5+ years of professional experience building backend systems and production services Demonstrated experience designing and operating large-scale data engineering pipelines
We are seeking an experienced IT/Lab Manager to lead the planning, deployment, and operations of our physical lab environment and IT systems. This role will focus on building and maintaining scalable, reliable, and secure environments to support engineering teams involved in research, quality assurance, validation, and related activities. It will also support internal collaborators. You will have an outstanding opportunity to drive innovation in a multidimensional, technology-focused company that is crafting the future of data-center and lab technologies. If you bring perfection and creative thinking while solving issues as they arise, and enjoy working with distributed teams – your place is with us! What You’ll Be Doing: Own day-to-day operations, planning, and roadmap for the engineering lab and IT infrastructure (servers, storage, networking, and related services). Lead and mentor an IT/Lab team, driving guidelines, standards, and a culture of ownership, partnership, and continuous improvement. Collaborate closely with R&D, QE, Verification, and other engineering teams to design, provision, and maintain environments that meet their performance, reliability, and security needs. Lead all aspects of running data center and lab operations, including rack layout, cabling, power and cooling, hardware lifecycle, and resource availability. Lead procurement and vendor management for hardware, software, and services, including evaluation, negotiation, and ongoing relationship management. Implement and maintain automation for system provisioning, configuration, and operations using tools such as shell/Perl/Ansible. Design and maintain monitoring, logging, and alerting for servers, network, and storage systems to ensure high availability and rapid incident response. Investigate and resolve sophisticated infrastructure issues across OS, networking, storage, virtualization, and appli
Technical Lead - Backend About Us: Paytm is India's leading mobile payments and financial services distribution company. Pioneer of the mobile QR payments revolution in India, Paytm builds technologies that help small businesses with payments and commerce. Paytm’s mission is to serve half a billion Indians and bring them to the mainstream economy with the help of technology. About the role: We are looking for a Tech Lead for the Gold Tech team to drive technical direction, architecture, and delivery of Paytm Digital Gold — a high-scale bullion platform serving consumer and merchant use cases at lakhs+ rpm. This role combines hands-on engineering with technical leadership — you will architect solutions, unblock the team, own critical money-path reliability, and partner with Product and Engineering leadership on roadmap and execution. We are building reliable, high-availability financial products and need a Tech Lead who can own modules end-to-end, raise the bar on code quality, and mentor junior engineers. Key Responsibilities: ● Define and evolve technical architecture for Gold services — scalability, reliability, security, and maintainability. ● Lead design and implementation of high-availability, high-volume transactional systems. ● Own end-to-end delivery of major initiatives — breakdown, estimation, risk management, and production rollout. ● Set engineering standards: code quality, review practices, testing strategy, observability, and incident response. ● Drive cross-service integration design. ● Guide the team on Kafka event design, scheduler orchestration, CDC pipelines, and data consistency patterns. ● Mentor and grow engineers (SSE and below); conduct reviews, pair on complex problems, and build team capability. ● Contribute hands-on to critical modules; unblock the team on complex bugs, performance issues, and production fires. ● Represent Gold Tech in architecture reviews, tech debt prioritization, and platform-wide initiatives. Skills Required: ●
Who We Are At Justworks, you’ll enjoy a welcoming and casual environment, great benefits, wellness program offerings, company retreats, and the ability to interact with and learn from leaders in the startup community. We work hard and care about our most prized asset - our people. We’re helping businesses get off the ground by enabling them to focus on running their business. We solve HR issues. We’re data-driven and never stop iterating. If you’d like to work in a supportive, entrepreneurial environment, are interested in building something meaningful and having fun while doing it, we’d love to hear from you. We're united by shared goals and shared motivations at Justworks. These are best summed up in our company values, which are reflected in our product and in our team. Our Values If this sounds like you, you’ll fit right in. Who You Are Justworks is looking for an experienced security engineer skilled in detection and response, who can help enhance and mature Justworks’ Security. As a Senior Detection Engineer, you’ll design, build, and maintain the detection logic that powers our platform, conduct proactive threat hunting, and drive continuous improvements across our detection and incident handling workflows. You’ll collaborate closely with IT, Engineering, Platform, and other members of the Security team to identify attacker behaviors, build high‑fidelity detections, and strengthen our defenses. You’ll also play a key role in designing and conducting table‑top exercises, improving processes, and building automation that reduces friction and accelerates response. You’ll help explore how AI can enhance detection, hunting, and operational efficiency. Your Success Profile What You Will Work On Build, tune, and deploy high‑quality detections across our platform Develop and refine detections using telemetry from EDR, threat intel, endpoint & cloud posture platforms and native AWS cloud services Conduct proactive threat hunting to uncover threat actor behaviors a
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: Responsible for leading the Cloud Automation Engineering function. Primary focus will be leading a team of other engineers in designing and implementing automation solutions to improve customer experience and increase productivity in our cloud estates. Responsible for maintaining and delivering automation solutions through infrastructure as code, ensuring security best practice, evangelising automation practice and tools, and supporting customer needs, both internal and external. What you'll be doing: Identify opportunities for improvement and automation of operations Design, build, test and implement use cases to drive automation adoption and improve operational efficiency Work closely with the IT Operations team to develop automated incident detection and response mechanisms. Implement proactive monitoring and alerting systems to quickly respond to and resolve critical issues, minimizing downtime and service disruptions Responsible for driving CSI initiatives to improve operations (processes/tools) working with various stakeholders Responsible for providing feedback at various leve
Position: Engineering Manager - Database Job Location: Noida Role Overview We are seeking a Database Engineering Manager (Individual Contributor) with deep expertise in MySQL and strong working knowledge of MongoDB, PostgreSQL, and Cassandra. This role combines hands-on database administration and optimization with strategic ownership of database reliability, automation, and cloud adoption. The candidate will lead by example—driving technical excellence, influencing best practices, and partnering cross-functionally with DevOps, SRE, and product engineering teams to deliver highly available, secure, and scalable database platforms. Key Responsibilities 1. End-to-End Ownership of MySQL databases in production & staging—availability, performance, and reliability. 2. Architect, manage, and support MongoDB, PostgreSQL, and Cassandra clusters for scale and resilience. 3. Define and enforce backup, recovery, HA, and DR strategies across all critical database platforms. 4. Drive database performance engineering—tuning queries, optimizing schemas, indexing, and partitioning for high-volume workloads. 5. Own replication, clustering, and failover architectures ensuring business continuity. 6. Champion automation & AI-driven operations—design self-healing scripts, predictive scaling, and proactive monitoring solutions. Collaborate with Cloud/DevOps teams on AWS database services (RDS, Aurora, DynamoDB, EC2, S3) to optimize cost, security, and performance. 7. Establish monitoring dashboards & alerting mechanisms for slow queries, replication lag, deadlocks, and capacity planning. Ensure compliance & security standards—encryption, auditing, and regulatory requirements. 8. Lead incident management & on-call rotations, ensuring rapid response and minimal MTTR. 9. Act as a strategic technical partner, contributing to database roadmaps, automation strategy, and adoption of AI-driven DBA practices. Required Skills & Experience 1. 6–10 years of p
Opportunity Overview: We’re looking for a Manager, Platform Engineering that can lead and grow a high-performing engineering team focused on Developer Experience, DevOps, SRE, and Quality. You will own the systems and processes that enable teams to build, test, release, and operate software with high velocity and reliability, driving engineering efficiency and operational excellence across the organization. What you’ll do: Lead a fast-paced, autonomous team of engineers focused on platform engineering, developer experience, DevOps, SRE, and quality engineering Own and drive the internal developer platform strategy and roadmap, improving how engineering teams build, test, deploy, and operate services Create transparency into engineering efficiency and system health through meaningful metrics across delivery, reliability, and quality Enable teams to move faster by improving CI CD pipelines, environments, tooling, and overall developer workflows Provide technical leadership across platform, infrastructure, and reliability, helping teams build scalable and resilient systems Ensure strong engineering practices across release processes, testing, quality, reliability, and security Define and enforce release guardrails, validation standards, and rollback mechanisms to improve production safety Improve environment stability and consistency across development, QA, and pre production environments Drive test strategy and automation maturity to improve overall product quality and confidence in releases Define and implement observability, monitoring, and alerting standards across systems Improve incident detection, response, and RCA practices, ensuring learnings translate into platform and system improvements Drive cloud infrastructure best practices across AWS, containers, and infrastructure as code Foster a culture of ownership, reliability, and continuous improvement within the team Provide innovative solutions for attracting, developing, and retaining top engineering talent I
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Security Engineer on the Detection and Response (D&R) team at Roblox, you’ll protect our user community alongside the underlying platform infrastructure. You’ll design high-fidelity detections, engineer security data platforms, and respond alongside the team during incidents. This is a hybrid in-office role in San Mateo. You Will: Deliver robust D&R capabilities: Engineer high-fidelity detections end-to-end. Lead partners through threat modeling and logging, to deploying actionable alerts, while keeping false positives low. Build security data pipelines: Develop security data pipelines and actively contribute to internal software and data platforms, collaborating across engineering teams. Ensure service reliability: Participate in an on-call rotation to keep detection and response services healthy. Embody security culture: Serve as a trusted security partner across Roblox, helping protect our community and enterprise while fostering a culture grounded in trust, ownership, and shared responsibility. You Have: 3+ years of experience in Security Data Engineering: You have built services that are efficient, reliable, and scalable using programming languages like Golang or Py
Get new security incident response engineer jobs by email
Daily job updates · Unsubscribe anytime