NVIDIA is transforming how the world uses AI, cloud, and accelerated computing, and trust is at the center of that mission. Our Attestation and Trust Services team builds the secure cloud services that show customers their NVIDIA platforms are healthy, resilient, and ready for their most important workloads. In this role, you help design and run services that sit at the intersection of hardware, security, and large-scale distributed systems. We partner closely with security, silicon, platform, and cloud teams to bring new ideas into reliable production services that people rely on every day. We care about building systems that last, supporting each other, and creating space for learning and experimentation. If you enjoy solving complex problems, keeping services running smoothly, and collaborating with teammates from many disciplines, we would love to talk with you! What you’ll be doing: Your main focus will be on building and managing our core attestation cloud services. Day-to-day responsibilities include crafting APIs and integrations, boosting reliability, and working alongside NVIDIA teams to convert hardware trust mechanisms and standards into production-ready solutions. You will contribute significantly to shaping how customers verify that NVIDIA platforms are secure and prepared for their workloads. Crafting and evolving attestation cloud services, APIs, and SDK/CLI integration points that confirm the integrity of NVIDIA platforms across data center, AI, networking, and partner environments. Improving reliability and operational maturity through SLOs/SLIs, alerting, runbooks, incident response, and safe rollout practices. Crafting resilient service behavior that handles dependency failures, caching challenges, regional issues, customer-side resilience needs, and graceful degradation. Architecting trust-material distribution for certificate status, re
Jobiba hiring network
Incident Commander Jobs
589 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current incident commander jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior software engineer Job Description Summary Overview: Provides support of applications software through programming, analysis, design, development and delivery of software solutions. Researches alternative technical solutions for changing business needs. Role: •Responsible for programming, testing, implementation, documentation, maintenance and support of systems application software in adherence with MasterCard standards, processes and best practices. •Develop high quality, secure, scalable software solutions based on technical requirements specifications and design artifacts within expected time and budget. •Research, create and evaluate technical solution alternatives for the business needs current and upcoming technologies and frameworks. •Perform feasibility studies, logic designs, detailed systems flowcharting, analysis of input-output flow, cost and time analysis. •Work with project team to meet scheduled due dates, while identifying emerging issues and recommending solutions for problems and independently perform assigned tasks, perform production incident management. Participate in on-call pager support rotation. •Document software programs per Software Development Best Practices. Follow MasterCard Quality Assurance and Quality Control processes. •Assist Seni
Work Flexibility: Onsite What you will do This position is responsible for environmental compliance, occupational safety, industrial hygiene, and support of sterilization operations while partnering closely with Operations, Quality, Engineering, and Maintenance teams. The role requires balancing operational needs with regulatory compliance, risk management, and employee safety. Own and lead site Environmental, Health & Safety programs and management systems. Establish and sustain a strong safety culture through visible leadership, coaching, employee engagement, and influence across key site functions. Maintain strong field presence through observations, safety walks, contractor oversight, and risk assessments to identify operational risks, escalate concerns, stop work when needed, and support safe and compliant operations. Serve as the site owner for environmental permitting and regulatory compliance, including permits, reporting, submissions, renewals, modifications, and agency interactions. Partner with Legal, Engineering, Operations, and Quality to assess process changes, sterilization validations, equipment modifications, facility expansions, and related EHS or regulatory impacts. Lead industrial hygiene and EO-related compliance programs, including exposure monitoring, respiratory protection, fit testing, occupational health surveillance, hazard communication, ergonomics, and employee notifications. Lead incident investigations, root cause analyses, corrective action programs, emergency response efforts, and required employee training. Conduct inspections, compliance assessments, audits, and analyze EHS metrics to drive continuous improvement and ensure timely closure of findings, actions, and regulatory commitments. What y
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Senior SOC Analyst to join our team . This is a remote role, reporting to the Global SOC Manager in the Enterprise Security department . As a key guardian of our infrastructure, you will monitor, detect, and analyze security incidents to protect our digital assets . You will be the first line of defense, ensuring that potential threats are mitigated promptly and effectively to maintain a secure environment . What you’ll do (Role Expectations) Monitor security alerts and events to identify potential threats and vulnerabilities Detect and analyze security incidents using multiple security tools Respond to security incidents promptly, following established procedures and protocols Conduct in-depth analysis of security events to determine the scope and impact of potential threats Perform phishing incident analysis Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: Own the service-management processes and quality for the M365 estate, ensuring consistent, well-governed operations and continuous improvement across the Kyndryl-delivered service. What you'll be doing: Define and maintain M365 service-management processes and standards. Govern Kyndryl process adherence, SLAs and service quality. Own service reporting, reviews and continuous improvement. Drive problem management and root-cause elimination. Support audits, compliance and governance. Coordinate day-to-day delivery with Kyndryl run teams across incident, request and problem management. Govern operational SLAs and KPIs and drive partner service improvement. Manage operational escalations and ensure low-risk, well-controlled change. What you'll need: Service management and process design for M365. ITIL and continual-service-improvement practice. Vendor governance and SLA reporting. Data analysis and reporting. Stakeholder management. Proven experience leading operational teams for the relevant platforms. ITIL / service-management experience and a vendor-coordination track record. Relevant
About the Role: As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams to build world-class machine learning inference platforms. These platforms power essential services like personalized recommendations, search, and content understanding across Tubi. A core responsibility of this team is developing and maintaining low-latency ML model serving systems that support Deep Learning, LLM, and Search models. This involves building self-service infrastructure and critical components such as the inference engine, feature store, vector store, and experimentation engine. You will improve the way we deploy and operate our services and even contribute to open-source projects. This role grants the architectural freedom to explore new frameworks, lead critical cross-functional projects, and transform the capabilities of our ML and Product teams. Responsibilities: Design and build scalable, high throughput, and low latency distributed systems using Scala Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art. Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary Mentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc. Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi. Your Background: Experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM b
About Wolt At Wolt, we create technology that brings joy, simplicity and earnings to the neighborhoods of the world. In 2014 we started with delivery of restaurant food. Now we’re building the delivery of (almost) everything and you’ll find us in over 500 cities in 30 countries around the world. In 2022 we joined forces with DoorDash and together we keep on dreaming big and expanding across the globe. Working at Wolt isn’t always easy, but it’s definitely exciting. Here you’ll learn more, build more, and ship more than in most other companies. You’ll be challenged a lot, but also have a lot of fun on the way. So, if you’re a self-starter with drive and entrepreneurial spirit, this could be the ride of your life. What you’ll do: Build and maintain high-throughput backend services using Go . Collaborate with product managers, designers, and frontend developers to ship features that support internal support agents across the globe. Design systems that are scalable , resilient , and easy to maintain. Lead and contribute to architectural discussions and technical decision-making. Write well-tested code and help the team maintain high code quality standards. Our humble expectations: 7+ years of professional software engineering experience, with a proven track record of building and scaling complex systems. 2+ years of production experience in Golang , with the ability to mentor others and drive best practices across the team. Strong hands-on experience with both SQL and NoSQL databases — especially Cassandra. Solid understanding of designing and operating low-latency, high-throughput distributed systems . Nice to have Background in Node.js or other backend languages. Familiarity with cloud infrastructure (AWS, GCP) and event-driven architectures. Previous on-call experience , with a pragmatic approach to reliability and incident management. What we value A product-oriented mindset — you think beyond the ticket, understand the “why” behind the work, and aim to create real
About Backblaze Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators. Together with our partners, we’re helping customers break free from the restrictive, overpriced legacy solutions that hold them back, and blaze forward with the full power of the open cloud in their hands. Founded in 2007, we scaled the business with less than $3 million in outside funding until 2021, when we did a traditional IPO on the Nasdaq stock exchange. Today, Backblaze generates over $136M ARR and is the leading specialized storage cloud, managing over three billion gigabytes of data storage for 500K+ customers in 175+ countries, including businesses, developers, IT professionals, and individuals. But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us. We’re seeking a Sr. Reliability Engineer ll (DBA) to join our team! About the Role We are seeking a Site Reliability Engineer (SRE) with a DBA (Database Administration) focus to help ensure the stability, scalability, and reliability of our production database systems - primarily Vitess (distributed MySQL) and Cassandra - alongside the rest of our services and infrastructure. This role operates within procedures and runbooks established by our senior DBA SREs, and focuses on building automation, maintaining observability, and supporting incident response to keep customer-facing systems performing at their best. The SRE will collaborate with engineering, product, and operations teams to embed reliability practices into day-to-day development and operations while contributing to tools and processes that improve efficiency and reduce manual effort Key Responsibilities Database Administration Operating and maintaining high-availability database systems — primarily Vitess (distributed MySQL) and Cassandra — against established architecture and runbooks. Op
Who are we? FalconX is a pioneering team of operators, investors, and builders committed to revolutionizing institutional access to the crypto markets. Operating at the intersection of traditional finance and cutting-edge technology, FalconX addresses the industry's foremost challenges: Navigating the digital asset market can be complex and fragmented, with limited products and services that support trading strategies, structures, and liquidity found in conventional financial markets. As a comprehensive solution for all digital asset strategies from start to scale, FalconX operates as the connective tissue empowering clients with seamless navigation through the ever- evolving cryptocurrency landscape. Responsibilities Be part of a trading systems engineering team, dedicated to building out the core trading platforms. Work closely with cross functional teams to improve the system reliability, scalability and security. Engage in and improve the quality supporting the platform. Build and manage systems, infrastructure and applications through automation. Provide operational support to internal teams working on the platform. Work on improvements to bring in high efficiency, reduce latency, deploy systems faster. Practice sustainable incident response and blameless postmortems. Together with your engineering team, you will share an on-call rotation and be an escalation contact for service incidents. Implement and maintain rigorous security best practices across all infrastructure, with a focus on minimizing attack surface and ensuring data integrity. Monitor system health and performance with a keen eye for identifying and resolving issues before they affect trading activity. Manage user queries and service requests (often requiring in depth analysis of the technical and/or business logic of our systems). Proactive approach to problem analysis and resolution of production incidents. Manage Issue tracking and prioritisation of day to day production incidents. Manage platf
Do work that matters. At AlertMedia , we help organizations protect their people, operations, and brand. Our modern Risk Intelligence and Response platform empowers teams to detect emerging threats, assess impact, and respond with confidence. We believe building resilience should be simpler—and it starts with bringing critical information and workflows together in one unified platform. Our core values drive us in our important mission of keeping people safe & informed: We’re humans not robots Customers always come first We work better together Simplicity is our strength Our reputation is priceless Hard work pays off As one of the fastest-growing software companies in the country, AlertMedia is looking for a Software Engineer II with full-stack experience. This role is on our Incident Response team. Customers use Incident Response during an active emergency — to declare an incident, assess who and what is affected, assign tasks, and coordinate the response across the organization. As an AI-forward company, we actively use modern AI tools such as Claude, Codex, and other emerging technologies to help our engineers work smarter, move faster, and build better products. Who you are: You thrive in a collaborative engineering environment that plays to everyone's strengths with a willingness to work across the entire stack. You design for people operating under pressure, and you think as carefully about the failure paths as the happy path. Ideally, you have JavaScript and Python experience (similar languages are great, as long as, you're willing to learn) but more importantly, you care about the quality of your work, the impactful product you help build, and the team you help build it with. You have AWS experience and are willing to learn from people with both more and less experience than you. What you will do: Design and d
Do work that matters. At AlertMedia , we help organizations protect their people, operations, and brand. Our modern Risk Intelligence and Response platform empowers teams to detect emerging threats, assess impact, and respond with confidence. We believe building resilience should be simpler—and it starts with bringing critical information and workflows together in one unified platform. Our core values drive us in our important mission of keeping people safe & informed: We’re humans not robots Customers always come first We work better together Simplicity is our strength Our reputation is priceless Hard work pays off As one of the fastest-growing software companies in the country, AlertMedia is looking for a Sr. Software Engineer with full-stack experience. This role is on our Incident Response team. Customers use Incident Response during an active emergency — to declare an incident, assess who and what is affected, assign tasks, and coordinate the response across the organization. As an AI-forward company, we actively use modern AI tools such as Claude, Codex, and other emerging technologies to help our engineers work smarter, move faster, and build better products. Who you are: You thrive in a collaborative engineering environment that plays to everyone's strengths with a willingness to work across the entire stack. You've developed scalable production web applications in Python/Django/Next.js/React. You design for people operating under pressure, and you think as carefully about the failure paths as the happy path. You care about the quality of your work, the impactful product you help build, and the team you help build it with. You have AWS (or other cloud) experience and are willing to learn from people with both more and less experience than you. You are a user of AI-assisted development tools that improve engineering workflows, code quality, and team productivity. What you
About Hexnode Hexnode, the Enterprise software division of Mitsogo Inc., was founded with a mission to simplify the way people work. Operating in over 100 countries, Hexnode UEM empowers organizations in diverse sectors. Fueling the transformation to a seamless ecosystem of connected tools, Hexnode is revolutionizing the enterprise software and cybersecurity landscape. Role Overview We are seeking a AWS Operations Specialist to manage and maintain our cloud infrastructure and device ecosystems. This is a highly operational, execution-focused role—not an architecture position. The ideal candidate has 2 to 4 years of experience executing infrastructure as code, monitoring environments, and following documented playbooks to keep our systems secure and resilient. Because this role handles secure environments, candidates must be US Citizens and capable of passing a comprehensive federal background check. Key Responsibilities Infrastructure Execution: Run, maintain, and execute existing Terraform and Ansible scripts to deploy and update infrastructure. GovCloud Monitoring: Actively monitor our AWS GovCloud dashboards, keeping a close eye on system health, performance metrics, and security baselines. Mobile Device Management: Manage Android Enterprise kiosk configurations, ensuring secure deployments and smooth device operations. Incident Response & Triage: Respond swiftly to operational alerts by strictly following our documented team playbooks. Escalation: Identify anomalies or issues that fall outside established, documented procedures and escalate them accurately to the engineering team. Required Qualifications & Profile Citizenship: Must be a US Citizen (required for GovCloud infrastructure management). Background: Must be able to successfully clear a rigorous federal background investigation. Experience: 2 to 4 years of hands-on experience in a technical operations, DevOps, or SysAdmin role. Technical Familiarity: Comfort executing/running Terraform and
About Hexnode Hexnode is a global leader in Unified Endpoint Management (UEM), trusted by over 100 countries and managing millions of devices worldwide. With a rapid pace of innovation, we have established ourselves as a dominant force across Apple, Windows, Android, macOS, Linux, and tvOS. Fuelling the transformation to a seamless ecosystem of connected tools, Hexnode is revolutionizing the enterprise software and cybersecurity landscape. Job Overview: In this role, you will not just oversee ticket resolution; you will lead the evolution of our support culture by implementing the Technology Acceptance Model (TAM) to ensure long-term product adoption and client success. As a linchpin of our 24/7 global operations, you will manage critical handovers and ensure our response times remain industry leading. Responsibilities: · Manage and mentor a shift of Technical Support Engineers, fostering a high-performance environment focused on technical growth and empathy. · Oversee the Shift Handover process to the incoming Lead, ensuring zero data loss and continuity of service for all high-priority issues. · Guarantee that all incoming tickets are triaged and dispatched within defined response times to meet and exceed global SLAs. · Conduct deep-dive analysis of incident patterns to identify root causes and provide strategic suggestions for incident management improvements. · Partner with Engineering and Product teams to function as the "Voice of the Customer," ensuring technical feedback directly influences the product roadmap. · Stay at the forefront of UEM trends and provide advanced training sessions for both high-value clients and internal technical stakeholders. Requirements: · 10+ years of total technical support experience, with at least 4–5 years in a formal leadership or supervisory role, preferably within a SaaS environment. · Mandatory, deep-seated expertise in Unified Endpoint Management (UEM) and its ecosystem. · Advanced Technical Exposure: Direct experie
Job Description: Data & Analytics Engineer (3-4 Years Experience) : 100% Remote Job Title: Data & Analytics Engineer Experience: 3-4 years Location: Remote Employment Type: Full-time Team: Data Engineering & Analytics About the Role Eltropy is a digital conversations platform for credit unions and community financial institutions in the US. The Data Engineering & Analytics team builds the AWS data pipelines and customer-facing dashboards that power analytics across the platform. We are looking for a Data & Analytics Engineer with 3-4 years of experience who can own dashboard delivery end to end along with the pipelines behind it. The ideal candidate learns fast, builds product context quickly, listens well, and collaborates effectively across product, engineering, DevOps, and customer-facing teams. Key Responsibilities Own dashboard changes end to end in QuickSight and ThoughtSpot - new metrics and filters, SPICE refresh management, internal-to-production promotion, and post-release validation. Build and maintain batch and streaming ETL pipelines on AWS using Glue (PySpark), S3, Redshift, and Airflow (MWAA) DAGs. Write and optimize Redshift SQL; debug query performance, connection contention, and data mismatches across sources. Support near-real-time ingestion (Kafka/MSK CDC → Glue Streaming → S3 → Redshift). Investigate customer-reported analytics discrepancies (Jira/support tickets), root-cause them in the data, and communicate findings clearly to support, product, and engineering. Set up and respond to pipeline monitoring - CloudWatch metrics and alarms, monitoring DAGs, refresh health - and participate in incident triage and RCA. Develop deep product knowledge: understand what each metric means to our credit union customers and translate product changes into data model and dashboard updates. Ensure data quality, validation, and consistency across systems. Required Skills & Qualifications 3-4 years of experience in data engineering and/or
Your wellbeing, our mission. Join a company shaping a healthier world. GET TO KNOW US At Wellhub we're revolutionizing workplace wellness. Our platform connects employees worldwide to the best partners for fitness, mindfulness, therapy, nutrition, and sleep—all in one simple subscription. Headquartered in NYC with team members in Europe, North America and South America, we’re on a mission to make every company a wellness company. We believe work should be fulfilling, inspiring, and balanced. Here, you’ll find a team that values wellbeing, collaboration, and different perspectives, where passion and creativity push boundaries to create real impact. Your contributions will help shape a healthier, more balanced world for you and millions of people globally. Join us in redefining the future of wellbeing! THE OPPORTUNITY We are hiring a Staff Security Engineer | AppSec to our Information Security team in Brazil! This is a Remote – Brazil position, meaning you can work from anywhere within the country. Please note that this role is only open to candidates in Brazil. The Information Security team is responsible for protecting our subscription-based product serving millions of users globally. As a Staff Security Engineer, you will own multiple security domains end-to-end — with your center of gravity in software security (secure SDLC, vulnerability management, threat modeling, pentesting, and red teaming) while reaching across incident response, threat intelligence, cloud security, and compliance as the team's mandate requires. You will become the organization's go-to authority for the hardest, cross-domain security trade-offs — the ones without an obvious owner. By connecting pentest findings, incident root causes, compliance requirements, and cloud misconfigurations into a unified risk strategy, you will shape baseline security standards, mentor engineering teams, and drive medium-to-large strategic initiatives that scale with our growth. YOUR IM
Get new incident commander jobs by email
Daily job updates · Unsubscribe anytime