ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently. You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. EXAMPLE INITIATIVES You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services. Building AI-assisted tooling for incident triage and response. RESPONSIBILITIES Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking. Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code. Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution. Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations. Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, con
Jobiba hiring network
Operations Support Senior Analyst Jobs
10,000 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current operations support senior analyst jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Us What if your work could drive change in a globally established industry, shaping processes that touch every corner of the world? At Forto, we are at the forefront of change, harnessing the power of AI to revolutionise logistics. We want to reinvent digital supply chains to be transparent, frictionless and sustainable. From day one, our mission has been to simplify global trade – creating a seamless and efficient logistics process. Your role & Mission As a Senior Backend Engineer in our Shipment team, you will build the "brain" behind our logistics operations. You’ll maintain and evolve a sophisticated event-driven architecture designed to answer one crucial question: "How do we move this shipment most efficiently?" From contract rate management to automated shipment nomination, you will lead the charge in creating reliable, data-heavy systems that optimize our global portfolio. You aren’t just writing code; you’re building a revenue-driving engine. Working alongside industry experts, you will mentor your team to deliver robust services that enable NVOCCs to thrive. If you are passionate about leading teams to solve massive logistical challenges, we want to hear from you. What you will do Design and implement robust backend services for capturing, storing, and validating procurement contract rates and allotment volumes, ensuring 100% data integrity for our supply chain logic. Lead the evolution of our event-driven microservices to process high-velocity data, enabling real-time updates on contract utilization and space availability. Refactor and optimize existing systems responsible for shipment-to-contract matching to ensure low-latency performance and high maintainability. Build and iterate on the logic that answers "Which contract should we use?" by automating the nomination process based on cost, volume commitments, and carrier performance. Streamline the ingestion and management of complex carrier agreements, moving from manual entry toward a fully a
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. As a Sr. AI GTM Engineer on Vanta's GTM Engineering team, you will design and ship internal AI products and processes that completely transform how our go-to-market teams operate and win. You'll embed with Sales, Customer Success, and Revenue Operations to build novel applications and platforms that accelerate pipeline generation, improve win rates, and scale how Vanta engages customers. You will also own end-to-end outcomes, move quickly from concept to production, and directly shape how customers experience Vanta in the field. This role is ideal for AI-pilled builders who want to be close to users, solve complex business problems with elegant technical solutions, and define entirely new categories of enterprise software. Vanta's GTM Engineering team operates as an internal incubator: applying AI at scale to revolutionize customer engagement. We build high-impact tools that reshape conversations, learn from every customer interaction, and demonstrate the value of our platform in real-world scenarios. This team sits at the intersection of engineering and GTM, shipping solutions that range from rapid prototypes to production-grade systems. What you'll do as a Senior AI GTM Engineer at Vanta: Ideate and build products that transform how Vanta generates new business, interfaces with prospects and customers, and grows existing customer relationships Own the full product lifecycle: prototype, iterate, ship, and maintain solutions that solve real-world customer and internal workflow problems Partner closely with Vanta’s CRO and executive leadership teams to architect and execute plans that bring Vanta to the forefront of GTM orgs lev
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role Supabase manages millions of Postgres instances and is growing. We have strong teams across observability, release engineering, and incident management — and we're concentrating our reliability efforts into a dedicated SRE practice that ties the discipline together across the platform. You'll be embedded within Service Operations, and your primary job is to make every engineering team more reliable — not by owning their infrastructure, but by establishing the practices, frameworks, and feedback loops that let them own reliability themselves. You'll work across the org: sometimes setting the standard, sometimes pair-programming a fix, sometimes helping a team define their error budget, sometimes telling them it's exhausted. This role is ideal for someone who has a strong vision for how SRE should work and thrives in async, fast-paced environments where influence matters more than authority. What You'll Own Partner with service teams to define meaningful SLIs and SLOs grounded in customer experience, and build the error budget policies that turn them into engineering decisions Own and evolve the Operational Readiness Review (ORR) process — conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation Strengthen the incident-to-improvement pipeline: connecting postmortem findings to operational readiness gaps, identifying repeat failure patterns, and driving systemic fixes Act as the reliability expert teams pull in for architecture reviews, failure mode analysis, dependency mapping, and resilience design Identify and quantify operational toil across the org, and build or advocate for automation that eliminates it
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We're looking for a Release Engineer (SRE) to join our Release Engineering team (part of EngOps) — a production-operations expert who brings an SRE mindset to how Supabase ships and runs, making deploys safe, observable, and recoverable at scale. Release Engineering's scope has grown well beyond build-and-ship: we increasingly own the operational reliability of the systems that deploy and run Supabase. In this role you'll treat our deployment pipelines, pre-production signal, and the control plane itself as production systems — with SLOs, error budgets, and on-call ownership — and you'll be the person teams lean on when reliability is on the line. This is not a "gatekeeper" role. You'll make the reliable path the easy path: standardising how we deploy, instrumenting what we ship, and ensuring that when something breaks, we detect it quickly and recover quickly. What You'll Be Responsible For In this role, you'll: Own the reliability of Supabase's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies) Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load Participate in on-call, lead blameless po
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We’re looking for an EngProd Engineer to join our Engineering Operations team and own the engineering experience from local setup to production deploy. You’ll build and integrate the tools our engineers rely on every day: local development, code search, testing, code review, CI, AI-assisted workflows, and the metrics that show where time is being lost. You’ll work closely with other engineering teams at Supabase to eliminate friction, optimize cycle times, and make reliably shipping to production effortless. This role is ideal for someone who thrives in async, fast-paced environments and is excited about building developer tools that engineers love and that empower them to do their best work. What You’ll Be Responsible for In this role, you’ll: Improve local development and the daily workflow Own local development environment stability; eliminate “works on my machine,” keep environments reproducible across teams, and make onboarding and setup self-service for every engineer Improve the daily editing experience, including IDE and editor tooling, extensions, and code search, indexing, and navigation that stays fast as the codebase grows Build paved roads: project scaffolding, templates, and golden-path workflows that make the right thing the easy thing Speed up build, test, and CI Improve our testing tooling and make sure it works well for engineers, including test orchestration, ephemeral end-to-end environments, and techniques like mutation and chaos testing that catch what ordinary tests miss Profile build and test performance, optimize CI pipelines, parallelize and shard tests, cache aggressively, and hunt down the bottlenecks that leave engineers waiting Integrate and optimize
About Pinecone Pinecone is the leading vector database for building accurate and performant AI applications at scale in production. Pinecone's mission is to make AI knowledgeable. More than 5000 customers across various industries have shipped AI applications faster and more confidently with Pinecone's developer-friendly technology. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: As a Business Development Representative, you’ll be the primary catalyst for user activation and expansion. You will work at the critical intersection of our Product-Led Growth (PLG) motion and our high-growth revenue engine. Your core mission is to identify high-potential users within our self-serve funnel and proactively convert them into qualified sales opportunities. You are part researcher and part automation hacker. You will partner closely with our Revenue Operations team to combine creative outreach with GTM (Go-To-Market) engineering workflows (using tools like Clay, n8n, Zapier, etc.) to deliver personalized, data-backed touchpoints at scale. You'll collaborate closely with Sales, Customer Success, and Solutions Engineering to accelerate our PLG motion, turning a robust sign ups & usage funnel into a predictable revenue engine and setting a new standard for AI-native GTM. Responsibilities: Drive Sales Opportunities: Proactively engage users exhibiting high-potential usage signals (e.g., specific API patterns, advanced feature experimentation, new workspace invites) to guide them toward opportunities for the sales team. Identify & Convert: Execute highly personalized, data-driven outreach sequences to connect with AI practitioners, technical leads, and business decision-makers. Your goal is to understand their platform behaviors and signals in order to generate qualified sales leads. Surface PQLs: Partner with RevOps and Growth to analyze product telemetry and surf
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to design and implement robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability and performance. You will: Design and Implement Observability Solutions : Develop comprehensive monitoring and alerting systems using modern observability tools. Create dashboards and metrics that provide real-time visibility into system health and performance. Implement logging strategies that enable quick problem identification and resolution. Drive Automation and Infrastructure as Code : Architect and implement infrastructure automation solutions using tools like Terraform, Ansible, or Pulumi. Design and maintain CI/CD pipelines that enable reliable and consistent deployments. Create self-healing systems that can automatically respond to common failure scenarios. Establish SLOs and SLIs : Work with product and engineering teams to define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to track and report on these metrics, ensuring we maintain high reliability standards while balancing innovation speed. Incident Management and Response : Lead incident response efforts, conducting thorough post-morte
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. As a GTM & G+A Sourcer , you won't just be filling seats; you will be the primary architect of the talent pipeline that builds the future of our business, operations, and growth teams. We are seeking a Fixed Term Full Time employee for a 6 - 12 month term . The Mission At Replit, we don’t look for average talent; we look for "10x" builders and creators. Your mission is to find the needles in the haystack—high-impact professionals across Go-to-Market (Sales, Marketing, Growth) and G+A (Finance, People, Ops) who are product-minded and excited by the challenge of democratizing coding. Key Responsibilities Strategic Pipeline Building: Identify and engage top-tier talent for GTM functions and G+A roles. Creative Outreach: Utilize AI tools to streamline workflows while crafting highly personalized, compelling outreach that resonates with the Replit ethos. Deep Research: Map out talent across LinkedIn, X (Twitter), competitive startups, and niche professional communities to uncover "undiscovered" talent. Full Lifecycle Partnership: Collaborate closely with Recruiters and Business Leads to calibrate profiles and refine search strategies in real-time. Data-Driven Iteration: Track conversion rates and leverage data to pivot sourcing strategies as needed. Who You Are Experience: 3+ years of sourcing experience with a focus on GTM or G+A roles, ideally within a high-growth startup environment. Business Acumen: You understand the AI/SaaS landscape and know what excellence looks like in non-technical candidates within our specific culture. Persistence: You possess a "hunter" mentality and enjoy the challenge of identifying talent for highly specific, specialized roles. High Agency: You are self-directed, comfortable operating wi
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. Replit is the fastest way to turn ideas into software. With Replit Agent, anyone can build and ship a real application in natural language, without setting up a single tool. We've grown explosively, from $2M to $400M+ revenue in under a year, with a global community of 60+ million users, and Replit Agent is now one of the fastest-growing products in history. That same wave is now hitting the enterprise. Companies like Visa, Ernst & Young and Comcast are rolling Replit out to thousands of employees so finance teams, operations leads, and product managers can build without waiting on an engineering backlog. When a company puts Replit in the hands of thousands of employees, they're trusting us with their data and their workflows. That trust has to be earned: SSO and SCIM that just work, governance their security team will sign off on, and real control over where their apps run and how their data is handled. Getting this right is what turns a pilot into a company-wide standard, and it's one of the biggest levers on Replit's next phase of growth. About the Role We're hiring a Staff Product Manager to own Replit's enterprise product end to end. We've built one of the best coding agents in the world. The work now is making it just as powerful inside a big company, where putting it in front of every non-technical employee unlocks enormous leverage. That means extending what Agent can do in an enterprise context and building the platform that lets large organizations adopt it safely and at scale. This is a role with unusual leverage. We're still early in our enterprise journey but the traction is remarkable, which means most of the platform is still to be built and you'll have an outsized hand in shaping it. You'll decide w
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role: Join our Infrastructure Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Infrastructure Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff Infrastructure Engineers who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios. Optimize Performance and Infrastructure: Collaborate with core infrastructure and product teams to performance tune and optimize our cloud deployments (Kubernetes, Docker, GCP). Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions. Elevate Developer Experience: Design and implement improvements to our build, test, and deployment systems to make software delivery faster, safer,
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering (SRE) team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust observability solutions, lead incident response, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Architect and Implement Observability: Design, build, and lead the implementation of comprehensive monitoring, logging, and tracing solutions. Create dashboards and metrics that provide real-time visibility into system health and performance, enabling proactive issue detection. Define and Drive Reliability Standards: Work with product and engineering teams to define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to monitor and report on these metrics, holding teams accountable and ensuring we maintain high reliability standards while balancing innovation speed. Lead Incident Management and Response: Act as a senior leader during high-impact incidents, guiding the team to rapid resolution
What you'll do Run human-subjects recording sessions end-to-end: participant prep and fitting, sensor configuration and calibration, stimulus delivery, real-time signal monitoring, and session documentation. Own daily system QC for a one-of-a-kind sensing instrument: baseline noise recordings, per-sensor health checks, log review, and escalation of anomalies before they touch a dataset. Own participant operations: recruitment coordination, scheduling, screening, consent, and compliant handling of participant records. Stand up and maintain the operational backbone of the program: SOPs for acquisition, QC, and participant workflows — documentation that survives you. Handle first-pass data operations: file conversion, organization, metadata, and quality review in Python. Partner daily with the program lead to refine acquisition protocols and improve system performance over time. What we're looking for Hands-on experience acquiring physiological or imaging data from human participants in a clinical or research setting Obsessive consistency and attention to detail: you notice when something is off and you don't let it slide. Professional, patient, and calm with research participants — sessions succeed or fail on how people feel in the room. Basic scripting ability (Python and/or shell) for QC and data-handling tasks, or clear aptitude and motivation to build it. High ownership of the unglamorous parts: scheduling, documentation, tidy data hygiene. Discretion — parts of this program are not yet public. We'll walk you through the specifics, including the exact technology you'd be operating, in the first conversation. Useful experience Controlled or low-noise recording environments and highly sensitive instrumentation. Physiological signal-processing tools in Python. Human-subjects research administration (IRB protocols, consent workflows). Research studies involving structured tasks or sensory protocols with human participants.
We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, incident.io, Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different s
About the Team At Trendyol Core Commerce, we build innovative, data-driven strategies that power sustainable growth and global expansion. From seller experience to new market launches, we turn insights into action—fast. Our cross-functional teams shape the future of commerce with bold ideas, real-time impact, and a deep sense of ownership. In a fast-paced, collaborative environment, we grow together — as individuals and as a team. As an Category Intern, you’ll step into the dynamic world of e-commerce, supporting our seller partners and contributing to their success on our platform. This hands-on internship within the Category Teams gives you a unique chance to gain real-world experience in seller operations and performance management. You’ll apply your data-driven mindset and strong communication skills to onboard new sellers, expand their product selection, and analyze performance data to deliver valuable insights.
Get new operations support senior analyst jobs by email
Daily job updates · Unsubscribe anytime