Jobiba hiring network

Remote Remote Remote Remote Integration Reliability Engineer 2c Technical Operations Manager Jobs

15 active opportunities · Updated for September 2026

Market range: $131.3K – $199.3K/yr

Fresh results

15 shown

Explore current remote remote remote remote integration reliability engineer 2c technical operations manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for a Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the Infra team, we build the scalable, secure, and robust backbone that enables this vision. Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads. We are seeking a talented engineers with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale’s cloud platform. You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers. A snapshot of projects you may work on Design, build, and scale services that orches

REMOTEpythonawsazure
View job →
S
Smartsheet
📍 -REMOTE, USA-Full-timeRemoteFrom $1.7M/yr
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and automations that power Smartsheet’s core business functions across Finance, Sales/GTM, and People & Culture. Our team owns mission-critical systems and workflows that enable how the company hires, sells, bills, pays, reports, and scales. We operate at the intersection of software engineering, enterprise platforms, and business-critical data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Finance Systems team engineers and operates the platforms that support financial operations, including ERP, procurement, billing, and compliance. We work across configuration, extensibility, and integration to ensure systems are scalable, auditable, and resilient, treating code, configurations, and controls with the same rigor as software. As a Senior Software Engineer I (Finance Systems), you will lead the design, build, and operation of systems and workflows that directly support business execution at scale. You will own complex technical initiatives, partner with Product Managers and stakeholders on technical roadmaps, and mentor junior engineers. You will report into a Manager, Enterprise Systems, and can be based in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Systems Architecture & Optimization: Engineer and lead the end-to-end lifecycle—analysis, prioritization, and tec

REMOTEvueawsrest
View job →
GA
greenhouse,EnCharge AI
📍 Remote - US, CanadaFull-timeRemote$200K – $250K/yr
8 hrs ago

EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing. EnCharge’s robust and scalable next-generation in-memory computing technology provides orders-of-magnitude higher compute efficiency and density compared to today’s best-in-class solutions. The high-performance architecture is coupled with seamless software integration and will enable the immense potential of AI to be accessible in power, energy, and space constrained applications. EnCharge AI launched in 2022 and is led by veteran technologists with backgrounds in semiconductor design and AI systems. Lead DFT Engineer Job Description: Developing silicon for edge-to-cloud computing isn't just about speed; it’s about balancing high-performance data processing with extreme power efficiency and reliability in remote environments. As the Design for Test (DFT) Lead, you will be the architect of our testing strategy, ensuring our data center chips are flawlessly manufacturable and resilient enough for edge deployment. Key Responsibilities: Architectural Leadership: Define and implement the end-to-end DFT architecture for complex SoCs, including Hierarchical DFT, Scan compression, Boundary Scan and MBIST. Edge-Specific Reliability: Develop strategies for In-System Test (IST) and power-on self-test (POST) to ensure chip health in remote edge data centers. Implementation & Flow: Oversee scan insertion, ATPG (Stuck-at, Transition, Path Delay), and Memory /Logic BIST. Cross-Functional Synergy: Collaborate with Design, Physical Design, and Yield teams to ensure high test coverage while minimizing area overhead and power impact as well as timing analysis . Post-Silicon Validation: Lead the bring-up and debug phase on ATE (Automated Test Equipment) to root-cause silicon failures and optimize test time. Technical Requirements: Experience: 12+ years in DFT, with at least 2 years in a leadership or principal role. Bachelor’s degree in a related field. Tools:

REMOTEaisem
View job →
A
Anyscale
📍 RemoteFull-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: Optimizing performance of large-scale workloads on Ray Stability and stress testing infrastructure Improving fault tolerance (HA) As part of this role, you will: Leading cross-team projects while mentoring junior team members Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements

restmachine learningai
View job →
A
Anyscale
📍 RemoteFull-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: - Optimizing performance of large-scale workloads on Ray - Stability and stress testing infrastructure - Improving fault tolerance (HA) As part of this role, you will: Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements to Ray core Improve the testing process for Ray to make re

restmachine learningai
View job →
S
Supabase
📍 RemoteFull-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution, including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the team: Management API , written in TypeScript , Nest.js and for other JavaScript technologies, is a central part of every product in the Supabase stack. It allows all Supabase services and Supabase Studio to communicate with each other, programmatically manage your own Supabase projects and organisations, as well as integrate with 3rd party services like Vercel , Resend , or Lovable . We are seeking someone to help us maintain the existing API, expand OAuth applications capabilities, enhance reliability, and improve the overall public API experience. You will: Design, implement, and maintain both internal and public-facing APIs used across various Supabase products, including Studio, CLI, management APIs, and OAuth applications. Integrate with third-party platforms and partners, either by developing custom integrations or providing clear API points for them to connect with Supabase. Collaborate closely with various teams across Supabase (DevOps, Frontend, etc.) to ensure smooth integration and implementation of API functionality for the rest of the platform. Build and enhance testing, debugging, and monitoring tools to ensure public APIs' stability, reliability, and performance. Work with the dev-workflows team to enhance and improve the Branching experience, making it easier and more valuable for Supabase users. You have: 5+ years of experience in backend API development, with strong expertise in TypeScript and JavaScript (Node.js) and familiarity with modern tools and frameworks (e.g., Nest.js, Express, Vitest, Zod). Expertise in designing robust, scalable, and maintainable APIs, and experience with API versioning, pagination, and error handling best practices. Experience with OAuth

javascripttypescriptjava
View job →
S
Supabase
📍 RemoteFull-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the role We're looking for an engineer to drive the evolution of OrioleDB and its collaboration with the upstream PostgreSQL community. OrioleDB is a next-generation storage engine for PostgreSQL, and this role sits at the intersection of database internals, open source community work, and Supabase's managed Postgres platform. This role requires strong working overlap with Americas time zones due to production support and on-call responsibilities. What you'll work on New OrioleDB features. Design and implement new capabilities in OrioleDB — for example, native index access methods (GiST, GIN, HNSW for pgvector), disaster recovery tooling, and other storage-engine-level features that expand what OrioleDB can do. Stability. Strengthen OrioleDB's reliability through deeper test coverage, fault injection, crash and recovery testing, and improvements to CI infrastructure. Ensure regressions are caught early and that OrioleDB behaves predictably under stress, replication, and failure scenarios. Upstream collaboration. Contribute changes directly to PostgreSQL core. Part of OrioleDB lives as a patch on top of PostgreSQL — moving the right pieces upstream shrinks what we maintain ourselves and benefits the wider community. This happens through PostgreSQL's open development process: the pgsql-hackers mailing list, public code review, and commitfests. Supabase integration. Work with Supabase's Postgres team to ensure OrioleDB fits naturally into Supabase's managed offering and roadmap. You Will: Design, implement, and test new OrioleDB features and integrate them cleanly with PostgreSQL's planner, executor, and surrounding subsystems. Build out and maintain test infrastructure: regression suites, f

sqlpostgresqlai
View job →
GU
greenhouse,DoorDash USA
📍 San FranciscoFull-timeFrom $2.3M/yr
8 hrs ago

About the Team DoorDash Labs is an independent team within DoorDash. We explore robotics and automation to transform last mile logistics in the long term. If you have a passion for applying robotics solutions in a service used by millions of people, then we want to talk to you! About the Role We are hiring a Firmware Validation & Integration Engineer for our autonomy software team. This is a critical role to build robust and scalable validation for our firmware and systems to ensure reliability at every level. In this role, you will work with our electrical, firmware, and autonomy engineers to build the infrastructure and test suites required to validate the system. This includes designing and implementing our Hardware-in-the-Loop (HIL) simulation environments and automation frameworks from the ground up. You will report to the Autonomy Platform Lead on our Autonomy Platform Team at DoorDash Labs. We expect this role to be hybrid with some time in-office and some time remote. You’re excited about this opportunity because you will… Play an integral role on a small and focused team. Design and build Hardware-in-the-Loop (HIL) systems to simulate vehicle dynamics and sensor data for comprehensive firmware and system-level validation. Develop automated test infrastructure and software tools to exercise multiple embedded platforms throughout our robot system. Interface many layers of our control system including vehicle controls, power management, and motion control to ensure seamless system integration. Implement low-level test sequences and validation algorithms to safely stress-test vehicle components such as batteries, drive-train, and thermal management devices. Collaborate with cross-functional teams to identify edge cases and hardware-software corner cases that impact vehicle safety and performance. We’re excited about you because… BS/MS degree in Computer Science, Robotics, Electrical Engineering, or related technical field. 5+ years of experience in validati

pythonawsgit
View job →

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b

typescriptpythonaws
View job →

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b

typescriptpythonaws
View job →
C
Coinbase
📍 - USAFull-timeRemoteFrom $186.1K/yr
28 days ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . We're hiring a Senior Software Engineer to join the CB Node team within the Platform organization. This team runs all the blockchain nodes that power every asset Coinbase offers to customers, currently spanning 55 unique protocols across Ethereum, Bitcoin, Solana, and beyond. You'll focus on performance optimization and reliability across the blockchain platform stack, diagnosing root causes instead of over-provisioning, reducing infrastructure spend, and building the observability and health-check systems that keep our nodes reliable at scale. What you'll do: Own deep performance analysis and optimization across the blockchain node stack, diagnosing root causes of resource consumption and implementing system-level fixes that reduce infrastructure spend while maintaining reliability SLOs. Build observability, health-check, and automated failover capabilities for blockchain nodes, ensuring the team can detect degradation and respond proactively rather than reactively scaling resources. Drive a right-sizing initiative across the node fleet, profiling workloads, identifying over-provisioned instances, and establishing capacity models that balance cost efficiency with headroom for reliability. Partner with protocol integration engineers, wallet teams, and indexer teams to extend performance improvements across the full blockchain platform stack, not just the node la

REMOTEreactawsai
View job →
C
Coinbase
📍 - USAFull-timeRemoteFrom $186.1K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Software Engineer on the Platform Core Automation team, you'll build the AI infrastructure and agentic systems that automate customer support and compliance operations at Coinbase. This team is reimagining what these processes look like when AI handles them end-to-end, replacing manual workflows with intelligent agents that resolve customer issues and execute compliance tasks autonomously. You'll own the design and delivery of LLM-powered systems, grounding techniques, and integration pipelines that directly reduce resolution times, cut costs, and improve accuracy across millions of customer interactions. What you’ll be doing (ie. job duties): Own the design and delivery of agentic AI systems that power Coinbase's customer support and compliance automation, from LLM orchestration through production deployment and monitoring. Build scalable, secure backend infrastructure in Python and Golang that serves AI workloads, including model integration pipelines, guardrails, grounding mechanisms, and measurement frameworks. Drive end-to-end project execution on complex AI initiatives, making technical trade-offs across latency, accuracy, cost, and reliability. Partner with Customer Experience, Compliance, and product engineering teams to identify high-impact automation opportunities and translate operational pain points into AI-powered solutions. Strengthen engine

REMOTEpythonmongodbaws
View job →
O
OpenAI
📍 San FranciscoFull-time
1mo ago

About the Role The Engineering Acceleration team builds and operates the foundational systems that engineers use to build, test, and ship ChatGPT, the API, and OpenAI's infrastructure. We are looking for an engineer to help evolve OpenAI's build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, and software quality. You will work on the systems that determine how quickly and confidently engineers can move: Bazel-based builds, Buildkite pipelines, test selection, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to make OpenAI one of the most productive engineering organizations in the world while preserving a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping useful systems instead of fighting infrastructure. In This Role, You Will Own and evolve Bazel-based build and test workflows across a large, polyglot monorepo. Design and maintain Starlark rules, macros, toolchains, and integrations that make builds reproducible, hermetic, and easy for product teams to adopt. Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, test sharding, retry behavior, and flake isolation. Build systems that reduce unnecessary CI work through affected-target detection, dependency graph analysis, test selection, caching, batching, and smarter scheduling. Improve local development workflows so engineers can reproduce CI behavior, debug build failures, and iterate quickly without learning every detail of the build stack. Operate and optimize build infrastructure across Docker/OCI images, Kubernetes-based runners, cloud resources, and remote cache/exec

typescriptpythonaws
View job →
D
Datadog
📍 Remote, FranceFull-timeRemote
4 days ago

We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, http://incident.io , Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different

REMOTEpythonnode.jsai
View job →
D
Datadog
📍 CaliforniaFull-timeRemote
1mo ago

We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, incident.io, Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different s

REMOTEpythonnode.jsai
View job →
🔔

Get new remote remote remote remote integration reliability engineer 2c technical operations manager jobs by email

Daily job updates · Unsubscribe anytime