Jobiba hiring network

Workload Porting And Performance Engineer Jobs

751 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current workload porting and performance engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

A
Anyscale
📍 Remote• Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: - Optimizing performance of large-scale workloads on Ray - Stability and stress testing infrastructure - Improving fault tolerance (HA) As part of this role, you will: Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements to Ray core Improve the testing process for Ray to make re

restmachine learningai
View job →

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for a Software Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the Infra team, we build the scalable, secure, and robust backbone that enables this vision. Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads. We are seeking a talented Software Engineer with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale’s cloud platform. You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers. A snapshot of projects you may work on Design, build, and scale services that orches

pythonawsazure
View job →

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale is seeking a Staff Software Engineer to lead the technical vision for our Infrastructure team. As a Staff Engineer, you will be responsible for the architectural evolution of our control plane and data plane, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. You will act as a force multiplier, setting the standards for Kubernetes-based cloud-native infrastructure while mentoring engineers and driving cross-functional alignment across the Ray open-source community and our proprietary product teams. Key Responsibilities Architectural Leadership: Define and drive the multi-year technical roadmap for services that orchestrate Ray clusters across diverse cloud and on-premises environments. Systemic Optimization: Lead the design and optimization of high-performance control plane components specifically tailored for large-scale, heterogeneous AI/ML workloads. Platform Reliability: Establish the organization-wide standards for the reliability, scalability, and observability of Anyscale-managed infrastructure. Strategic Integration: Direct the long-term strategy for accelerator integration (GPUs, TPUs) and container management to ens

pythonawsazure
View job →
S
Synthesia
📍 United Kingdom• Full-time
1mo ago

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. We’re looking for an Engineer to join the ML Platform team at Synthesia. Our team builds and operates the systems that allow researchers and product teams to train, serve, and deploy generative models reliably and efficiently . This includes research infrastructure, production serving systems, internal tooling, and the platform interfaces that connect them. A growing part of our mission is making these systems more automation-friendly and agent-oriented , so that workflows can increasingly be operated through reliable tooling rather than manual effort. We’re looking for a strong generalist with a systems mindset: someone who is comfortable working across infrastructure, backend systems, and tooling, and who has seen ML systems in practice. this is not a pure ML Engineer role. We’re especially interested in people who think deeply about reliability, scalability, performance, and resource efficiency in complex production environments. This is a hands-on IC role with significant ownership. You’ll help shape how our ML platform evolves as we scale the number of models, workloads, tools and teams relying on it. What you’ll do Design and improve the platform systems that support model training, evaluation, and product

pythonkubernetesgit
View job →
F
1mo ago

About Us What if your work could drive change in a globally established industry, shaping processes that touch every corner of the world? At Forto, we are at the forefront of change, harnessing the power of AI to revolutionise logistics. We want to reinvent digital supply chains to be transparent, frictionless and sustainable. From day one, our mission has been to simplify global trade – creating a seamless and efficient logistics process. Your role & Mission The Site Reliability Engineering team at Forto is responsible for reliability and developer experience. We enable our development teams to write complex business logic by providing best-in-class tooling and infrastructure. We have a production environment based on GCP, Kubernetes, Terraform, and Helm. On top of that, we have self-service tooling written in TypeScript. “You build it, you run it” - our job is to make that real. This is a high-ownership role on a lean team that directly shapes how 70+ engineers build and ship software. If you care about platform quality and want your work felt immediately across an engineering org, this is a great match for you. What you will do Build out our runtime platform as a self-service product that enables our engineering teams to write code, run workloads, and drive engineering culture forward. Bring software development skills and practices into platform engineering, such as code quality, domain-driven design, and test-driven development. Own the developer portal and internal platform roadmap, including leading this year's overhaul of our CI/CD pipelines in collaboration with all product teams. Ensure site reliability by building observability solutions, deployment, and disaster recovery capabilities. Own reliability standards end-to-end through SLOs and error budgets — shaping how teams balance velocity and risk. Drive infrastructure cost optimisation across Kubernetes, MongoDB, and Datadog at scale. Improve our security posture through tooling, compliance work, and

typescriptmongodbaws
View job →
S
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We’re looking for an experienced Database Support Engineer to join our Support Team and help developers unblock complex issues, build more reliably, and get the most out of the Supabase platform. You’ll work closely with customers and engineering, helping us improve product quality and developer experience based on real-world usage. This role is ideal for someone who thrives in async, fast-paced environments and is excited about building developer tools that scale to millions. What You’ll Be Responsible For In this role, you’ll: Resolve Technical Issues: Investigate and resolve advanced customer issues across Postgres, Auth, RLS, Storage, Realtime, Edge Functions, and client libraries. Provide Consultative Advice: Offer proactive guidance to optimize workloads, avoid common pitfalls, and provide tailored advice to our most advanced users. Bridge the Gap to Engineering: Reproduce bugs, isolate root causes, propose workarounds, and escalate to engineering with solid reproduction steps. Investigate Deeply: Ask smart, targeted questions that turn incomplete reports into actionable cases. Communicate Effectively: Communicate clearly and empathetically with users of all levels—especially during high-impact or time-sensitive situations. Drive Product Improvement: Spot patterns across tickets and recommend improvements to documentation, tooling, and the product itself. Mentor and Lead: Mentor junior team members and help raise the overall bar for support quality. You Might Be a Good Fit If Experience: You have 7+ years in technical support, databases, backend engineering, SRE, or similar roles. Postgres Expertise: You know PostgreSQL deeply—including autovacuum behavior, WAL growth, long

sqlpostgresqlaws
View job →
S
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We’re looking for an experienced Database Support Engineer to join our Support Team and help developers unblock complex issues, build more reliably, and get the most out of the Supabase platform. You’ll work closely with customers and engineering, helping us improve product quality and developer experience based on real-world usage. This role is ideal for someone who thrives in async, fast-paced environments and is excited about building developer tools that scale to millions. What You’ll Be Responsible For In this role, you’ll: Investigate and resolve advanced customer issues across Postgres, Auth, RLS, Storage, Realtime, Edge Functions, and client libraries. Provide consultative, proactive advice to optimize workloads, avoid common pitfalls, and offer tailored advice to our most advanced users. Reproduce bugs , isolate root causes, propose workarounds, and escalate to engineering with solid reproduction steps. Ask smart, targeted questions that turn incomplete reports into actionable cases. Communicate clearly and empathetically with users of all levels—especially during high-impact or time-sensitive situations. Spot patterns across tickets and recommend improvements to documentation, tooling, and the product itself. Mentor junior team members and help raise the overall bar for support quality. You Might Be a Good Fit If Experience: You have 7+ years in technical support, databases, backend engineering, SRE, or a similar field. Postgres Expertise: You know PostgreSQL deeply—autovacuum behavior, WAL growth, long-running transactions, table bloat, and other internals. Performance: You’re strong in high-performance SQL and query optimization. Troubleshooting: You’re comfortable tr

sqlpostgresqlaws
View job →
S
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We’re looking for an experienced Database Support Engineer to join our Support Team and help developers unblock complex issues, build more reliably, and get the most out of the Supabase platform. You’ll work closely with customers and engineering, helping us improve product quality and developer experience based on real-world usage. This role is ideal for someone who thrives in async, fast-paced environments and is excited about building developer tools that scale to millions. What You’ll Be Responsible For In this role, you’ll: Investigate and resolve advanced customer issues across Postgres, Auth, RLS, Storage, Realtime, Edge Functions, and client libraries. Provide consultative, proactive advice to optimize workloads, avoid common pitfalls, and offer tailored advice to our most advanced users. Reproduce bugs , isolate root causes, propose workarounds, and escalate to engineering with solid reproduction steps. Ask smart, targeted questions that turn incomplete reports into actionable cases. Communicate clearly and empathetically with users of all levels—especially during high-impact or time-sensitive situations. Spot patterns across tickets and recommend improvements to documentation, tooling, and the product itself. Mentor junior team members and help raise the overall bar for support quality. You Might Be a Good Fit If Experience: You have 7+ years in technical support, databases, backend engineering, SRE, or a similar field. Postgres Expertise: You know PostgreSQL deeply—autovacuum behavior, WAL growth, long-running transactions, table bloat, and other internals. Performance: You’re strong in high-performance SQL and query optimization. Troubleshooting: You’re comfortable tr

sqlpostgresqlaws
View job →
S
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About this role We’re hiring experienced performance engineers to make performance measurement at Supabase rigorous, repeatable, and actionable by product teams. This role owns the benchmarking and profiling systems we use to detect regressions, explain performance changes, and publish credible numbers for customers and the market. You’ll build the tooling and the methodology and make it easy for product teams to use it correctly. You’d join a new performance team as one of its first members, working alongside a deeply seasoned performance engineer, with the team growing to roughly six this year. There’s no legacy process to inherit — you’ll help define how we measure and communicate performance metrics at Supabase. What you’ll do Build and evolve benchmarking, profiling, and load-testing tooling (cloud, database, and end-to-end). Define benchmark suites and reference workloads that reflect real customer behavior (and evolve them as the products change). Establish performance baselines and regression gates across products (variance tracking, environment control, and detection of changes). Build mechanisms that turn benchmark results into action: triage workflows, owner routing, and clear, reproducible reports. Partner closely with database teams (e.g. Multigres, OrioleDB) to identify bottlenecks and land concrete performance improvements. Enable the company to publish credible, repeatable performance numbers for customers and the market. Help teams self-serve performance testing and make performance a first-class engineering concern. Problems you might work on Build our continuous benchmarking infrastructure so product teams can spot performance regressions early. Ideally, teams can run benchma

aigorust
View job →
C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of committed researchers and engineers. The mission of the team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will work on advancing core audio model serving metrics, including latency, throughput, and quality by diving deep into our systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You’ll collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. You may

pythongitrest
View job →
C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of researchers and engineers focused on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques that improve how models execute in production, driving lower latency, higher throughput, and consistent quality across diverse workloads. As an engineer on this team, you’ll work across the inference stack to improve core performance metrics by diving deep into model execution, identifying bottlenecks, and developing innovative optimizations. You’ll collaborate closely with modeling and systems teams to experiment, measure, and ship improvements that meaningfully accelerate inference. As the team evolves, you’ll have opportunities to build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large-scale architectures. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, e

pythongitrest
View job →
C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc

awsazuregcp
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Role Overview We are seeking a Platform Experience and Developer PM to own how developers and enterprise technical teams build on, integrate with, and operate Cohere's model platform. This is a high-leverage role sitting at the intersection of three domains: Managed Services and Models as a Service. Own Cohere's managed service offerings as a product. This is broader than model serving alone. It includes the full range of how enterprises consume and operate Cohere's capabilities, from shared multi-tenant model access to dedicated single-tenant deployments, and from synchronous real-time inference to high-volume asynchronous workloads. You will own the product thinking around deployment models, data residency and regional compliance requirements, self-serve provisioning, and the operational controls that give enterprises confidence in running production workloads on Cohere’s infrastructure. API and SDK. Own the roadmap for how developers build on Cohere. This means setting the direction for our APIs and SDKs, thinking carefully about interface design and ergonomics, and ensuring we ship developer primitives that are stable, well-

C
Clickup
📍 United States• Full-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 We're looking for a Staff Data Engineer to own the architecture and technical vision of our data platform. This is a high-leverage, high-autonomy role where you'll set the technical bar for the team, drive cross-functional alignment on data infrastructure strategy, and solve our hardest engineering problems. You'll operate across AWS serverless technologies, Snowflake, dbt, and Terraform, but your impact goes well beyond any single tool: you'll shape how we think about reliability, scalability, cost, and developer experience at the platform level. This role is for someone who doesn't just build great systems, but makes the engineers around them better. The Role: Own the technical architecture of ClickUp's data platform, making design decisions that balance scalability, cost, reliability, and velocity. Define and drive the technical roadmap for data infrastructure in partnership with leadership. Design systems at scale : build frameworks, abstractions, and patterns that other engineers use daily. Lead complex, cross-team technical initiatives spanning data engineering, analytics engineering, data science, and data analytics. Drive cost optimization across cloud infrastructure and compute, turning efficiency into a competitive advantage. Build and evolve our data pipelines using AWS serverless (Lambda, Fargate, Step Functions, Kinesis, S3, DynamoDB, Aurora), Snowflake, and dbt. Establish and champion engineering standards : observability, testing, CI/CD, code review, and documentation practices. Design and maintain infrastructure for AI/ML workloads , including LLM frameworks, feature pipelines, training

pythonsqlaws
View job →
C
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 This role bridges infrastructure and product engineering: you'll build genuine partnerships across product, AI, and enterprise teams so that ownership is shared and velocity is never blocked by platform constraints. As Director, you'll set a forward-looking cloud vision, proactively align with stakeholders across the business, and ensure the platform scales for multi-shard, multi-region growth while meeting security and compliance commitments (SOC 2, business continuity/disaster recovery, and enterprise security frameworks). The Role: Cost Efficiency: Drive significant annual infrastructure savings by migrating data workloads to EKS and self-hosting key services such as OpenSearch. Ingress Convergence: Deprecate legacy frontend ALBs and consolidate to a single EKS-managed ALB per shard, unblocking faster deployments across the org. Coverage & Bench Depth: Eliminate single points of ownership across Networking, OpenSearch, and Terraform through cross-training and targeted hiring into coverage gaps. Stakeholder Alignment: Stand up a recurring alignment cadence with Product, AI, Enterprise, and Security so infra planning is driven by demand, not ad-hoc interrupts. Automation: Ship automated shard buildout via Backstage to remove manual toil from enterprise scaling. Reliability: Cut P0/P1 incidents attributed to Cloud Platform (DNS, ALB misconfiguration) through hardened ingress patterns and Terraform-policy guardrails, including blocking unauthenticated public endpoints. Roadmap Ownership: Deliver a roadmap covering Agent enablement, centralized IaC, and deployment rollout acceleration, tied to AI and

awsmachine learningai
View job →
🔔

Get new workload porting and performance engineer jobs by email

Daily job updates · Unsubscribe anytime