We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Are you passionate about technology and ready to dive into the world of Site Reliability Engineering? We are looking for enthusiastic individuals to join our team as Site Reliability Engineer II, where you'll play a crucial role in ensuring the reliability and performance of our platform. This is a top-notch opportunity to learn from skilled engineers and contribute to solving complex technical challenges. You will be involved in monitoring systems, responding to incidents, and developing automation to streamline operations. We are looking for someone eager to grow their skills, learn new technologies, and contribute to a culture of reliability. The Site Reliability Engineering (SRE) team integrates software and systems engineering to design and manage large-scale, distributed, and fault-tolerant systems. The team is responsible for ensuring high reliabili
Jobiba hiring network
Ai Systems Engineer Jobs
10,000 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current ai systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. We are building and maintaining a highly scalable asynchronous platform that empowers our organization to handle critical business cases. As a software engineering team, our mission is to create robust and innovative solutions that drive the success of our business and deliver unparalleled value to our customers. We adopt Infrastructure as Code practice to automate the provisioning and configuration of our resources, which helps reduce manual configuration and improve consistency. Our team culture is built on collaboration, open communication, and a supportive environment where each member's ideas are valued and contributions are recognized. We believe in the importance of fostering a positive workplace culture that inspires innovation and creativity. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and production readiness reviews, platform management and capacity planning ceremonies with cross-functional teams Document Infrastructure operations process and insights, identify repeatable actions and ruthlessly automate repetitive tasks Participate in our teams on-call rotations, respond to incidents and support other teams mitigate customer impacting events Experience: 5+ years experience working on teams responsible for software development, automation and systems engineering Experience building large-scale infrastructure, distributed systems or networks. Knowledge with SQS,
About the Team The Workload Networking team is responsible for the collective communication stack used in our largest training jobs. Using a combination of C++ and CUDA we work on novel collective communication techniques that enable efficient training of our flagship models on our largest custom built supercomputers. The models we train are key ingredients to the AI research progress at OpenAI and the field as a whole, and we continually incorporate learnings from our entire research org into our training platform. About the Role As a Software Engineer, Networking you will design and implement custom networking collectives that are tightly integrated into our training stack. We’re looking for people who have a background in low level performance critical software. Experience with collective communication is a bonus. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Collaborate closely with ML researchers to design and implement efficient collective operations in C++ and CUDA. Ensure that our largest training jobs take full advantage of the different network transports used in our supercomputers. Work on simulations to inform our future supercomputer network designs. You might thrive in this role if you: Have written distributed algorithms using RDMA in the past. Are comfortable writing low level performance sensitive CPU and/or GPU code. Are familiar with network simulation techniques. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voic
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
About the Team The RL and Reasoning team drives the core reasoning paradigm and has created groundbreaking innovations such as o1 and o3. They focus on pushing the boundaries of reinforcement learning research, building next-generation generative models, and deploying them at scale. About the Role As a Research Engineer/Research Scientist at OpenAI, you will advance the frontier of AI alignment and capabilities through cutting-edge RL methods. Your work will sit at the heart of training intelligent, aligned, and general-purpose agents, including the systems that power various models. We’re looking for people who have a background in reinforcement learning research, are able to iterate quickly, and are proficient at coding. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. You might thrive in this role if: You love being on the cutting edge of RL and language model research. You’re a self-starter who takes initiative and ownership of ideas, driving them to completion. You value principled approaches, simple experiments in tightly-controlled settings, and reaching trustworthy conclusions which stand the test of time. You thrive in a fast-paced, dynamic, and technically complex environment where rapid iteration is key. You’re comfortable diving into a large ML codebase to debug and improve it. You have a deep understanding of machine learning and machine learning applications. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the ful
About the team OpenAI’s Forward Deployed Engineering (FDE) team partners with global pharma and biotech, CROs, and research institutions to deploy production-grade AI systems across the R&D value chain. We operate at the intersection of customer delivery and core platform development, converting early deployments into repeatable system standards and evaluation practices that scale across regulated environments. About the role As a Life Sciences FDE Manager, you’ll lead a team of FDEs delivering production AI systems across drug discovery and development workflows. You’ll own delivery outcomes and team leverage while staying hands-on as a player-coach. This includes building and shipping alongside the team, setting technical direction, and maintaining a high bar for production-grade systems in regulated environments. We measure success through the health and quality of your FDE team, production adoption and measurable workflow impact, the quality of eval-driven feedback delivered back to Product and Research, and the repeatability of deployment patterns across life sciences customers. This role is based in New York City We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. This role will require travel up to 25%. In this role you will Lead and grow a team of FDEs delivering production AI systems across regulated life sciences environments Be accountable for your team’s end-to-end delivery outcomes, balancing scope, speed, robustness, and risk in high-stakes deployments Coach and develop engineers through direct feedback, high technical standards, and clear expectations for execution and ownership Operate as a player-coach, directly contributing to production systems while leading, coaching, and setting technical direction Guide teams through ambiguous, multi-workstream engagements spanning data, workflows, infrastructure, security, and scientific stakeholders Run evaluation loops that measure model and system quality against
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. About Fern, a Postman Company Fern helps software companies build a world-class API experience. Our customers include industry leaders like Nvidia, Square, and Twilio, as well as fast-growing AI companies like ElevenLabs and OpenRouter. In the next year, the majority of API integrations will be implemented by AI agents. Agents don’t read marketing pages. They need strongly typed schemas, structured endpoints, deterministic contracts, and documentation that can be fed into a context window. Our team is small and talent-dense. We’re a team of builders from Google, Palantir, Amazon, and Uber, working together in New York City. About the Role As a Senior Software Engineer at Fern, you’ll build APIs, scale AI infrastructure, and design developer experiences that reach millions of people. Scale infrastructure to keep up with growth: Work on AI systems deployed on Vercel, AWS, and Turbopuffer that must be performant, reliable, and secure under real-world load. Stay close to customers: We work directly with API teams at companies like Nvidia, Twilio, and Square. You’ll debug real production edge cases, shape APIs ba
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Staff Engineer in Micron’s HBM Product Engineering Component Validation team, you will be an individual contributor responsible for the validation of advanced High Bandwidth Memory (HBM) products. You will use proven knowledge in dynamic random-access memory and high-bandwidth memory architecture to compose validation coverage, implement and analyze validation runs, lead product debug and failure analysis toward a zero-escape validation goal across test modes and product achievements. You will also adopt AI-enabled ways of working, using enterprise AI tools to accelerate your engineering tasks and improve efficiency, work within Micron’s global validation organization and collaborate with groups at different locations. Responsibilities: Translate DRAM/HBM develop intent and architecture into effective verification scope and test cases. Contribute to validation of new and high-risk product features, working with Design Validation, and Systems Engineering. Review design specifications, schematics, and datasheets, provide design feedback and represent the Product Engineering perspective in Design and Verification reviews. Read, analyze, and debug circuit blueprints to detect design-related failure mechanisms, and recommend approaches for verification and fault diagnosis. Perform root-cause analysis and examination of electrical failures (debug) of validation findings, distinguishing developed marginality from process and test-related issues. Apply automation and modern tooling (including
Bloomreach is building the world’s premier agentic platform for personalization .We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey. We're taking autonomous search mainstream, making product discovery more intuitive and conversational for customers, and more profitable for businesses. We’re making conversational shopping a reality, connecting every shopper with tailored guidance and product expertise — available on demand, at every touchpoint in their journey. We're designing the future of autonomous marketing , taking the work out of workflows, and reclaiming the creative, strategic, and customer-first work marketers were always meant to do. And we're building all of that on the intelligence of a single AI engine — Loomi — so that personalization isn't only autonomous…it's also consistent.From retail to financial services, hospitality to gaming, businesses use Bloomreach to drive higher growth and lasting loyalty. We power personalization for more than 1,400 global brands, including American Eagle, Sonepar, and Pandora. About the Team: AI Search The AI Search team owns Bloomreach’s core search platform, serving hundreds of millions of queries per day across enterprise customers. We build and operate highly scalable, low-latency systems that combine traditional information retrieval with modern ML-driven ranking and semantic understanding — all in production at scale. This team sits at the intersection of systems engineering, search relevance, and applied ML , with direct impact on customer revenue and experience. The Role: As a Senior Staff Engineer , you will be a technical leader responsible for shaping the architecture and long-term evolution of Bloomreach’s Search platform. You will lead complex initiatives, set technical direction, and mentor engineers, while remaining deeply hands-on. This role is ideal for someone
OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT We are looking for a Software Engineer, Platform & Integrations . You will own the end-to-end delivery of complex features, optimize backend performance, and collaborate on architectural decisions that impact more than 1,000 enterprise customers worldwide. This is an ideal role for an experienced engineer who thrives on technical autonomy, enjoys tackling complex cloud-native challenges, and wants to have a direct impact on product direction. Feature Ownership & Architecture: Design, build, and maintain scalable, secure, and highly observable backend services and microservices using Java and Spring Boot. Platform Evolution: Actively contribute to upgrading our core platform infrastructure, focusing on system resilience, performance tuning, and seamless component communication. Security & Compliance: Implement rigorous secure coding practices to safeguard data exchange and ensure compliance across our multi-tenant SaaS environment. Collaborative Execution: Work closely with Product, QA, and senior le
About AlphaSense: The world’s most sophisticated companies rely on AlphaSense to remove uncertainty from decision-making. With market intelligence and search built on proven AI, AlphaSense delivers insights that matter from content you can trust. Our universe of public and private content includes equity research, company filings, event transcripts, expert calls, news, trade journals, and clients’ own research content. The acquisition of Tegus by AlphaSense in 2024 advances our shared mission to empower professionals to make smarter decisions through AI-driven market intelligence. Together, AlphaSense and Tegus will accelerate growth, innovation, and content expansion, with complementary product and content capabilities that enable users to unearth even more comprehensive insights from thousands of content sets. Our platform is trusted by over 6,000 enterprise customers, including a majority of the S&P 500. Founded in 2011, AlphaSense is headquartered in New York City with more than 2,000 employees across the globe and offices in the U.S., U.K., Finland, India, Singapore, Canada, and Ireland. Come join us! About the team: The role of Engineer II, Information Engineering is responsible for designing, building, and operating scalable enterprise platforms that power internal teams across the organization. This role approaches technology challenges with an engineering mindset - developing automated, reliable, and secure solutions that reduce manual work and enable teams to operate efficiently at scale. Our mission is to design, build, and operate secure, scalable IT platforms and identity services that power the company. We leverage engineering, automation, and deep systems expertise to reduce operational overhead, increase reliability, and support long-term organizational growth. This engineer will focus on identity and access management, enterprise systems engineering, enterprise security, and automation-driven platform development, while supportin
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! As a Manager of Security Engineering, your key responsibilities include: Serve as trusted advisor to team’s leadership and partner teams by clearly articulating business risks associated with security issues Execute the long-term vision for the Security team in alignment with Cohere’s product and business goals. Collaborate closely with leadership to prioritize high-impact initiatives and strategic customer engagements. Vulnerability Management: Develop and implement enterprise-wide vulnerability management processes and tooling, including identification, prioritization, remediation tracking, and reporting, including customer artifacts Static Application Security Testing (SAST): Establish SAST programs, integrate tools into CI/CD pipelines, and analyze results to identify and remediate security flaws in source code Dynamic Application Security Testing (DAST): Implement DAST methodologies, configure scanning tools, and conduct regular assessments of running applications Penetration Testing: Lead and oversee internal and external penetration testing engagements, including web application, API, network and agentic AI platform inclu
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX About the Role You’ll help define how machine learning models run across Cloudflare’s global network, from frontier open LLMs and real-time voice models to customer-deployed models served on heterogeneous GPUs and next-generation accelerators. You’ll work with systems engineers, product teams, hardware partners, and AI/ML engineers to bring models into production with low latency, strong reliability, and effic
The data management software market is transforming how organisations build and run applications. MongoDB is the leading developer data platform and the first database provider to IPO in more than 20 years. Join us at the forefront of data and application development. MongoDB Technical Services Engineers combine deep technical expertise with exceptional problem-solving and customer-service skills. You’ll advise customers and resolve complex challenges across MongoDB Core, drivers, Atlas, Cloud Manager, cloud platforms, and infrastructure. We’re looking for candidates based in Dublin to join our vibrant office and collaborative in-office team. This is a five-day role with one of the following schedules: Tuesday–Saturday, Sunday–Thursday, or a five-day pattern covering both Saturday and Sunday. Under our hybrid model, employees on weekend schedules are expected to work from the office two days per week. Cool things you’ll do You’ll help customers troubleshoot complex issues and run critical MongoDB workloads with confidence. You’ll: Solve customer challenges across architecture, performance, recovery, and security Lead investigations from diagnosis to resolution, providing clear, actionable guidance Partner with Product Management and Engineering to advocate for customers and improve MongoDB Build tools, documentation, and training while mentoring peers and raising technical excellence What you need We value curiosity, adaptability, strong technical foundations, and a genuine desire to help customers. You should bring many of the following: 5–6 years of experience in technical support, systems engineering, database administration, SRE, or a related field Experience running complex, mission-critical production database systems Strong Linux and systems engineering skills, including performance, memory, I/O, storage, networking, security, clustering, and troubleshooting A solid understanding of networking concepts and protocols, including DNS, TCP/IP, and SSL/TLS Ability
The worldwide data management software market is massive, forecasted to grow from approximately $82 billion in 2023 to approximately $137 billion in 2027. At MongoDB, we are transforming industries and empowering developers to build amazing apps that people use every day. We are the leading developer data platform and the first database provider to IPO in over 20 years. Join our team and be at the forefront of innovation and creativity. We are looking for a Technical Services Engineer to join our team in Dublin for a dedicated Tuesday through Saturday shift. We are looking to speak to candidates who are based in Dublin for our hybrid working model. About the Team MongoDB Technical Services Engineers (TSEs) use their exceptional problem-solving and customer service skills, along with deep technical experience, to advise customers and solve complex MongoDB problems. TSEs are experts in the entire MongoDB ecosystem, including the database server, drivers, and services like Atlas and Cloud Manager. Our engineers combine MongoDB expertise with passion, initiative, and teamwork to achieve exceptional results. Cool Things You’ll Do Work alongside our largest customers and partners to solve complex issues involving architecture, performance, recovery, and security Serve as an expert resource on standard methodologies for running MongoDB at scale Act as an advocate for customer and partner needs by collaborating with product management and development teams Contribute to internal projects, such as the software development of support tools for performance, benchmarking, and diagnostics Mentor and ramp up new team members while building knowledge of new product lines within the MongoDB ecosystem Work with top ISV partners to ensure excellence for joint customers What You Need Solid hands-on experience with systems engineering, including Linux performance, memory management, I/O tuning, configuration, security, networking, and troubleshooting A strong understanding of net
Get new ai systems engineer jobs by email
Daily job updates · Unsubscribe anytime