We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity As a Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with a proven track record of building and scaling resilient infrastructure. What you'll do Platform Orchestration: Work on a large-scale K8s infrastructure platform, ensuring high availability and performance. Automation: Drive the evolution of internal tooling to streamline platform delivery. This role requires Experience: Solid hands-on background in DevOps, Site Reliability, or Infrastructure Engineering. Kubernetes Mastery: Deep internal knowledge of K8s primitives (Deployments, StatefulSets, Services) and hands-on experience writing custom Kubernetes Operators. Golang Proficiency: Proficiency in Go, specifically for infrastructure automation and systems programming. Operations-Heavy Mindset: A proven track record of managing production environments and handling high-severity incidents. Cloud Infrastructure: Hands-on experience with cloud-native scaling tools (e.g., Karpenter, Cluster API) and Day 1/Day 2 operations of K8s clusters. Tooling: Familiarity with Helm and GitOps workflows (e.g., ArgoCD or Flux). Please note that visa sponsorship is not available for this position. Fostering a diverse, welcoming and inclusive environment is important to us. We work hard to make everyone feel comfortable bringing their best, most authentic selves to work every day. We cele
Jobiba hiring network
Software Engineer Internal Systems Salary India Jobs
6,428 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current software engineer internal systems salary india jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work
About the Team The Cooperative AI team is scaling OpenAI with OpenAI. We are building a model powered knowledge system that evolves and learns as our products, systems and customers evolve. We leverage our state of the art models, technologies, and products (some external, some still in the lab) to assist or completely automate robust operations supporting both internal and external customers. We support OpenAI customers and internal partners globally, powering systems from customer support to integrity to product insights. We are a self-contained multi-disciplinary team, who enjoy a lightning fast feedback loop with customers at scale, some of whom sit just a few pods away. We iterate fast, and engineer for reliable long-term impact. We're constantly looking for the similarities and patterns in different types of work, and focus on building simple primitives, to apply world class knowledge to many domains. The work of this team exemplifies use of OpenAI technologies. We build systems so everyone can see the leverage that is possible with well designed AI-based implementations. We do this by working through internal use cases focused on Customers (specifically knowledge systems, automation systems, and automated agent systems) to prove impact, then we scale. About the Role We’re looking for a Backend Software Engineer to help architect and scale the infrastructure that powers our knowledge systems. This is a deeply technical and highly cross-functional role where you’ll build robust systems and backend services that serve as the foundation for how knowledge is created, accessed, and applied across OpenAI. In this role, you will: Design, build, and maintain backend services and APIs to support intelligent automation and knowledge systems Integrate and structure data across internal platforms, transforming it into formats optimized for use by downstream systems and AI workflows. Collaborate closely with product, research, and engineering teams to integrate OpenAI mode
About the Team The Cooperative AI team is scaling OpenAI with OpenAI. We are building an AI powered knowledge system that evolves and learns as our products, systems and customers evolve. We leverage our state of the art models, technologies, and products (some external, some still in the lab) to assist or completely automate robust operations supporting both internal and external customers. We support OpenAI customers and internal partners globally, powering systems from customer support to integrity to product insights. We are a self-contained multi-disciplinary team, who enjoy a lightning fast feedback loop with customers at scale, some of whom sit just a few pods away. We iterate fast, and engineer for reliable long-term impact. We're constantly looking for the similarities and patterns in different types of work, and focus on building simple primitives, to apply world class knowledge to many domains. The work of this team exemplifies use of OpenAI technologies. We build systems so everyone can see the leverage that is possible with well designed AI-based implementations. We do this by working through internal use cases focused on Customers (specifically knowledge systems, automation systems, and automated agent systems) to prove impact, then we scale. About the Role We’re looking for Software Engineers who're passionate about blending production-ready platform architecture with new tech and new paradigms. You’ll push the boundaries of OpenAI’s newest technologies to enable interactions and automations that are not only functional, but delightful. We value proactive, customer-centric engineers who can get the foundational details right (data models, architecture, security) in service of enabling great products. In this role, you will: Own the end-to-end development lifecycle for new platform capabilities and integrations with other systems Collaborate closely with engineers, data scientists, information systems architects, and internal customers to understand th
Scale AI is seeking a highly skilled and motivated Software Engineer, ARC (Architecture, Reliability, & Compute) to join our dynamic Public Sector Engineering team. As a part of this team, you will define how the company ships software, establishing the patterns for deploying into complex government and high-security environments, rather than just running Terraform scripts. You will build and maintain internal CLIs/tools that standardize testing, deployment, environment management and are tools that engineering relies on to prevent downstream breakages. You will execute on automated deployment efforts to pay down tech debt, creating fully functional staging/testing environments, and defining the company's standard for safe deployments. You will: Design and implement secure scalable backend systems for Public Sector customers, leveraging Scale's modern and cloud-native AI infrastructure. Own services or systems and define their long-term health goals, while also improving the health of surrounding components. Re-architect the stack to run in compliant or restrictive environments. This requires designing swappable components (auth, storage, logging) to meet government/security mandates without breaking the product. Collaborate with cross-functional teams to define and execute the vision for backend solutions, ensuring they meet the unique needs of government agencies operating in secure environments. Participate actively in customer engagements, working closely with stakeholders to understand requirements and deliver innovative solutions. Contribute to the platform roadmap and product strategy for Scale AI's Public Sector business, playing a key role in shaping the future direction of our offerings. Must have: At least an active secret clearance and the ability & willingness to up level to TS/SCI with CI Poly. This is a requirement and candidates will not be considered who do not hold at least a secret clearance Ideally you'd have: Full Stack Development: Prof
About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team An exciting opportunity to join a new team within the Software Operations group. The Build Engineering team is a new function within Software Infrastructure, which focuses on the overall process of building and integration of the Machine Learn ing S oftware S tack. You will work closely with the QA and development teams to get an understanding of how our ML SW stack is built, helping to ensure good build practices, and proving that the stack works together and is reproducible in secure, sandboxed environments. Responsibilities and Duties Developing our internal t
About the Team Come help us build and develop tools serving hundreds of engineers internally! We’re looking for a Fullstack Software Engineer to join our Developer Insights team. About the Role Our mission is to improve the developer experience of engineers at DoorDash by building various internal products, including our internal developer portal, Developer Insights. Our success as a platform team depends on the success of the product teams we serve. Because of this, we invest in building a strong community that encourages participation and promotes best practices. You’re excited about this opportunity because you will… Introduce cutting edge technologies to our engineering organization, including tools built on LLMs Build new features for Developer Insights (using Backstage.io) Improve the developer experience for all of our engineers Work and collaborate across team boundaries. Contribute features and bug fixes to upstream open-source projects. Mentor and educate your peers. Lead the team in a technical fashion and assist in roadmap planning and measurement of existing features. Represent the team at large in OKR and engineering all-hands presentations. Context switch from frontend to backend to data depending on the need that arises. We’re excited about you because… You have at least 2 years of experience in web technologies using Typescript with React on the frontend with Java, Kotlin, Python or Go backend experience. You have a product mindset and apply that to how you would build out platform services. You love systems and software, and you're proficient in both. You’re curious and dive deep into different system architectures. You are an organized and excellent written and verbal communicator. You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software Compensation The successful candidate’s starti
About the Team Our mission is to provide a world-class development experience that makes DoorDash's web engineers among the most productive in the industry. We achieve this by creating the tools that enable all teams at the company to ship features quickly and reliably. Because their success is our success, we are deeply invested in building a strong, collaborative web community that champions best practices and welcomes participation. About the Role As a Software Engineer on the Developer Experience team, you will build the foundational pieces for all DoorDash, Wolt and Deliveroo Web applications. These include monorepos, build & CI systems and agent-first development tooling. You will work closely with engineers and other internal stakeholders to deliver large and impactful initiatives. Additionally, you will be a culture carrier for our Web engineers through mentorship, education, and engagement of your peers. You will report into the Engineering Manager of our Web Developer Experience team in our Developer Platform organization. You must be located in either San Francisco, CA, Sunnyvale, CA, Los Angeles, CA, Seattle, WA, or New York, NY. You’re excited about this opportunity because you will… Shape the future of Web Development. You will have a direct and meaningful impact on the daily workflows of every web engineer at the company, enhancing their productivity and overall developer experience. Build from the ground up. You will architect and implement foundational libraries, cutting-edge build systems, and innovative development tools that serve as the bedrock for all of our web applications. Solve complex, high-impact challenges. You will tackle some of the most significant technical hurdles in web engineering, and the solutions you deliver will be leveraged by hundreds of employees across numerous product teams. Act as a force multiplier. Your work will directly empower product teams to build, test, and release new features to our customers fa
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build cloud-native data and storage services for hybrid and multi-cloud infrastructure, including dataset discovery, ingestion, governance, checkpointing, observability, and low-latency access. Develop scalable cloud-native services and APIs that support exabyte-scale, high-performance GPU training and inference workflows. Work closely with product managers, internal AI teams, platform teams, and partner engineering teams to understand requirements and turn them into reliable production systems. Collaborate with SRE, operations, and support teams to improve service reliability, performance, observability, on-call readiness, and operational scale. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, and verification. What we need to see: BS in Computer Science, Information Systems, Computer Engineering, or equivalent experience, with 5+ years of software engineering experience. Strong foundation in algorithms, data structures, distributed systems, and practi
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys
Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. At Discord, we believe everyone can find a place where they belong. Our mission is to help make it easy for everyone to find and join meaningful conversations and to make every part of our product feel smart and delightful. Discord is a rapidly scaling company that puts Data at the heart of its decision-making and future growth. The Experimentation team at Discord is responsible for building the tools, systems, and processes that empower Discord employees to continually test hypotheses and improve our product in a data-informed way. We are looking for a Staff Software Engineer to lead the development of Discord's next-generation experimentation platform, including building state-of-the-art AI assisted experimentation workflows. You will lead and mentor engineers, fostering their growth and enhancing their skills. What you'll be doing Design, build, and support Discord's next generation experimentation platform: scope of work would include core services, data pipelines, configuration tooling, statistical analysis tools and UX. Collaborate with external (sometimes non-technical) stakeholders and customers throughout the company Understand and influence product telemetry practices to support experimentation needs. Use your technical expertise to build delightful user experience and high-scale data systems that power internal teams and product features used by many millions of users every day Deliver business results by collaborating with stakeholders across Discord to experiment on features, machine learning models and more. What you should have You have 7+ years of experience as a full-stack SWE You have
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand, and tackle issues, analyze performance, and get the most out of their software and infrastructure. This is a unique opportunity to shape the future of observability while pioneering the next generation of engineering: we are actively revolutionizing how we build software by embedding modern, agent-powered workflows directly into our daily development lifecycle. You’ll tackle complex distributed systems problems while helping drive internal innovation on the frontlines of AI-assisted engineering. About the team This position is for the Service Architecture Intelligence team. You will be building, improving, and maintaining a distributed service architecture capable of ingesting large volumes of data, analyzing spans and traces, and publishing them downstream so the UI can offer an exceptional experience to our customers. We work with data at scale, and our pipeline is built with a diverse tech stack (Java, Kafka, Redis, public cloud services, and more). You will work alongside a team of talented engineers solving complex distributed systems challenges. If you're passionate about performance and scale, and want to contribute to one of the largest and fastest-growing observability platforms while co-crea
Get new software engineer internal systems salary india jobs by email
Daily job updates · Unsubscribe anytime