Inference Engineer

Bengaluru, Karnataka, India on site Until 9/27/2026 5+ years exp First posted July 29, 2026 Last posted July 29, 2026
Job description

𝗧𝗵𝗶𝘀 𝗿𝗼𝗹𝗲 𝗶𝘀 𝗳𝗼𝗿 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗪𝗲𝗲𝗸𝗱𝗮𝘆'𝘀 𝗰𝗹𝗶𝗲𝗻𝘁𝘀

𝗦𝗮𝗹𝗮𝗿𝘆 𝗿𝗮𝗻𝗴𝗲: 𝗥𝘀 𝟮𝟰𝟬𝟬𝟬𝟬𝟬 - 𝗥𝘀 𝟯𝟲𝟬𝟬𝟬𝟬𝟬 (𝗶𝗲 𝗜𝗡𝗥 𝟮𝟰-𝟯𝟲 𝗟𝗣𝗔)

Experience: 5+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are seeking a highly skilled Inference Engineer with strong expertise in Large Language Models (LLMs) and vLLM to build, optimize, and scale high-performance AI inference platforms. This role is ideal for professionals who are passionate about deploying production-grade generative AI systems, optimizing model serving, and improving inference efficiency for large-scale applications.

As an Inference Engineer, you will be responsible for designing, implementing, and maintaining scalable LLM inference infrastructure that delivers low-latency, high-throughput AI experiences. You will work closely with AI Researchers, Machine Learning Engineers, Platform Engineers, and Product teams to deploy state-of-the-art language models, optimize serving pipelines, and improve resource utilization across GPU-based environments. This role offers the opportunity to work with cutting-edge AI technologies while solving complex performance, scalability, and infrastructure challenges.

Requirements

Key Responsibilities

  • Design, deploy, and maintain scalable inference infrastructure for Large Language Models (LLMs) in production environments.
  • Build and optimize high-performance model serving pipelines using vLLM and other modern inference frameworks.
  • Improve inference latency, throughput, memory utilization, and overall system performance across GPU clusters.
  • Deploy, monitor, and manage open-source and proprietary LLMs while ensuring reliability and scalability.
  • Optimize GPU resource allocation, batching strategies, caching mechanisms, and model parallelism for efficient inference.
  • Collaborate with Machine Learning Engineers to transition trained models into production-ready inference services.
  • Develop APIs, microservices, and deployment workflows for AI-powered applications.
  • Implement monitoring, logging, benchmarking, and performance profiling for inference workloads.
  • Troubleshoot production issues related to model serving, infrastructure, and distributed inference systems.
  • Contribute to the continuous improvement of AI platform architecture, automation, and deployment best practices.

What Makes You a Great Fit

  • 5+ years of experience in Machine Learning Infrastructure, AI Platform Engineering, MLOps, or Inference Engineering.
  • Strong hands-on expertise with Large Language Models (LLMs) and vLLM for production-scale inference.
  • Experience deploying and optimizing transformer-based models using modern inference frameworks.
  • Solid understanding of GPU computing, CUDA, distributed inference, model quantization, and performance optimization techniques.
  • Experience with Python and deep learning frameworks such as PyTorch, Hugging Face Transformers, or similar ecosystems.
  • Knowledge of containerization and orchestration technologies including Docker and Kubernetes.
  • Familiarity with cloud platforms such as AWS, Azure, or GCP for AI infrastructure deployment.
  • Strong understanding of REST APIs, microservices, distributed systems, and scalable backend architectures.
  • Experience implementing monitoring, benchmarking, and observability for AI inference workloads.
  • Excellent analytical, debugging, and problem-solving skills with the ability to optimize complex AI systems.
  • Strong collaboration and communication skills, with the ability to work effectively across AI, platform, and product engineering teams.
  • Passion for advancing generative AI infrastructure and delivering reliable, high-performance inference solutions for real-world applications.
About this role

Summary

Build, optimize, and scale high-performance inference platforms for large language models.

Job title

Inference Engineer

Experience level

5+ years

Minimum experience

5+ years exp

Industry

software

Location requirements

Bengaluru, India; remote work not specified

Salary

$2400k–$3600k

Management role

No

Skills & keywords

Required skills

Large Language ModelsvLLMPythonPyTorchCUDADockerKubernetes

Preferred skills

cloud platformsREST APIsmicroservicesmonitoringperformance optimization

Specializations

LLMsvLLMinferenceGPUAI
Locations

Structured locations inferred from the posting.

Bengaluru, Karnataka, India

On-site City