AI Inference Engineer

Perplexityai

Apply to this job
New York City Palo Alto Until 8/21/2026 H-1B sponsor history First posted March 26, 2025 Last posted March 26, 2025
Job description

Perplexity is an AI-powered answer engine founded in December 2022 and growing rapidly as one of the world’s leading AI platforms. Perplexity has raised over $1B in venture investment from some of the world’s most visionary and successful leaders, including Elad Gil, Daniel Gross, Jeff Bezos, Accel, IVP, NEA, NVIDIA, Samsung, and many more. Our objective is to build accurate, trustworthy AI that powers decision-making for people and assistive AI wherever decisions are being made. Throughout human history, change and innovation have always been driven by curious people. Today, curious people use Perplexity to answer more than 780 million queries every month–a number that’s growing rapidly for one simple reason: everyone can be curious. 

We are looking for an AI Inference engineer to join our growing team. Our current stack is Python, Rust, C++, PyTorch, Triton, CUDA, Kubernetes. You will have the opportunity to work on large-scale deployment of machine learning models for real-time inference.

Responsibilities

  • Develop APIs for AI inference that will be used by both internal and external customers
  • Benchmark and address bottlenecks throughout our inference stack
  • Improve the reliability and observability of our systems and respond to system outages
  • Explore novel research and implement LLM inference optimizations

Qualifications

  • Experience with ML systems and deep learning frameworks (e.g. PyTorch, TensorFlow, ONNX)
  • Familiarity with common LLM architectures and inference optimization techniques (e.g. continuous batching, quantization, etc.)
  • Understanding of GPU architectures or experience with GPU kernel programming using CUDA

The cash compensation range for this role is $190,000 - $250,000.

Final offer amounts are determined by multiple factors, including, experience and expertise, and may vary from the amounts listed above.

Equity: In addition to the base salary, equity may be part of the total compensation package.
Benefits: Comprehensive health, dental, and vision insurance for you and your dependents. Includes a 401(k) plan.

 

About this role

Summary

Develop APIs for AI inference and optimize machine learning models.

Job title

AI Inference Engineer

Experience level

3+ years

Industry

software

Location requirements

New York City, Palo Alto, San Francisco; remote work not allowed

Salary

$190,000 - $250,000

Visa sponsorship

H-1B sponsor history

Management role

No

Skills & keywords

Required skills

ML systemsdeep learningPyTorchTensorFlowONNXLLM architecturesinference optimizationGPU architecturesCUDA

Preferred skills

None specified

Specializations

AIinferencemachine learningdeep learningGPU
Locations

Structured locations inferred from the posting.

New York, NY, USA

On-site City

Palo Alto, CA, USA

On-site City

San Francisco, CA, USA

On-site City