Research Scientist - Audio

San Francisco Bay Area on site Until 8/21/2026 First posted April 14, 2026 Last posted April 14, 2026
Job description

ABOUT RETELL AI

Retell AI is using first-principles thinking to reimagine the call center with cutting-edge voice AI. Thousands of companies now use Retell's AI voice agents to handle sales, support, and logistics calls that once required large teams of human agents. Backed by Y Combinator, Alt Capital, and other leading investors, we've scaled to $80M in ARR with a team of 50, up from $5M at the start of 2025, and are now valued at over $1.5B.

Our vision for 2026 is to build a modern CX platform where entire contact centers are powered by AI. Instead of basic automation that needs constant human tuning, we're creating intelligent AI “workers” that act as frontline agents, QA analysts, and managers, continuously executing, monitoring, and improving every customer interaction.

We're growing fast and looking for ambitious builders who want to tackle hard technical problems, move quickly, and have a real impact on one of the fastest-growing voice AI companies in the world.

Let's build the future together.

Recent recognition:

 

ABOUT THE ROLE

Retell AI transforms customer experience with voice AI for enterprises, including customers like CVS/Aetna, American Airlines, Lenovo, and Grab. We have more customer stories than we can tell!

This is a research-driven, high-impact role for ML researchers who want to push the boundaries of real-time AI. As a Founding Machine Learning Research Engineer at Retell, you’ll focus on advancing model capabilities for human-like voice agents operating in complex, real-world environments.

You’ll explore new approaches across LLMs and audio models, design novel evaluation methods, and prototype systems that improve reasoning, latency, and conversational quality. Your work will directly influence production systems, bridging cutting-edge research with real-world deployment.

If you’re excited about solving open-ended ML problems, experimenting rapidly, and shaping how voice AI systems think and perform, this is a unique opportunity to do so at scale.

KEY RESPONSIBILITIES

  • Research & Experimentation – Explore and develop new techniques across LLMs and audio models to improve reasoning, latency, and conversational quality in real-time systems.

  • Model Training – Rapidly build and iterate on models and pipelines, turning research ideas into working prototypes. Innovate on paradigms, training methods, and inference.

  • Evaluation & Benchmarking – Design novel evaluation frameworks, datasets, and metrics to measure performance on complex, real-world voice tasks.

  • Bridge Research to Production – Collaborate closely with engineering to translate research insights into deployable systems.

  • Human Feedback Loops – Develop methods to incorporate human evaluation into model improvement, especially for subjective conversational quality.

  • Advance the Frontier – Stay at the cutting edge of ML research and bring new ideas into Retell’s product and infrastructure.

REQUIRED

  • Strong ML Research Background – You've worked on advanced ML problems (like LLM pre-training and post-training, transcription model training, TTS, or multimodal systems), either in industry or academia.

  • Deep Technical Foundation – Comfortable with PyTorch, model architectures, and the math behind modern machine learning.

  • Top Academic Background – Master's degree in CS, ML, AI or related field required; PhD preferred. Equivalent research-level engineering experience also considered.

 

YOU MIGHT THRIVE IF YOU

  • Published or Awarded – First/co-author publications at top-tier venues (NeurIPS, ICML, ICLR, ACL, Interspeech, etc.) or notable competition awards are a strong plus.

  • Experimental Mindset – You enjoy exploring open-ended problems and iterating quickly on ideas.

  • Bridge Theory & Practice – You can translate research into systems that work in real-world environments.

  • Startup-Ready – You thrive in fast-paced environments with high ownership and ambiguity.

  • Collaborative & Clear Communicator – You can explain complex ideas and work cross-functionally to drive impact.

JOB DETAILS

  • Cash: $225,000 - $400,000 base salary

  • Equity: Offers Equity

  • Location: Redwood City, CA, US (100% Relocation Provided)

  • US Visas: Retell AI is open to sponsoring work authorization for qualified candidates, including H1B/H-1B, TN, L-1, E-3, F-1 (OPT/CPT).

OTHER BENEFITS

  • 100% coverage for medical, dental, and vision insurance

  • $70/day DoorDash credit for unlimited meals and snacks

  • $200/month wellness reimbursement

  • $300/month commuter reimbursement

  • $75/month phone bill reimbursement

  • $50/month internet reimbursement

COMPENSATION PHILOSOPHY

  • Best Offer Upfront: Choose from three cash-equity balance options, no negotiation needed

  • Top 1% Talent: Above-market pay (top 5 percentile)

  • High Ownership: Small teams, >$1M revenue/employee, significant equity

  • Performance-Based: Offers tied to interview performance, not past salaries

INTERVIEW PROCESS

  • Talent Screen (15min): chat with our recruiter to get a better sense of the role, the team, and what it’s like to work here.

  • Technical Interview (45 min): LLM theory specific coding Interview (PyTorch)

  • Technical Interview (45 min): Live Practical Systems Design and Coding Interview.

  • Onsite/Virtual Interviews (3 hrs): Hosted in our office if located in the Bay Area or virtual, with three rounds:

    • ML System Design

    • ML Question Deep Dive

    • Backend + AI Practical

Compensation (from employer):
$225K – $400K • Offers Equity

About this role

Summary

Research and develop voice AI models, improve reasoning, latency, and conversational quality.

Job title

Research Scientist - Audio

Experience level

master's or PhD in CS, ML, AI

Industry

software

Location requirements

Redwood City, CA, US; remote work not specified

Salary

$225k–$400k

Management role

No

Skills & keywords

Required skills

pytorchml researchmodel architecturesmathematics

Preferred skills

published papersawardsexploring open-ended problemssystem translation

Specializations

audio modelsllmsmodel trainingevaluationml research
Locations

Structured locations inferred from the posting.

Redwood City, CA, USA

On-site City