Kernel Engineer (Internship and Full-time)

Tilde Research

Apply to this job
San Francisco on site Until 8/22/2026 First posted March 18, 2026 Last posted March 18, 2026
Job description

Tilde Research is a moonshot AI lab advancing mechanistic interpretability, new architectures, and pretraining science. We build foundational understanding of models to advance the frontier of intelligence.


About the role:

As a Kernel Engineer at Tilde, you'll design, implement, and optimize high-performance GPU kernels that are critical to scaling our training and inference workloads. Your work will enable faster iteration cycles, higher throughput, and lower latency. You'll work closely with ML researchers and engineers to co-design models and infrastructure that are deeply performance-aware, and help push the limits of what current hardware can support.

What you might work on:

  • Design, develop, and tune custom GPU kernels for core model operations

  • Work with ML engineers to prototype and scale novel model architectures

  • Contribute to system-wide efforts to improve efficiency and throughput, beyond just kernel-level optimizations


You're a good fit if you:

  • Have experience in deep learning or related research areas

  • Have demonstrated exceptional capability in working on ML kernels. This can include:

    • Strong open source contributions

    • Thoughtful technical blog posts/work logs

    • Previous experience working on hardware-aligned algorithms

  • Deep familiarity with PyTorch, Triton/TK/TileLang (>1 of), basic familiarity with CUDA, and knowledge of GPU architecture.

  • Communicate clearly and effectively, both verbally and in writing

  • Strong algorithmic thinker

  • Are able to learn quickly

About this role

Summary

Design, develop, and optimize GPU kernels for ML workloads in a research setting.

Job title

Kernel Engineer (Internship and Full-time)

Experience level

Industry

technology

Location requirements

San Francisco, on-site work required.

Salary

Not specified

Management role

No

Skills & keywords

Required skills

deep learningPyTorchCUDAGPU architecture

Preferred skills

TritonTKTileLangopen source contributions

Specializations

GPU kernelsdeep learningCUDAPyTorchhardware algorithms
Locations

Structured locations inferred from the posting.

San Francisco, CA, USA

On-site City