Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs

Apply to this job
San Francisco, CA on site Until 8/22/2026 First posted March 18, 2026 Last posted March 18, 2026
Job description

About Us

Gimlet is building the first multi-silicon neocloud designed for fast, efficient inference.

As AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.

Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.

We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI. This gives our team access to systems research problems grounded in frontier models, cutting-edge production workloads, and emerging hardware architectures.

About the role

At Gimlet, we believe every hire changes the company.

As a an early-stage company, talent density matters more than headcount. The engineers we hire today will shape the systems, culture, and standards that define Gimlet for years to come.

The future of AI infrastructure will not be built on a single hardware platform. It will be built on software capable of extracting maximum performance from increasingly diverse compute architectures.

Kernel engineers sit at the center of that challenge.

This role is an opportunity to help build the execution layer that transforms theoretical hardware performance into production reality.

You will work close to accelerators and execution hardware, designing, optimizing, and validating kernels that power large-scale AI workloads across both established and emerging architectures.

This is not a traditional GPU optimization role.

We are building systems that must operate efficiently across heterogeneous hardware environments, where performance, efficiency, and correctness directly influence the economics of AI infrastructure.

What success looks like

In the first 12-18 months, you will help:

  • Build and optimize kernels that improve latency, throughput, and hardware utilization for production AI workloads

  • Develop execution strategies that unlock performance across both established and emerging accelerator architectures

  • Improve memory efficiency, scheduling behavior, and execution characteristics across the inference stack

  • Partner with compiler, runtime, and distributed systems engineers to ensure end-to-end performance optimization

  • Influence how heterogeneous hardware is deployed and utilized within the next generation of AI infrastructure

  • Help establish performance engineering standards that shape the future of Gimlet's execution platform

You may be a good fit if

  • Strong software engineering fundamentals

  • Experience working on performance-critical systems close to hardware

  • Comfort reasoning about low-level execution behavior, memory hierarchies, and performance tradeoffs

  • Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience.

Strong candidates may also have

  • Experience with CUDA, Triton, CUTLASS, or other accelerator programming models

  • Deep understanding of GPU execution models (warps/wavefronts, blocks, grids)

  • Experience optimizing memory access patterns (coalescing, shared memory, cache behavior)

  • Familiarity with occupancy, latency hiding, and instruction-level parallelism

  • Experience using profiling and performance analysis tools

  • Familiarity with multi-GPU or distributed execution is a plus

Why join now?

Gimlet is at the very beginning of its journey, and that's what makes this moment special. Most AI infrastructure companies are focused on deploying more compute. We are focused on making increasingly diverse compute work together, and that ambition touches every part of how we build and run this company.

As an early member of the team, you will have significant ownership over your work, partner directly with a small group of highly capable people, and help shape not just what we build, but how we scale the company.

We value people who are excited to work across domains, take ownership of meaningful problems, and help define what Gimlet becomes over the next several years.

Agency Policy: Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.

Compensation (from employer):
$150K – $350K • Offers Equity

About this role

Summary

Design, optimize, and validate kernels for AI workloads across hardware architectures

Job title

Member of Technical Staff - Kernels & GPU Performance

Experience level

null

Industry

software

Location requirements

San Francisco, CA, no remote work allowed

Salary

$150k–$350k

Management role

No

Skills & keywords

Required skills

software engineeringperformance-critical systemslow-level executionmemory hierarchies

Preferred skills

CUDATritonCUTLASSGPU execution modelsprofiling tools

Specializations

performance optimizationGPUkernelshardware architecturesperformance analysis
Locations

Structured locations inferred from the posting.

San Francisco, CA, USA

On-site City