Member of Technical Staff

Morph
Morph

IT

San Francisco, CA, USA

USD 175k-350k / year + Equity

Posted on Aug 2, 2026

Morph

Fast Models Optimized for Coding Agents

Member of Technical Staff

$175K - $350K•0.40%•San Francisco, CA, US
Job type
Full-time
Role
Engineering, Machine learning
Experience
6+ years
Visa
US citizen/visa only
Connect directly with founders of the best YC-funded startups.
Apply to role ›
Tejas Bhakta
Founder
Tejas Bhakta
Founder

About the role

The best candidates would be top 1% at multiple parts of the inference stack.

work on PD disaggregation research

Morph builds the inference infrastructure behind the fastest open models. Our stack spans kernels, model serving, routing, autoscaling, and capacity. We are hiring a performance engineer to make the entire system faster, cheaper, and more reliable.

What you’ll do

  • Find the gap between theoretical hardware performance and production performance
  • Trace latency and throughput regressions from the API layer down to individual kernels
  • Optimize batching, scheduling, routing, quantization, and distributed execution
  • Build benchmarks and observability that make bottlenecks obvious
  • Validate that every optimization preserves model quality and correctness
  • Stack-rank opportunities and ship the highest-impact fixes yourself

You might be a fit if you

  • Have optimized complex production systems
  • Understand GPU performance, memory bandwidth, collectives, and inference serving
  • Are strong in Python and comfortable navigating unfamiliar codebases
  • Can turn profiling data into clear engineering decisions
  • Care about tokens per second, tokens per dollar, and correctness equally

You will work directly with the founders on problems that determine how efficiently frontier-scale models can be served. Small team, enormous compute, immediate production impact.

About the interview

2 day work trial

About Morph

Morph builds specialized code-generation models and serves them on a custom inference stack.

Technical work involves autoresearch for kernels and custom speculative-decoding models.