Back to full agenda
Track 2: Push Models Further
Breakout session

How Grammarly Scales Real-Time Inference with vLLM on CoreWeave

Session
208
Day & Date
Wednesday, Sep. 30, 2026
Time
Location
Moscone South, Level 3, Room 302

About this session

At Superhuman (formerly Grammarly), a fine-tuned language model corrects grammatical errors for millions of monthly users, served with vLLM on CoreWeave's dedicated inference. Consumer scale stresses everything: throughput, latency, correctness, and cost. This talk covers how the team pressure-tests a serving setup against those requirements, vLLM tuning, and where dedicated inference and fully managed inference platforms converge or diverge. Beyond benchmarks, the team also validates in production with A/B tests. Expect an overview of how the model works and how serving is architected at Superhuman.

Speakers

CS
Christoph Stuber
Software Engineer - ML Platform, Grammarly

Share this session

Get in the room. San Francisco, September 29.

Fully Connected 2026 is where the engineers, leaders, and operators running AI in production come together for three days of depth, access, and real conversation.