How Grammarly Scales Real-Time Inference with vLLM on CoreWeave
About this session
At Superhuman (formerly Grammarly), a fine-tuned language model corrects grammatical errors for millions of monthly users, served with vLLM on CoreWeave's dedicated inference. Consumer scale stresses everything: throughput, latency, correctness, and cost. This talk covers how the team pressure-tests a serving setup against those requirements, vLLM tuning, and where dedicated inference and fully managed inference platforms converge or diverge. Beyond benchmarks, the team also validates in production with A/B tests. Expect an overview of how the model works and how serving is architected at Superhuman.
Speakers
Share this session


Get in the room. San Francisco, September 29.
Fully Connected 2026 is where the engineers, leaders, and operators running AI in production come together for three days of depth, access, and real conversation.