TLDRocket
Sign in

Serving DeepSeek-V4: why million-token context is an inference systems problem

Together AI

DeepSeek-V4's million-token context window relies on architectural changes using compressed sparse attention, heavily compressed attention, and sliding window attention that reduce key-value cache requirements. Together achieved 3.7M tokens of capacity on an NVIDIA HGX B200 node through cache management policies, compared to 1.2M without optimization. Serving V4 efficiently requires inference engines to handle multiple cache types, implement context-aware prefix caching policies, and choose endpoint configurations matched to workload characteristics—making long-context serving primarily a systems problem rather than just a model capability.

Why it matters

DeepSeek-V4 makes million-token context a serving-systems problem. Together AI explores the inference work behind V4 on NVIDIA HGX B200, including compressed KV layouts, prefix caching, kernel maturity, and endpoint profiles for long-context workloads.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.