IndexCache, a new sparse attention optimizer, delivers 1.82x faster inference on long-context AI models
Processing 200,000 tokens through a large language model is expensive and slow: the longer the context, the faster the costs spiral. Researchers at Tsinghua University and Z.ai have built a technique called IndexCache that cuts up to 75% of the redundant computation in sparse attention models, delivering up to 1.82x faster time-to-first-token and 1.48x faster generation throughput at that context length.The techniqu…
AI brief
Pulse reads the full article- What happened
- Why it matters
- What to watch
Sign in to get the AI brief. Pulse explains what happened, why it matters, what to watch and who's exposed. Free for members.
Sign in to read the briefMore like this





