CoWindow and MassAlloc Attention: collective causal coverage and distribution-adaptive compute [R]
I m one of the authors of two recent papers exploring different sources of redundant computation in attention. I d like to share the ideas and hear feedback from people working on long-context models
At a glance
- reddit.com: CoWindow and MassAlloc Attention: collective causal coverage and distribution-adaptive compute [R]
- arxiv.org: CoWindow Attention: Full Causal Coverage Is a Collective Property
The story
reddit.com: I m one of the authors of two recent papers exploring different sources of redundant computation in attention. I d like to share the ideas and hear feedback from people working on long-context models and attention kernels. CoWindow Attention (CoWA) distributes distant context across KV heads using complementary windows, while sharing local and prefix-sink windows. Each head attends sparsely, but the union of their visible positions covers the full causal history. The pattern is position-defined and requires no learned router or indexer. Paper: https://arxiv.org/abs/2609.32704 MassAlloc Attention (MALA) retains full causal QK scoring, then uses attention s own softmax statistics to decide whether to execute subsequent computation for a tile. It reduces low-contribution post-score work, using a common tolerance across training and inference. Paper: https://arxiv.org/abs/2609.32712 Both support training forward/backward and inference prefill/decoding. At 128K tokens on 8 H100 GPUs with TP=8, attention-operator speedups relative to FullAttn are: Method Forward Backward Decode CoWA 7.4x 8.6x 3.0x MALA 2.2x 3.0x 1.6x These measurements are for the attention operators, not end-to-end model speedups. We evaluated scaling from 0.6B to 14B and conducted separate continued-training experiments at 32B. At 14B with 32K context, total training FLOPs decreased by 28.5% for CoWA and 23.1% for MALA, with model capabilities comparable to FullAttn on the reported evaluations. Two distinctions that matter: collective coverage does not imply identical head-wise interactions or outputs to FullAttn, and MALA still pays for full causal QK scoring. Neither result establishes universal lossless equivalence to dense attention. I d be interested in feedback on workloads that might stress collective coverage, or attention distributions where adaptive post-score allocation could be less effective. Happy to discuss implementation and evaluation details. submitted by /u/BitExternal4608 [link] [com
arxiv.org: arXiv:2609.32704v1 Announce Type: new Abstract: FullAttn repeatedly exposes the complete causal history to every attention head, creating substantial redundant computation and memory traffic even with IO-efficient dense kernels. We introduce CoWA, a structured attention architecture that distributes access to the causal history across KV heads. All heads share near-diagonal and prefix-sink windows, while complementary long-range windows partition the remaining history. Their union provides full causal coverage although each head attends sparsely to distant tokens. This position-defined attention pattern requires no learned router or indexer, is used consistently during training and inference, and aligns with KV-head tensor parallelism. A window-matched ablation at 8K isolates the effect of complementary long-range allocation: CoWA with 100% collective coverage reaches 89.73% accuracy, compared with 89.97% for FullAttn, while duplicated long-range windows perform substantially worse. Across a broader controlled associative-recall comparison with matched token budgets, CoWA closely tracks FullAttn as the context grows, whereas other sparse patterns lose a substantial fraction of the associations. In an attention-operator benchmark at 128K tokens with tensor parallelism, CoWA reduces forward and backward latency during training by 7.4x and 8.6x and decoding latency during inference by 3.0x over FullAttn. Its per-rank peak operator memory matches FullAttn during training and is 7.6x lower during decoding. Across scaling-law training from 0.6B to 14B parameters, CoWA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results show that full causal coverage can be a collective property of the head ensemble rather than a duplicated property of every head.