We're introducing our latest research paper HydraHead, a new attention hybridization architecture that fuses Full Attention and Linear Attention at the head level.
Motivated by insights from mechanistic interpretability, HydraHead treats the attention head—not the layer—as the natural granularity for attention hybridization to build more efficient long-context models.
A short thread 🧵
显示更多