Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
Hierarchical Memory Mamba integrates a lightweight working memory that extracts slow paragraph-level semantics from the fast sensory memory embedded in the backbone's hidden states and endows HMM cross-task generalization through parametric learning, which is not observed in other long-context enhanced Mamba variants.