How Effectively Does a Model Use Its Memory | Blog | Liquid AI

Even with the same state/cache size, models can differ significantly in how well they utilize memory—impacting recall, compression, and trainability.

We introduce Effective State-Size (ESS): A proxy metric for memory utilization.

Read the paper

At its core, many deep learning sequence models—attention, SSMs, gated convs—can be expressed as: y = T(u)u, where T(u) is an input-dependent matrix.

By extending classic signal processing results, we show that any equivalent recurrence must materialize a state whose size is at least the rank of the submatrices of T(u). We define this rank as the ESS and interpret it as a measure of the model’s memory utilization.

Our analysis of ESS reveals several key insights:

This work was accepted at ICML 2025.

For all the details refer to the paper: “ Quantifying Memory Utilization with Effective State-Size”.

Citation

If you use this work, please cite the technical report:

Rom N. Parnichkun, Neehal Tumma, Armin W. Thomas, Alessandro Moro, Qi An, Taiji Suzuki, Atsushi Yamashita, Michael Poli, and Stefano Massaroli (2025). Quantifying Memory Utilization with Effective State-Size. arXiv:2504.19561.

@inproceedings{parnichkun2025quantifying,
  title     = {Quantifying Memory Utilization with Effective State-Size},
  author    = {Parnichkun, Rom and Tumma, Neehal and Thomas, Armin W. and Moro, Alessandro and An, Qi and Suzuki, Taiji and Yamashita, Atsushi and Poli, Michael and Massaroli, Stefano},
  booktitle = {Proceedings of the 42nd International Conference on Machine Learning},
  pages     = {48276--48334},
  year      = {2025},
  volume    = {267},
  series    = {Proceedings of Machine Learning Research},
  publisher = {PMLR},
}