Zipf-like Statistical Regularities in Molecular Sequence Representations for Chemical Language Models
Large language model (LLM)-based approaches increasingly use molecular strings such as SMILES and SELFIES for molecular generation and property prediction. However, the statistical properties of molecular token distributions have not been systematically characterized. Here, we analyzed rank–frequency distributions of...