Recurrent Transformers increase computational depth through temporal recurrence, feeding each token's high-level hidden state into the computation of the next. This raises a natural question: is additional computation better spent on more temporal steps or greater physical depth? We investigate this question using Late...
Ze-Yi Huang, Xuehai He, Yong Jae Lee et al.· 0 citations
Latent Recurrent Transformer is studied, a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden state from the previous token as recurrent memory for the next token, while retaining one-forward-per-token decoding with 9% latency overhead over the standard Transformer.
The TaSe (Talk in Pieces, See in Whole) framework is introduced with three main contributions: a hierarchical synthetic captioning dataset spanning three tiers from category names to descriptive sentences; the three-component disentanglement module guided by a novel disentanglement loss function, transforms text embedd...
Sojung An, Kwanyong Park, Yong Jae Lee et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.