A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person's record, and such records are private, so benchmarks generate the person and the questio...
Arman Behnam, Sunglyoung Kim, Liang-Wei Yang· 0 citations
Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all; across four models and six reasoning tasks, it matches the strongest baseline in task performance while delivering 32-43% higher throughput than that method when deployed with vLLM.
Heng-Yi Wang, Jie-Lin Qiu, Wenting Zhao et al.· 0 citations
It is suggested that inference-time harness design is an important complement to conventional training-time distillation, enabling strong models to transfer cognitive structure to weaker models without retraining.
Cheng Qian, Wenting Zhao, Liang-Wei Yang et al.· 5 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.