As agents rapidly evolve, existing benchmarks can become saturated, limiting their ability to distinguish capabilities and reveal remaining failure modes. Particularly in scientific domains, constructing and updating benchmarks requires substantial time, labor, and domain expertise, making it difficult to keep evaluati...
Dongki Kim, Namkyeong Lee, Surag Nair et al.· 0 citations
Many biological discovery problems require experiments to be selected sequentially under constrained budgets. CRISPR screening is a prominent example, as exhaustive perturbation testing is often infeasible and candidate perturbations must instead be prioritized over multiple experimental rounds. Despite the importance...
Carl N. Edwards, E. De Brouwer, Xiner Li et al.· 0 citations
UVE-PCoT improves affinity prediction and cue-level explanation over general-purpose multimodal large language models and ablations and operationalizes this perceptual-conflict-oriented perspective into an interpretable framework, advancing UVE evaluation from black-box scoring to explanatory analysis and providing cue...
Xiner Li, Yi Xiao, Jinhao Qiao et al.· Electronics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.