A new benchmark, KIFI, is designed, which comprises 1032 carefully selected instances from the TRUE and ScreenEval datasets, with key information annotated, and it is shown that LLMs frequently fail to use the appropriate information to make correct decisions.
Xindi Guo, Zhen Xie, Patrick H. Chen· Annual International ACM SIG...· 0 citations