The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure
The results show that models often do not affirm culmination but nevertheless accept the corresponding simple-past hypothesis, a pattern the authors characterize as Sufficiency Bias, and show that prompting interventions produce a Decision Shift among labels without reliably improving the underlying semantic understand...