Evaluating reasoning-tuned large language models for clinical decision-making in spine surgery.
PURPOSE Most clinical evaluations of large language models assess factual recall rather than the multi-step reasoning behind operative plans. Reasoning-tuned models, post-trained to generate explicit intermediate reasoning before answering, may better approximate surgical decision-making. We compared two such models fr...