Skip to content
Open access

SLIM-KT: Transferring Large Language Model Item Semantics into a Lightweight, Cold-Start-Robust Knowledge Tracing Model

2026 · IEEE Access · Vol 14, pp. 150519-150535 · 0 citations · 38 references

Abstract

Knowledge tracing (KT) estimates a learner’s evolving mastery and underpins adaptive learning and intelligent tutoring systems. Deep KT models are fast, but they encode each question by an identity (ID) and therefore degrade sharply on newly authored questions with little response history (the cold-start problem). Large language model (LLM) methods read the question text and reduce cold-start error, yet they keep a large model on the inference path and are costly to serve online. We present SLIM-KT, a semantic-transfer framework that keeps text-based cold-start robustness at the inference cost of a small deep model. During an offline preprocessing stage, a frozen text encoder (RoBERTa or MiniLM) embeds each question stem, and a frozen instruction-tuned LLM (Qwen2.5-7B-Instruct) assigns difficulty scores and per-option adequacy labels; both are cached once per item and never recomputed. A lightweight SAKT student replaces its learnable question-ID embedding with the cached embedding (semantic representation transfer), and the LLM-derived labels can optionally be distilled into auxiliary difficulty/option heads. At test time no LLM or encoder is called: an unseen question, even one with no logged responses, needs only a one-time offline encoding of its text. On three datasets spanning rich to coarse item text, SLIM-KT improves cold-start AUC over a matched ID-based baseline by 12.4 points on DBE-KT22 and 4.1 on XES3G5M (learner-clustered bootstrap, two-sided $p\lt 0.001$ ); the gains persist at zero shot (+ 12.3 and + 4.8) and latency stays at 0.045 ms per interaction. Overall AUC also rises on the text-rich datasets (0.8033 vs. 0.7937 on XES3G5M; 0.7700 vs. 0.7605 on DBE-KT22), but a four-arm ablation shows that this overall gain is a representation-uniqueness effect—a fixed random per-item vector matches or slightly exceeds the true embeddings—whereas the content-specific claim is limited to cold start. One-time offline semantic caching is thus a practical route to cold-start-robust, deployable knowledge tracing.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.