Agent frameworks increasingly delegate work by forking sub-agents; a common default makes the child inherit the parent's full working context. We measure how the effect of inherited state changes with capability, where $C_m$ denotes clean fork-fresh accuracy. We compare 3 inheritance policies: Reset (fork fresh: base e...
Whether off-the-shelf SLMs meet practitioner-defined thresholds and, when they fail, why, and whether quantization changes the answer are asked, and an eligibility gap is found.
This work evaluates a frozen, closed-set, action-scored benchmark with 2 suites that represent 2 different meanings of "no memory", finding that at the 3 smaller scales, models trust a stale document more than a stale memory; at 8B, the difference is not significant.
Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional precision budget. Using causal mixed-precision intervention as ground...
Jun-Hao Hu, S. Ramachandran· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.