MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents
This work presents MemRiskBench, a five-category risk taxonomy operationalized by deterministic trace grounded checks, instantiated as a 120-episode scripted benchmark with full trace logging and no LLM-as-judge on the pass/fail path, evaluated on five locally run quantized instruction-tuned models.