LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned an...
Alessandro Bondielli, Lucia C. Passaro, D. Bacciu et al.· 0 citations
A hybrid deep learning framework for learned one-step tracklet filtering of radar measurements that achieves lower filtering error and improved robustness under severe non-Gaussian disturbances, compared with EKF- and UKF-based analytical baselines.
A. Cabras, Niccolò Pilloni, Victor Mustieles-Pérez et al.· AI Sensors· 0 citations
This work proposes Z-PEFT, a lightweight meta-classifier that relies exclusively on layer-wise spectral measures for classification, and achieves the best performance while maintaining low and scalable computational cost among weight-space detectors.
Nicola Pitzalis, Donald Shenaj, Giacomo Cignoni et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.