Auditing a KB Elicitation of Frontier LLM Knowledge: A Multi-dimensional Analysis of GPTKB v1.5
It is found that the models'factual knowledge differs quite significantly from established knowledge bases, and that its accuracy is significantly lower than indicated by previous benchmarks, shedding light on future research opportunities in neuro-symbolic AI concerning extraction, consolidation and verification of fa...