Benchmark Answers Leak into LLMs, Inflating Test Scores
New analysis highlights how answers from common AI benchmarks can unintentionally appear in training data, making models appear smarter than they are. This leakage undermines the validity of performance comparisons across large language models. The finding matters because it calls into question the reliability of current evaluation methods in AI development.
Sources (1)
technology