HaluEval
HaluEval is a benchmark built to answer a question that matters more every day: when a large language model produces a confident, fluent answer, can it also tell you which parts are wrong? Most of the time, the answer is no — and HaluEval quantifies exactly how badly models fail at catching their own mistakes.…