Researchers publish studies on methods to improve LLM evaluation and to measure internal inconsistencies in large language model uncertainty
Research publication Provisional 35% confidence first seen
One study proposes a dependence-aware approach for aggregating multiple LLM “judges” so that agreement is adjusted for correlated biases among the judges. Another study introduces a metric for how closely an LLM’s internal probability updates follow Bayes rule, quantifying deviations from consistent Bayesian reasoning.