AI Intelligence // signal over noise
← back to feed
HuggingFace Papers 7/10 signal

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

evalresearch
Summary
This paper introduces a two-level meta-rubric framework to evaluate factual completeness in open-ended generation, an aspect of factuality distinct from the more commonly measured precision. The authors instantiate this framework as GAMUT, a new benchmark with 1,813 questions grounded in wearable imagery across 10 domains, also available in a text-only version. Evaluating 14 models, the benchmark proves challenging and discriminative, with the top-performing model, Gemini 3.1 Pro, achieving a score of 58.7%. The framework uses a structured meta-rubric compiled into a machine-gradable checklist that an LLM judge scores reliably.
Problem
Current methods for evaluating the factuality of long-form generation focus predominantly on precision, measuring if a model's claims are correct. This approach fails to assess factual completeness—whether a response contains all the information it should. Measuring completeness is difficult because the required facts are often not a flat list but involve structured relationships, open-ended sets, and ordered processes that simple boolean checks cannot capture.
Method
The authors propose a two-level meta-rubric framework. A high-level, structured meta-rubric captures the organization and importance of the required content. This is then mechanically compiled into a flat checklist of binary, machine-gradable rubrics. An LLM judge scores a model's generation against this checklist. This framework was instantiated as the GAMUT benchmark, comprising 1,813 questions grounded in real wearable imagery, each with an expert-verified, evidence-backed rubric.
Details

Benchmark Performance and Characteristics:

  • The GAMUT benchmark was evaluated on 14 frontier and open-weight models.
  • The benchmark was found to be genuinely challenging, with the best score being 58.7% achieved by Gemini 3.1 Pro.
  • The evaluation framework is highly discriminative and robust to the choice of LLM judge used for scoring.

Benchmark Composition:

  • GAMUT contains 1,813 questions distributed across 10 diverse domains.
  • Each question is grounded in real wearable imagery and is paired with an evidence-backed rubric that has been verified by expert human annotators.
  • A text-only variant of the benchmark is also available, as the framework is modality-agnostic.
What's new
This work introduces a novel two-level meta-rubric framework for evaluating factual completeness, a dimension of factuality distinct from the more commonly measured precision. It also provides GAMUT, the first benchmark specifically designed to measure this structured, multi-faceted completeness in long-form generation.
Conclusion

The authors conclude that their two-level meta-rubric framework and the GAMUT benchmark successfully address the gap in evaluating factual completeness for open-ended generation. The benchmark is challenging for current models, highly discriminative, and robust, offering a new tool to assess a critical aspect of model factuality that goes beyond simple claim verification.

Don't read this site daily. Get it in your inbox.

The daily brief and Sunday deep dive — distilled, scored, and opinionated. For builders only.