I was recently working with some bagging and boosting methods and realized that current LLM evaluation has quite a lot in common with them.
We increasingly use LLMs to evaluate the output of other LLMs. But a single LLM judge can be biased or worse, confidently wrong (especially when using smaller models). One obvious response is to ask several judges and combine their verdicts. Another is to give smaller evaluators narrow specialties: one checks factual accuracy, another instruction-following, another style, and another safety.
Put like that, these new evaluation systems start to look remarkably similar to classical ensemble learning.
A jury is an ensemble
In bagging, we train several versions of a predictor and aggregate their outputs. The models will not make exactly the same "mistakes", so averaging or voting can produce a more stable result than trusting any single one.
An LLM judging panel follows the same broad intuition. Several judges evaluate the same response independently, after which their scores or votes are combined.
This is better described as a voting ensemble than literal bagging. Proper bagging involves bootstrap samples of training data; an LLM jury does not train anything. Still, the family resemblance is hard to miss: create variation, aggregate the results, and reduce the influence of one unstable prediction.
The idea already has a wonderfully direct name. In Replacing Judges with Juries, researchers proposed a Panel of LLM Evaluators, or PoLL. Their panel of smaller, diverse models performed better than one large judge in their experiments, while also reducing cost and some forms of model bias.
A jury asks the same question several times and trusts the aggregate.
Small specialists are not automatically boosting though
My second instinct was to call a collection of tiny, specialized evaluators a form of boosting. But is it really?
In boosting, weak learners are added sequentially. Each new learner focuses on the examples, or residual errors, that the existing ensemble still gets wrong. The important property is that the next learner is chosen in response to the mistakes that remain.
If one LLM always checks facts and another always checks tone, that is closer to decomposed evaluation or a mixture of experts. Each critic owns a part of the problem. A final evaluator can combine their findings, much like a meta-model in stacking.
To make it boosting-like, the evals need to be adaptive. For example, a first critic finds obvious problems. The response is revised. A second critic is then asked to inspect what the first pass missed. Later critics spend their effort on unresolved weaknesses rather than repeating a complete review from scratch.
A boosting-style evaluator asks what the previous evaluators still failed to notice.
Old ideas at inference time
The analogy is not perfect. Classical bagging and boosting are training algorithms, while most LLM evaluation panels are inference workflows built from prompts and model calls. Debate is not necessarily boosting. A panel is not necessarily bagging.
But I find the lens useful. It gives us a vocabulary for asking better design questions.
- Are the judges diverse enough to make different mistakes?
- Are their judgments independent, or do they simply reproduce the same bias?
- Should every evaluator inspect everything, or should some specialize?
- Does a later critic target the "residual errors" from earlier rounds?
- Are we voting, routing, stacking, or actually boosting?
Much of the current conversation treats multi-LLM evaluation as a new organizational problem: should the models vote, debate, critique, or defer to a final judge? Another way to see it is that we are rebuilding familiar ensemble methods at inference time, only now the base learners speak in paragraphs.