Provisional

Benchmark studies

This page reports what has actually been executed so far, against the design of the finished study. Nothing here is a completed benchmark result. The numbers will change as further models and repeats are run, and a partial result should not be read as a finding.

The study design

This is the target the study is being run against. It is fixed in advance so the results below can be read as a fraction of a known whole rather than as an open-ended collection of runs.

Why repeats. A single pass confounds run-to-run variation with model-to-model variation: if two models differ by four points on one execution each, that difference is not separable from the noise of the scorer. Several executions per model make the two sources of variation distinguishable. Until those repeats exist, every number on this page is a single draw.

Completion

doneexecuted and published partialstarted, not complete pendingnot yet run

Results

Loading the published report…

Provenance

What produced the numbers above, and how to get the data behind them.