Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Tool · Evaluation tool · EleutherAI

lm-evaluation-harness

A framework for few-shot evaluation of language models.

License
PermissiveMIT License
Language
Python
Latest release
31 Aug 2026lm-evaluation-harness v0.4.13
Measurements
0

Relationships

  • BenchmarkGPQA DiamondImplements
  • BenchmarkIFEvalImplements
  • BenchmarkMMLU-ProImplements

Releases & news

  1. lm-evaluation-harness v0.4.13
    A fix-focused release. The main fixes are for few-shot leakage, a multiple-choice filter bug, and group stderr, alongside two new ONNX backends and eight new benchmark suites. Also updated most configs for datasets>=4, which accounts for much of the diff by volume.
  2. lm-evaluation-harness v0.4.12
    New release with four new model backends, tensor parallel support for transformers based models (hf), new benchmarks, a TaskManager refactor, and a long tail of task correctness fixes.
  3. lm-evaluation-harness v0.4.11
    Minor release. Stay tuned for bigger changes next release.
  4. lm-evaluation-harness v0.4.10
    The big change this release: the base package no longer installs model backends by default. We've also added new benchmarks and expanded multilingual support.
  5. lm-evaluation-harness v0.4.9.2
    This release continues our steady stream of community contributions with a batch of new benchmarks, expanded model support, and important fixes. A notable change: Python 3.10 is now the minimum required version.
  6. lm-evaluation-harness v0.4.9.1
    This v0.4.9.1 release is a quick patch to bring in some new tasks and fixes. Looking aheas, we're gearing up for some bigger updates to tackle common community pain points. We'll do our best to keep things from breaking, but we anticipate a few changes might not be fully backward
  7. lm-evaluation-harness v0.4.9
    Enhanced Backend Support:
  8. lm-evaluation-harness v0.4.8
    New Backend Support:
  9. lm-evaluation-harness v0.4.7
    This release includes several bug fixes, minor improvements to model handling, and task additions.
  10. lm-evaluation-harness v0.4.6
    This release brings important changes to chat template handling, expands our task library with new multilingual and multimodal benchmarks, and includes various bug fixes.

Reviews

Contributions are not open yet, so there is nothing here from members.

Sources & history

Sources

  • GitHub (live source, 2 records, 13 Sept 2026)
  • Mutinai illustrative fixtures (illustrative fixture, 1 record, 13 Sept 2026)

External identifiers

Field history

  • homepageUrl = "https://www.eleuther.ai" GitHub (current)
  • homepageUrl = "https://www.eleuther.ai" GitHub
  • primaryLanguage = "Python" GitHub (current)
  • primaryLanguage = "Python" GitHub
  • summary = "A framework for few-shot evaluation of language models." GitHub (current)
  • summary = "A framework for few-shot evaluation of language models." GitHub