Tool · Evaluation tool · EleutherAI
lm-evaluation-harness
A framework for few-shot evaluation of language models.
- License
- PermissiveMIT License
- Language
- Python
- Latest release
- 31 Aug 2026lm-evaluation-harness v0.4.13
- Measurements
- 0
Relationships
- BenchmarkGPQA DiamondImplements
- BenchmarkIFEvalImplements
- BenchmarkMMLU-ProImplements
Releases & news
- lm-evaluation-harness v0.4.13A fix-focused release. The main fixes are for few-shot leakage, a multiple-choice filter bug, and group stderr, alongside two new ONNX backends and eight new benchmark suites. Also updated most configs for datasets>=4, which accounts for much of the diff by volume.
- lm-evaluation-harness v0.4.12New release with four new model backends, tensor parallel support for transformers based models (hf), new benchmarks, a TaskManager refactor, and a long tail of task correctness fixes.
- lm-evaluation-harness v0.4.11Minor release. Stay tuned for bigger changes next release.
- lm-evaluation-harness v0.4.10The big change this release: the base package no longer installs model backends by default. We've also added new benchmarks and expanded multilingual support.
- lm-evaluation-harness v0.4.9.2This release continues our steady stream of community contributions with a batch of new benchmarks, expanded model support, and important fixes. A notable change: Python 3.10 is now the minimum required version.
- lm-evaluation-harness v0.4.9.1This v0.4.9.1 release is a quick patch to bring in some new tasks and fixes. Looking aheas, we're gearing up for some bigger updates to tackle common community pain points. We'll do our best to keep things from breaking, but we anticipate a few changes might not be fully backward
- lm-evaluation-harness v0.4.9Enhanced Backend Support:
- lm-evaluation-harness v0.4.8New Backend Support:
- lm-evaluation-harness v0.4.7This release includes several bug fixes, minor improvements to model handling, and task additions.
- lm-evaluation-harness v0.4.6This release brings important changes to chat template handling, expands our task library with new multilingual and multimodal benchmarks, and includes various bug fixes.
Reviews
Contributions are not open yet, so there is nothing here from members.
Sources & history
Sources
- GitHub (live source, 2 records, 13 Sept 2026)
- Mutinai illustrative fixtures (illustrative fixture, 1 record, 13 Sept 2026)
External identifiers
Field history
- homepageUrl = "https://www.eleuther.ai" — GitHub (current)
- homepageUrl = "https://www.eleuther.ai" — GitHub
- primaryLanguage = "Python" — GitHub (current)
- primaryLanguage = "Python" — GitHub
- summary = "A framework for few-shot evaluation of language models." — GitHub (current)
- summary = "A framework for few-shot evaluation of language models." — GitHub