랭킹으로 돌아가기

huggingface/evaluation-guidebook

Jupyter Notebook

Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

evaluationevaluation-metricsguidebooklarge-language-modelsllmmachine-learningtutorial
스타 성장
스타
2.1k
포크
125
주간 성장
이슈
4
1k1.5k2k
2024년 10월2025년 5월2025년 12월2026년 7월
README

The LLM Evaluation guidebook ⚖️

!! THIS GUIDEBOOK IS NO LONGER MAINTAINED. THE LATEST AND MOST UP TO DATE VERSION (as of Dec 2025) OF IT LIVES HERE: https://huggingface.co/spaces/OpenEvals/evaluation-guidebook


If you've ever wondered how to make sure an LLM performs well on your specific task, this guide is for you! It covers the different ways you can evaluate a model, guides on designing your own evaluations, and tips and tricks from practical experience.

Whether working with production models, a researcher or a hobbyist, I hope you'll find what you need; and if not, open an issue (to suggest ameliorations or missing resources) and I'll complete the guide!

How to read this guide

  • Beginner user: If you don't know anything about evaluation, you should start by the Basics sections in each chapter before diving deeper. You'll also find explanations to support you about important LLM topics in General knowledge: for example, how model inference works and what tokenization is.
  • Advanced user: The more practical sections are the Tips and Tricks ones, and Troubleshooting chapter. You'll also find interesting things in the Designing sections.
  • User coming back to the site: Every year I do a dive on a topic, check them out!

In text, links prefixed by ⭐ are links I really enjoyed and recommend reading.

Table of contents

If you want an intro on the topic, you can read this blog on how and why we do evaluation!

Automatic benchmarks

Human evaluation

LLM-as-a-judge

Troubleshooting

The most densely practical part of this guide.

General knowledge

These are mostly beginner guides to LLM basics, but will still contain some tips and cool references! If you're an advanced user, I suggest skimming to the Going further sections.

Yearly dives

Resources

Links I like

Community translations

This guide has been kindly community translated!

Thanks

This guide has been heavily inspired by the ML Engineering Guidebook by Stas Bekman! Thanks for this cool resource!

Many thanks also to all the people who inspired this guide through discussions either at events or online, notably and not limited to:

  • 🤝 Luca Soldaini, Kyle Lo and Ian Magnusson (Allen AI), Max Bartolo (Cohere), Kai Wu (Meta), Swyx and Alessio Fanelli (Latent Space Podcast), Hailey Schoelkopf (EleutherAI), Martin Signoux (Open AI), Moritz Hardt (Max Planck Institute), Ludwig Schmidt (Anthropic)
  • 🔥 community users of the Open LLM Leaderboard and lighteval, who often raised very interesting points in discussions
  • 🤗 people at Hugging Face, like Lewis Tunstall, Hynek Kydlíček, Guilherme Penedo and Thom Wolf, and of course my teammate Nathan Habib with whom I've been doing evaluation and leaderboards since 2022

and of course to all the contributors :)

Citation

CC BY-NC-SA 4.0

@misc{fourrier2024evaluation,
  author = {Clémentine Fourrier and The Hugging Face Community},
  title = {LLM Evaluation Guidebook},
  year = {2024},
  journal = {GitHub repository},
  url = {https://github.com/huggingface/evaluation-guidebook)
}
관련 저장소
langfuse/langfuse

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23

TypeScriptnpmOtheranalyticsllm
langfuse.com
31.6k3.3k
mlflow/mlflow

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.

PythonPyPIApache License 2.0machine-learningai
mlflow.org
27.1k6k
promptfoo/promptfoo

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

TypeScriptnpmMIT Licensellmprompt-engineering
promptfoo.dev
23.5k2.1k
comet-ml/opik

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

PythonPyPIApache License 2.0open-sourcelangchain
comet.com/docs/opik/
20.8k1.6k
Tencent/WeKnora

Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.

GoGo ModulesOtheragentagentic
weknora.weixin.qq.com
18.7k2.6k
vibrantlabsai/ragas

Supercharge Your LLM Application Evaluations 🚀

PythonPyPIApache License 2.0llmllmops
docs.ragas.io
14.9k1.6k
mrgloom/awesome-semantic-segmentation

:metal: awesome-semantic-segmentation

semantic-segmentationbenchmark
10.8k2.5k
oumi-ai/oumi

Easily fine-tune, evaluate and deploy Gemma 4, Qwen3.5, Qwen3.6, gpt-oss, DeepSeek-R1, or any open source LLM / VLM!

PythonPyPIApache License 2.0dpoevaluation
oumi.ai
9.4k783
explodinggradients/ragas

Supercharge Your LLM Application Evaluations 🚀

PythonPyPIApache License 2.0llmllmops
docs.ragas.io
8.4k862
open-compass/opencompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

PythonPyPIApache License 2.0evaluationbenchmark
opencompass.org.cn
7.2k814
Helicone/helicone

🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓

TypeScriptnpmApache License 2.0large-language-modelsprompt-engineering
helicone.ai
6k632
coze-dev/coze-loop

Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to monitoring.

GoGo ModulesApache License 2.0agentai
5.6k777