← All work
Python
llm-eval-harness
Deterministic evaluation of model answers: exact match, contains, token F1, ROUGE-L, and retrieval metrics, plus an optional LLM judge and a report renderer.
llm-eval-harness
Highlights
- Deterministic metrics: exact match, F1, ROUGE-L
- Retrieval metrics
- Optional LLM judge
- Report renderer
Want the code?
View on GitHub