← All work

Python

llm-eval-harness

Deterministic evaluation of model answers: exact match, contains, token F1, ROUGE-L, and retrieval metrics, plus an optional LLM judge and a report renderer.

CI passing · Python

llm-eval-harness

Highlights

  • Deterministic metrics: exact match, F1, ROUGE-L
  • Retrieval metrics
  • Optional LLM judge
  • Report renderer

Want the code?

View on GitHub