Evaluation
Loading…
A RAG system should be measured, not assumed. Each retrieval strategy is scored on 30 labelled questions over the fictional demo documents — including four questions whose answer is deliberately not in the documents.
Loading results…
Runs 7 demo questions through the full Ask pipeline with your language model and measures faithfulness (sentences supported by cited passages), answer relevance, context precision and correct refusals — no paid “LLM judge” needed. These are proxies, not perfect judges.