Benchmarking Faithfulness, Answer Relevance, and Context Recall with Automated Synthetic Testsets
Distributed Architecture Takeaway
Building a Retrieval-Augmented Generation (RAG) prototype takes an afternoon; validating that it doesn’t hallucinate in production takes rigorous mathematical evaluation across Faithfulness, Relevance, and Context Recall.
Empirical Architecture Comparison: RAG Evaluation Triad Metrics and Mathematical Formulations
| Metric Name | Evaluated Dimension | Formula / Mathematical Definition |
|---|---|---|
| Faithfulness | Are generated claims supported by retrieved context? | $\frac{|\text{Verifiable Claims in Answer}|}{|\text{Total Claims Made in Answer}|}$ |
| Answer Relevance | Does the answer directly address the user query? | $\frac{1}{n} \sum_{i=1}^n \cos(\vec{E}(\text{Original Query}), \vec{E}(\text{Generated Question}_i))$ |
| Context Recall | Did the retriever find all information needed to answer? | $\frac{|\text{Ground Truth Sentences Attributed to Context}|}{|\text{Total Sentences in Ground Truth}|}$ |
| Context Precision | Are relevant chunks ranked higher than irrelevant ones? | Mean Average Precision (mAP) over retrieved text chunks |
1. The Failure Modes of Production RAG Systems
Retrieval-Augmented Generation is commonly treated as a solved problem: embed text, store in Pinecone, query top-k chunks, pass to GPT-4. In production, however, two independent failure modes arise:- Retrieval Failure: The vector search returns top-5 chunks that are semantically similar but lack the specific factual entity requested (Low Context Recall).
- Generation Failure: The correct context is present, but the model hallucinates external assumptions or confuses numbers (Low Faithfulness).
2. Automated Synthetic Testset Generation
Manually labeling thousands of evaluation Q&A pairs is financially and operationally prohibitive. RAGAS solves this by generating synthetic test datasets directly from corpus documentation. Using evolution techniques (Reasoning, Multi-Context, and Conditional questioning), the generator synthesizes diverse query profiles:from ragas.testset.generator import TestsetGenerator
from ragas.testset.evolutions import reasoning, multi_context
# Generates 100 enterprise benchmark test cases from documentation
generator = TestsetGenerator.from_langchain(generator_llm, critic_llm, embeddings)
testset = generator.generate_with_langchain_docs(
documents,
test_size=100,
distributions={reasoning: 0.4, multi_context: 0.4, conditional: 0.2}
)