RAG evaluation: fix retrieval before you blame the model
Most RAG failures are retrieval failures. A practical playbook for retrieval evaluation: golden sets, recall@k, chunking tests and drift monitoring.
Every article, newest first. Filter by pillar:
Most RAG failures are retrieval failures. A practical playbook for retrieval evaluation: golden sets, recall@k, chunking tests and drift monitoring.
A better embedding model is only an upgrade if you can install it. The sizing, dual-write and cutover plan for re-indexing 10M+ vectors with no downtime.
Most agents have one validator and it passes well-formed nonsense. The four validation layers, what each can prove, and how to test guardrails in CI.
Agent evaluation breaks when a flaky score reads as a regression. Mine a golden set from real traces, measure the noise floor, gate on the right thing.
Most Redis sharding goes wrong before the first shard exists. Diagnose the real ceiling, design keys around hash slots, and reshard without a p99 spike.