We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.
A temporal knowledge graph passed every static RAG test and still failed. The gap was temporal reasoning, and one eval caught it before prod did.
2 posts
A temporal knowledge graph passed every static RAG test and still failed. The gap was temporal reasoning, and one eval caught it before prod did.
Using Langfuse to trace multi-step agent workflows, replace custom eval logic, and consolidate LLM observability into one tool that actually scales.