Skip to content
Sadiq Khan

may 19, 2026 · 2 min read · rag, knowledge-graphs, retrieval

The oldest algorithms in your RAG stack

Louvain, PageRank and a 50-token trick: how much classical computer science is quietly powering 2026 retrieval.

The prettiest part of your RAG stack is also the most honest part.

I spent some time this week deep in knowledge-graph land for an AI project, and came out with a renewed appreciation for how much classical computer science is quietly powering the 2026 AI boom.

Graphs unlock "global" reasoning that embeddings can't. Recent research on graph-augmented RAG shows entity graphs plus community detection winning 70–80% of multi-document sense-making queries against vanilla similarity search. Once your corpus crosses roughly 50 documents, graphs start paying for themselves.

A 17-year-old algorithm is doing the heaviest lifting. Louvain modularity (Blondel et al., 2008) powers community detection in most modern graph-RAG approaches. The fanciest retrieval pipelines in production are running on a paper older than the iPhone 4.

HippoRAG is the most underrated 2024 paper. Gutiérrez et al. (NeurIPS 2024) ran personalised PageRank over a knowledge graph and beat vanilla RAG on multi-hop QA by about 20%, with cheaper retrieval. PageRank, from 1998. Still undefeated.

Depth is a precision dial, not a UI knob. Most multi-hop QA benchmarks (MuSiQue, HotpotQA, 2WikiMultiHop) are dominated by 2-hop reasoning. One hop gives you entity facts. Two hops answer "how does X relate to Y", the sweet spot for almost every real question users ask. By three hops, precision usually collapses. Defaults matter.

Contextual chunking is the cheapest 49% win in RAG. Published work this past year showed that prepending about 50 tokens of document-level context to each chunk before embedding lifts retrieval quality by around 49%, and around 67% with a reranker. Almost nothing else in this field has that kind of return.

The aesthetic is a load-bearing feature. Soft glow, neighbour fade, labels on focus: that's Bret Victor's ladder of abstraction in motion. Declutter by default, reveal on intent. The visualisations users can read are the ones they trust.

The graph view isn't decoration. It's the receipt the retrieval layer hands to a human, proof the answer is grounded in something real.

Build the substrate, make it legible, and the trust follows.


Originally posted on LinkedIn.

← All writing

Contact

sadiqkhan795@gmail.com

Say hello. I read everything.