MechInterp & AI Safety — Research Landscape

— papers
— topic clusters
Aug 2024 – Aug 2026

Abstract
Links
How to read this. These signals are heuristics computed over the collected corpus (topic-cluster size, year-over-year growth, and embedding-space adjacency between clusters) — not a definitive claim about what's understudied. Small or fast-declining clusters may point to under-explored niches, or simply less-searchable phrasing on my part while collecting papers. Topically-adjacent-but-distinct clusters (high cosine similarity between cluster centroids, appearing in different regions of the graph) often mark promising intersections nobody has fully connected yet. Use "Ask the Corpus" to interrogate these signals further.

Smallest topic clusters

Fewest papers collected — potential niches or emerging sub-areas

Fastest-growing clusters (2024 → 2025)

Where research volume is accelerating

Slowing / shrinking clusters

Relative attention may be moving elsewhere

Adjacent-but-disconnected topics

High embedding similarity between cluster centroids — candidate unexplored intersections
What looks understudied right now?
What's known about sparse autoencoders?
Steering ↔ jailbreak defense overlap?
CoT faithfulness — recent work?

About this assistant

Runs entirely in your browser. It retrieves the most relevant papers by lexical/topical match against the collected corpus, quotes real abstract text and extracted findings, and (with an optional key below) asks an LLM to synthesize a grounded answer.

Without a key, you still get a solid extractive answer built directly from matched papers — no data ever leaves your machine.

Optional: bring your own key

Key is used only for direct browser→API calls. Nothing is stored or sent anywhere else.