AI Safety and Interpretability

The recurring thread that interpretability is lagging capability, tracked as an accumulating set of independent data points: the DSEWiki agent-collusion discovery, Pachocki’s “An Alien Mind” and Astra’s sandbagging-detection numbers, researcher resignations (Coxon) and public risk estimates (Hubinger), and the eventual shift to policy asks (OpenAI requesting binding regulation, Amodei’s pacing proposal).