The question
Change exactly one belief (“The Eiffel Tower is in Paris” becomes “…in Rome”) so that the model gives the new answer, generalizes it to rephrasings, leaves unrelated facts alone, and uses the new fact when reasoning. As of 2026, no method does this reliably at scale.
What I found so far
- A single naive rank-one edit worked, but raised held-out perplexity by 148%. Adding C⁻¹ preconditioning brought that drift to roughly 0%.
- The remaining failures had a cause: nearby facts (the Louvre, the Colosseum) flipped to “Rome” too, because their internal representations share directions with the edited fact. That’s superposition, and there’s no perfectly clean “one fact” to grab at the weight level.
- Next: 30 sequential CounterFact edits across methods, and whether SAE/transcoder features on Gemma-2-2B (with Gemma Scope) predict where a fact lives better than causal tracing does.