surgical-forget
Exact, targeted removal of individual facts from a trace-based memory. No retraining. No index rebuild. No per-fact parameters.
surgical-forget is a tiny pure-NumPy library for storing keyโvalue facts in a
holographic memory substrate and removing any one of them with a single
complex-vector subtraction. Because the memory is a sum of bindings, forgetting
a fact is the algebraic inverse of storing it: subtract the binding, and the fact
is gone โ exactly, at the trace level, at O(D) cost.
The library ships as a single file, has no dependencies beyond NumPy, and is
transformers-style (save_pretrained / from_pretrained).
Why
RAG systems forget by deleting from an index, retraining, or rebuilding. None of these is exact, and all of them are slow. A trace-based memory that stores facts as explicit algebraic terms can remove one term without touching the others, and can prove it did.
This is the substrate-level primitive for machine unlearning in retrieval-augmented generation: remove a fact at inference time, verify the removal, keep the rest of the memory intact.
How it works
Facts are stored as complex phasor bindings:
trace = ฮฃ bind(key(label, role), value(v))
Removal:
trace = trace - bind(key(label, role), value(v))
That is the whole mechanism. There is nothing to train, nothing to re-derive, and nothing to approximate.
Install
pip install numpy
# copy surgical_forget.py into your project
Usage
from surgical_forget import SurgicalMemory
mem = SurgicalMemory(D=512)
# Store
fid_france = mem.store("France", "Paris", role="capital_of")
fid_japan = mem.store("Japan", "Tokyo", role="capital_of")
# Query
mem.query("France", ["Paris", "Tokyo", "Berlin"], role="capital_of")
# {'Paris': 0.400, 'Tokyo': 0.023, 'Berlin': -0.014}
# Forget โ exact subtraction, O(D)
mem.forget(fid_france)
# {'ok': True, 'fid': 0, 'pre_sim': 0.400, 'post_sim': 0.052,
# 'verdict': 'FORGOTTEN', 'query_sim': 0.052}
# Verify
mem.verdict("France", "Paris", role="capital_of")
# ('FORGOTTEN', 0.052)
What the numbers show
| Metric | Value |
|---|---|
| Stored facts (8, D=512) | 8 |
| Retrieval accuracy (before forgetting) | 8 / 8 |
| Retrieval accuracy (after forgetting 3) | targets โ chance, non-targets unchanged |
| Non-target interference | 0 / 5 facts affected |
| Trace pre-removal similarity | +0.400 |
| Trace post-removal similarity | +0.052 (cross-talk floor) |
| Storage | 4.5 KB total (trace 4 KB + fact list) |
| Removal cost | one complex-vector subtraction, O(D) |
The 0.052 residual is cross-talk from the remaining facts, not imperfect
removal. Removing all facts leaves a trace norm of 4.45e-06 โ the mechanism
is exact at the trace level.
Contrast with noise
| Noise injected (phasors) | Facts affected |
|---|---|
| 0 | 0 / 8 |
| 8 | 0 / 8 |
| 40 | 0 / 8 |
| 200 | 7 / 8 |
| 1000 | 8 / 8 |
Noise degrades all facts non-specifically. Subtraction removes exactly the target.
API
| Method | Description |
|---|---|
store(label, value, role=None) โ fid |
Add a fact. Returns an integer id. |
query(label, candidates, role=None) โ dict |
Similarity to each candidate. |
verdict(label, value, role=None) โ (str, float) |
PRESENT / FORGOTTEN at 2/โD. |
forget(fid, verify=True) โ dict |
Exact removal. Returns before/after metrics. |
history() โ list |
Full audit ledger. |
live() โ list |
Currently-stored facts. |
bytes_used() โ int |
Trace plus fact-list bytes. |
save_pretrained(path) |
Write trace + config. |
from_pretrained(path) |
Load a memory. |
What it cannot do
- Remove knowledge baked into transformer weights. That influence is distributed and nonlinear. Surgical removal requires a trace-structured memory.
- Prevent reconstruction from residuals. An adversary with the trace and codebook can still attempt inference. Removal removes the direct contribution, not the statistics.
- Handle concurrent writes. Single-threaded.
- Guarantee removal under ongoing writes. If new facts are stored while a removal is in flight, the trace's pre-removal state must be tracked.
Intended use
- Machine unlearning in retrieval-augmented generation.
- GDPR / right-to-be-forgotten workloads where a specific fact must be provably removed.
- Continual memory where a task or topic must be dropped without disturbing unrelated knowledge.
- Auditable memory: the ledger records every store and forget.
Out-of-scope use
- Claims of cryptographic security. The trace is not encrypted.
- Adversarial settings where the attacker has unbounded codebook access.
- General model unlearning. This operates on a trace, not on weights.
Citation
@software{q2026surgical,
author = {zeechimp},
title = {surgical-forget: Exact Trace-Level Fact Removal},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/zeechimp/surgical-forget}
}
License
MIT
Evaluation results
- Retrieval accuracy (8 facts, D=512) on Synthetic capital-city factstest set self-reported1.000
- Non-target interference (0 / 5 facts affected) on Synthetic capital-city factstest set self-reported1.000