YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
SpecMem: Accelerating Agentic Tool Calling via Live Memory Management
Agents spend much of their latency decoding tool calls token by token, yet the calls a user needs are highly repetitive across sessions. SpecMem turns that repetition into speed: it keeps a small, per-user, capacity-bounded memory of past tool calls that is updated live as the agent runs, and retrieves the closest past call as a draft for speculative decoding. The served model verifies the draft, so outputs are exactly those of standard decoding β wrong drafts cost only compute, never correctness.
The key finding is that liveness drives the gain: a store that keeps ingesting and evicting stays fresh as the query distribution drifts, while a frozen datastore (built once, then fixed) degrades over sessions.
Highlights
- Live per-user memory as a drafter. Top-1 cosine retrieval over
lightweight query embeddings (
all-MiniLM-L6-v2, CPU), per-user partitioning, LRU eviction at a small fixed capacity, online write-back after every verified call. No training, no extra GPU. - End-to-end wall-clock speedups of 1.62x / 1.74x / 1.18x / 1.69x over vanilla autoregressive decoding on API-Bank, ToolAlpaca, BFCL v4, and ToolBench, matching or exceeding a faithful frozen-datastore (ToolSpec-style) baseline on all four.
- Verified across serving stacks and architectures:
gpt-oss-120b(MoE) andgemma-4-31B-it(dense) on sglang,Nemotron-3-Super-120B(hybrid-SSM) on vLLM β anything with an OpenAI-compatible endpoint works. - Safety-aware speculation: an idempotency gate defers speculative execution of irreversible tools (payments, deletes), keeping the speedup while avoiding side effects a verifier cannot undo.
How it works
user query ββ> embed ββ> per-user memory (capacity-bounded, LRU)
β top-1 cosine β₯ Ο
βΌ
drafted tool call ββ> served model verifies
β² (token-level accept)
β
write-back of the verified call (live update)
Every query is answered once by the served model with greedy decoding (the target). Each memory policy ("arm") drafts from its own store and is scored by the token-level longest common prefix between its draft and the target β the accepted-token count a speculative decoder would realize. The compared arms:
| Arm (code name) | Description |
|---|---|
no_memory |
schema-only draft; lower bound |
static_global |
one global store frozen after warmup (ToolSpec-style) |
personal_memory |
SpecMem: per-user, live-updating, capacity-bounded |
toolspec |
faithful ToolSpec reimplementation (frozen kNN-vote + schema FSM) |
Installation
git clone <this-repo> specmem && cd specmem
pip install -r requirements.txt
bash scripts/download_data.sh # fetches BFCL, Seal-Tools, ToolAlpaca, API-Bank
Benchmark data is downloaded from the official sources, never redistributed
here; two datasets need a small manual step (ToolBench, tau2-bench) β see
data/README.md.
Quickstart
1. Serve a tool-calling model behind any OpenAI-compatible endpoint, e.g.
python -m sglang.launch_server --model-path openai/gpt-oss-120b --port 30000
2. Run the main acceptance experiment (3 arms x 40 users x 12 sessions, 3 stream seeds β the paper's headline setting):
python -m harness.run_accept \
--users 40 --tasks-per-user 15 --sessions 12 --queries-per-session 6 \
--capacity 48 --n-seeds 3 --url http://localhost:30000/v1 --tag main
Targets are cached by exact query string, so re-runs and all memory-arm
replays are GPU-free. Results land in results/main_accept_results.json
(per-session and overall MAT / accepted fraction / exact rate).
Add --benchmark sealtools for Seal-Tools. For other served models, point
--url/--model at the endpoint and --model-path (or the
SPECMEM_TOKENIZER env var) at the model's tokenizer so acceptance is
measured in that model's own tokens.
Reproducing the paper
| Experiment | Command |
|---|---|
| Main acceptance table (BFCL / Seal-Tools) | python -m harness.run_accept ... (above) |
| 4-benchmark main table + wall-clock speedups | python -m harness.phase4_maintable |
| Freshness-over-sessions curve | python -m harness.phase4_partb, then python -m harness.phase4_freshness_fig |
| Memory-capacity sweep | python -m harness.capacity_sweep |
| Ablations (eviction, sharing, perturbation) | python -m harness.run_ablation |
| Warmup-fraction sweep | python -m harness.review_r1_warmup |
| Provenance (shared vs per-user) | python -m harness.review_r2_provenance |
| Retrieval-threshold sweep | python -m harness.review_r3_confidence |
| Reset / TTL memory-policy arms | python -m harness.reset_arm, python -m harness.ttl_arm |
| Suffix-decoding baseline | python -m harness.phase4_suffixdecoding_maintable |
| Throughput / overlap under load | python -m harness.phase4_throughput, python -m harness.phase4_overlap |
| Speculative-execution safety gate | python -m harness.safety |
| Bootstrap confidence intervals | python -m harness.bootstrap_ci |
| tau2-bench live traces + scoring | python -m harness.tau2_live generate / extract / score |
The tau2-bench generate mode runs the served model as the agent against a
live GPT-4.1 user simulator and requires OPENAI_API_KEY (and optionally
OPENAI_BASE_URL) in the environment, plus a
tau2-bench install
(TAU2_BIN, TAU2_DATA_DIR). Credentials are read from environment
variables only and a leak check aborts if a key ever appears in an artifact.
Repository layout
harness/ all experiment code (run as python -m harness.<module>)
memory.py memory arms: NoMemory, StaticGlobal, PersonalMemory (SpecMem),
ToolSpecBaseline, suffix-decoding baseline
simulate.py multi-session, multi-user query-stream generator
data.py benchmark loaders (BFCL, Seal-Tools, ToolAlpaca, API-Bank,
ToolBench, tau2)
client.py OpenAI-compatible client + tool-call parsers (harmony, XML)
metrics.py canonicalization + token-LCP acceptance scoring
run_accept.py main 3-arm acceptance experiment
... see the table above for the per-experiment entry points
scripts/ data download
data/ benchmark data (downloaded; see data/README.md)
results/ experiment outputs (created at runtime)
Environment variables
| Variable | Purpose | Default |
|---|---|---|
TOOL_SERVER_URL |
served-model endpoint | http://localhost:30000/v1 |
SPECMEM_TOKENIZER |
tokenizer for the accept metric | openai/gpt-oss-120b |
OPENAI_API_KEY / OPENAI_BASE_URL |
tau2 user-simulator credentials | β |
TAU2_BIN / TAU2_DATA_DIR |
tau2-bench CLI and data locations | tau2 / β |
Citation
The paper is currently under review. If you use this code, please cite:
@article{specmem2026,
title = {SpecMem: Accelerating Agentic Tool Calling via Live Memory Management},
author = {Anonymous},
note = {Under review},
year = {2026}
}
License
This repository is released under the Apache License 2.0.
Benchmark datasets and served models keep their own licenses (see
data/README.md).