Four things the build settled that this document left open:
- The survival threshold is 0.6 Jaccard over stopword-stripped tokens, and
the spec named none. It is nowhere near the decision boundary on real data:
re-measured over 452c23fd's eight rounds, the highest overlap between any
dropped sentence and any surviving one is 0.29, and five rounds peak
below 0.10. Summaries paraphrase; they do not quote. The threshold earns its
place only on the case that does quote, which is a manual
/compactcarrying a verbatim instruction. - A round's span starts after the previous round's summary, not its boundary. The summary was never part of the conversation the next round dropped, and folding it in makes every round after the first quote the previous summary back at the caller.
limitdefaults to 20, not 40. The output is read by a session that has a context budget; the totals still describe the whole session.- The three "I don't know" cases are named strings, not shapes a caller has
to infer from an empty array:
no_compactions_in_this_session,nothing_decision_shaped_was_dropped, andevery_dropped_decision_survived_in_a_summary.
Still open, and unchanged by shipping: the false-positive rate of because,
whether 400 characters is the right ceiling, and whether a reversed decision
should still come back. The sequencing's step 4 -- precision measured by hand
on 50 returned sentences -- has not been done.
The one-line version
AgentWorth holds the transcript compaction threw away. Diff it against the summary that replaced it, and hand back the decisions the session no longer remembers making.
The problem, stated by the person who has it
The session confidently re-proposes something it already tried and rejected three hours ago. I remember the rejection. It does not, because the turn where it happened was summarised away.
Then I retype the reason. Then it happens again.
Why only this product can do it
Compaction is destructive from inside the harness and lossless from outside. The model's view is replaced by a summary; the full JSONL is untouched on disk. So the pre-compaction span and the summary that replaced it both exist, side by side, in a file nothing reads.
compaction.md measured the loss in bytes: about a third of one percent
survives per round, eight rounds deep, 28 KB left of a 68.6 MB session. It also
asserted an asymmetry — summaries keep conclusions and lose reasons — and
marked it as an observation, not a measurement.
This spec measures it.
The measurement
One session, 452c23fd-6e9b-4948-8e8f-6a31f1c3f7dd, 19,955 JSONL lines,
29,642 indexed events, 2.99M tokens, eight compaction rounds. Compaction is
found by isCompactSummary / compactMetadata on the summary line; each round
appears as a marker line followed by the summary itself.
A decision-shaped sentence is assistant text, 25–400 characters, matching one of three patterns. Deterministic, no model:
| Class | Pattern |
|---|---|
| decision | decided, decide to, chose, choosing, opted for |
| rejected | instead of, rejected, ruled out, won't, will not, not going to |
| reason | because |
For each round, count matches in the span about to be dropped, and in the summary that replaces it.
| round | dropped: dec / rej / why | summary: dec / rej / why |
|---|---|---|
| 1 | 6 / 14 / 19 | 0 / 4 / 0 |
| 2 | 0 / 28 / 35 | 0 / 3 / 0 |
| 3 | 6 / 29 / 10 | 0 / 0 / 1 |
| 4 | 8 / 31 / 21 | 0 / 3 / 0 |
| 5 | 6 / 29 / 10 | 1 / 3 / 1 |
| 6 | 1 / 19 / 19 | 1 / 1 / 1 |
| 7 | 15 / 13 / 42 | 1 / 1 / 0 |
| 8 | 10 / 19 / 18 | 5 / 2 / 0 |
| total | 52 / 182 / 174 | 8 / 17 / 3 |
Survival rate: decisions 15.4%, rejected alternatives 9.3%, reasons 1.7%.
Counting distinct sentences rather than class matches, since one sentence can hit two patterns: 402 decision-shaped sentences went into the eight rounds and 28 came out. 374 the session decided something in and no longer has. (Spec measurement, earlier regex — the shipped extractor counted 405 in, 28 out on the same session; see CHANGELOG 0.1.14, current truth.)
The asymmetry is real and it is worse than stated. 174 sentences explaining why something was done went into compaction. Three came out. A session that has compacted keeps roughly one in six of its conclusions and one in fifty-eight of its reasons — which is exactly the shape that makes it re-litigate a settled question, because it kept the answer's shadow and lost the argument.
Round 2 is the clearest single row: 63 decision-shaped sentences in, three out, and not one of them a reason.
Scope
Compaction is rare. compaction.md measured 22 of 543 sessions over 50 KB
compacted at least once, at a median 23.2 MB against 0.3 MB for the rest. This
tool serves the 4% of sessions that are 70× the size of a normal one — which is
also the 4% where a day of work is at stake.
The index copy used here predates #62, so it has no compaction_count column;
the numbers above come from the raw JSONL, not from SQLite. On a current index
compaction_count > 0 selects the population directly.
The MCP tool
forgotten_context(session_id?, round?, classes?, limit?)
| Param | Type | Default |
|---|---|---|
session_id |
string | the caller's most recent session for its cwd |
round |
integer, 1-based | all rounds |
classes |
subset of decision | rejected | reason |
all three |
limit |
integer | 40, ceiling 200 |
Returns:
{"session_id": "452c23fd-…", "rounds": 8,
"dropped_total": 402, "survived_in_summary": 28,
"forgotten": [
{"class": "rejected", "round": 3,
"text": "Going with a marker table instead of a second pass — the second
pass re-reads 68 MB for one boolean.",
"sequence": 8441, "timestamp": "2026-09-01T15:12:03Z",
"model": "claude-opus-5",
"followed_by": ["tool_call:Edit crates/storage/src/lib.rs",
"shell_command:cargo test -p agentworth-storage"]}],
"receipt": {"source_path": "~/.claude/projects/…/452c23fd-….jsonl",
"extracted_at": "2026-09-02T09:41Z",
"method": "regex_v1", "no_model": true}}
followed_by is what makes a sentence checkable. A stated decision with a tool
call after it was acted on; one with nothing after it is a claim. Both are
returned, labelled, and the caller decides.
The header line a session actually reads:
Things you decided and no longer remember — 374 sentences dropped
across 8 compaction rounds, 28 survived in the summaries.
The "I don't know" cases, all three of them:
- Session never compacted →
{"rounds": 0, "forgotten": []}. Not an error, and not an empty answer dressed as a finding. - Session compacted but no pattern matched →
forgotten: []withdropped_total: 0andmethod: "regex_v1"in the receipt, so the caller can tell "nothing was decided" from "the regex found nothing." - Raw JSONL missing or unreadable → refuse.
sessions.source_pathcan point at a file that has since been deleted; returning a partial diff from an index row would be inventing content.
No model, on purpose
A model could extract decisions better than three regexes. It is still the wrong v1.
The output is fed to an agent that cannot verify it. If a model paraphrases the dropped span, the tool becomes a second summariser — the exact lossy step this spec exists to undo — and the receipt stops pointing at words anyone said. A regex returns the sentence verbatim with a sequence number, which is a quotable fact. 374 verbatim sentences with false positives in them beats 40 fluent ones nobody can check.
Revisit when there is a measured precision number for the regex to beat.
New work
All three built in #83.
- Compaction round boundaries as a stored artifact.
compaction_countandcompaction_tokens_droppedexist since #62; the line offsets of each round do not, and re-scanning a 68 MB JSONL per call is not acceptable. One table:session_compaction(session_id, round, start_seq, end_seq, summary_seq, tokens_before, summary_tokens), written by the scanner and backfilled once for sessions indexed before it existed (Storage::needs_backfill, #74). Derivation isagentworth_schema::compaction_rounds, not an adapter's job. - The extractor, in
agentworth-outcomesbeside the loose-ends detector. Same sentence splitter, same length bounds — one implementation, not two. - The MCP tool, plus
agentworth forgottenand a handoff section.
Extraction runs on demand from the raw trace, not at scan time. Storing 402
sentences per compacted session in SQLite would duplicate transcript content
into the index, which AGENTS.md forbids.
Deliberately not built
- No model in v1. Argued above.
- No re-injection. The tool returns sentences. It does not write them into
a context, a
CLAUDE.md, or a prompt. - No cross-session diff. One session's own compaction rounds. "What did the
last session decide" is
carry_forwardinhandoff.md. - No warning before compaction.
compaction.mdis right that the useful version is a warning, and right that it needs a correlation nobody has measured yet.
Sequencing
session_compactionboundaries, written by the scanner.- The extractor plus a fixture test on a small compacted session.
- The MCP tool, regex only.
- Measure precision by hand on 50 returned sentences. Only then consider a model, and only if the number is bad.
Open questions
- What is the false-positive rate of
because? 174 matches in one session is a lot, and "because" appears in narration as readily as in reasoning. - Do the summaries drop reasons, or do the models simply state reasons more often in the dropped spans than the summariser has room for? The measurement above cannot separate those.
- Should a decision that was later reversed still be returned? It is forgotten either way, and returning it might re-suggest a dead end.
- Is 400 characters the right sentence ceiling? A long decision paragraph is the most valuable thing here and the current bound excludes it.