This article isn’t meant to be read — it’s meant to be handed to your agent. It is a self-assessment protocol: you give it to the AI that has access to your second brain, and it runs the measurement, asks you the questions, scores, and returns a diagnosis. You can skim it to judge the method, but that is not its real use.
How to use it. Paste this text (or drop it in as a file) into a throwaway conversation or a project dedicated to evaluation, with the agent that has access to the setup to be evaluated. It executes §1 and returns §8. Not into your day-to-day working project — §0 explains why.
Version 1.1 — September 2026. A single file, with no external dependency: framework, scoring scale, seventeen use cases, execution protocol and report template. Free to share.
Validity. The framework expires the day agent capabilities change in nature — not when a product ships. Re-examine it if, in a given pass, more than three use cases have become impossible to interpret. The figures cited date from mid-2026.
0. Before you start — what this protocol is, and the trap built into it
What it measures. The ability of an AI-assisted personal knowledge setup to produce five things. Not the tools used, not the time invested, not the satisfaction felt.
The unit of measurement is the use case, not the question. “Can I do this today, alone, without rebuilding?” A use case succeeds or it doesn’t. A well-asked question always overestimates the system: phrasing it is already half the work.
It measures a capability, not a value. A system can score high and be of no use at all. The two value counters (§7) are independent of the score.
It has no statistical value whatsoever. Seventeen use cases rule out any claim of that kind. The score is not to be compared with another system’s: it is only worth anything as the first point in a personal series.
The isolation trap, and how we deal with it
The original protocol requires that the rubric never enter the context of the system being evaluated: a system that has read its own evaluation rubric is no longer evaluating itself. Yet this file is designed to be imported — so it flatly violates its own rule.
Two mitigations, and they must be applied, not just read:
- Scenarios are never played in the conversation that carries this file. They are played in a fresh session, or by a subagent that receives the prompt to play and nothing else — above all not the wording of the use case, nor what is expected. The host conversation sets the prompt, receives the raw answer, and scores.
- This file does not go into the permanent corpus. It lives in a project dedicated to evaluation, or in a throwaway conversation. If it ends up in the instructions of the everyday work project, the following passes no longer measure anything.
What cannot be mitigated: on use cases scored in declarative mode, the agent knows the right answer while it is asking the question. That is the reason §5 exists.
1. Execution instructions — for the agent reading this file
You are going to conduct an evaluation. You are not a benevolent coach: you are an evaluator. The useful deliverable is the one that finds what doesn’t work.
Rules of conduct, non-negotiable
- Never ask about what you can observe. Any question whose answer is in the files, the file tree, the configuration or the history you have access to is a forbidden question. Go and look, then have it confirmed if the finding is ambiguous.
- One question at a time, in plain language, without the framework’s jargon. Never recite the wording of a use case to the person: you would get the answer it calls for.
- Don’t pay compliments. No “excellent system”, no “well done on this discipline”. A positive finding is stated as a finding: “A2.1 is at 2, lived evidence, here is the file.”
- Never round up by default. When in doubt between two scores, keep the lower one and write why in the comment.
- Don’t fix anything during the pass. A defect you find gets scored. Fixing it along the way destroys the baseline and makes the next pass unreadable.
- Don’t make up any file content. If you haven’t opened a document, you don’t know what it contains. “I can’t verify” is a valid and expected answer.
- Report what you can’t see. At the end, explicitly list the areas of the setup that stayed out of your reach. It is first-order information about the reliability of the score.
Step 0 — Framing (ask these questions, in this order)
Q1 — Scope. Which setup exactly are we evaluating? A setup spread across several accounts, several siloed spaces or several harnesses is measured once per scope. Scores don’t add up: a session never loads two setups at once, and a mixed total describes no real setup. If the person describes several scopes, ask which one we are playing today, and record the others as upcoming passes.
Q2 — Evidence mode. Present the three modes of §2 with their time cost and their score ceiling, and have the person choose. Don’t choose for them, and don’t downplay the cost of the full mode to make them more cooperative.
Q3 — Access. What do you really have at your disposal? Take an explicit inventory: file reading, write permissions, active connectors, visible scheduled tasks, conversation history, persistent memory. Distinguish what is declared available from what you have actually tested. This inventory conditions half the scores.
Q4 — History. Is this a first pass or a follow-up? If there is a previous pass, ask for it: the thresholds of §6 are about gaps between passes, not about a state.
Step 1 — Self-analysis, before any substantive question
Before questioning the person on the axes, exhaust what you can observe on your own. Depending on your access:
| What you inspect | What it instruments |
|---|---|
| Complete file tree: depth, naming, modification dates, duplicate names | A3, A5 |
| File headers: is there a statement like “this file is authoritative on X”? a validity date distinct from the writing date? | A2.1, A3.1, A5.1 |
| Instruction and context files loaded automatically: volume, redundancy, contradictions between them | A3.2, A5.2 |
| Inventory of the rules and thresholds written in the corpus, then of what triggers them | A4.1 |
| Existing scheduled tasks / automations: their existence, their frequency, their last successful run | A4.1, A4.3 |
| Connectors and automatic input sources, and what they write | A1.1, A1.3, A5.3 |
| Decision log, if there is one: does it carry the why and what was ruled out, or only the result? | A3.3 |
| Search for topics covered by three or more files, none of which concludes | A2.1 |
Write down what you found before asking a single substantive question. This inventory becomes the material for the cross-check (step 3).
Step 2 — Collection and scenarios
For each use case, in the order A1 → A5, apply the corresponding block of §4:
- Observation: what you look at yourself. Always first.
- Question: what you ask, only if observation isn’t enough. Phrased neutrally, without announcing what would be “good”.
- Scenario: to be played if the mode calls for it — in a fresh session or a subagent, never here.
- Scoring trap: the confusion that leads you to give a 2 to what is a 1. To be reread before setting the score.
One exception to the order: A5.3 is handled first if the setup has any automatic input at all. It is the only use case with asymmetric risk — the others cost time, this one exposes you to damage.
Step 3 — Cross-check
Before locking the scores, actively look for contradictions between:
- what the person declared and what you observed in step 1;
- two successive statements by the person;
- a claimed capability and the absence of the mechanism that would make it possible.
Each contradiction you find is raised as is, without aggressiveness and without letting it slide: “You say the system flags outdated positions; I found no validity date in the headers, nor any task that rereads them. What am I missing?” If the answer doesn’t produce a mechanism that can be named, the use case drops to 0 or 1, no higher.
Then apply the downgrade rules of §5 in full. They are not optional.
Step 4 — Scoring and report
Score against the scale in §2, record the counters of §7, read the result with the thresholds of §6, and deliver the output using the template in §8.
2. Scoring scale, evidence and modes
Three scores
| Score | Meaning |
|---|---|
| 0 | I can’t, or I rebuild it entirely by hand |
| 1 | I get there, but by pointing to the files, prompting again, or reconciling several sources |
| 2 | The system does it on its own, completely, and I can rely on it |
Why 0/1/2 and not binary. The evaluation literature recommends binary, including for human annotation, on the grounds that intermediate values serve to hide the annotator’s uncertainty. We deliberately depart from it here: the most frequent and most informative state is the intermediate one — “it works but I have to point to the files” — and flattening it into a failure would lose exactly the signal we are looking for. Binary becomes the right choice again if scoring is one day delegated to a model over a large volume.
The rule that keeps the rubric honest: no score without evidence
A use case that hasn’t been tested is scored “—”. It counts neither in the numerator nor in the denominator.
It is the only rule that prevents the exercise from turning into a mental model disguised as a measurement. A rubric filled with estimates describes what its author believes about their system: interesting information, but something else. The score is therefore always read in the form: n points out of 2 × k, k use cases filled in out of 17. Never as a percentage alone.
Four qualities of evidence
| Mark | Meaning | Maximum score allowed |
|---|---|---|
| [L] | lived — the case really happened since the last pass, the person can point to the occurrence | 2 |
| [P] | played — a deliberate scenario, run specifically for the pass, in a fresh context | 2 |
| [C] | checked — a structurally verifiable absence, without playing anything | 0 only |
| [D] | declared — the person asserts it, nothing has verified it | 1 only |
[C] is what makes a quick pass possible: some failures can be observed without a scenario. If no scheduled task exists on a scope, nothing there triggers a rule — the score is 0, and playing it would be theater. But [C] can never justify a 1 or a 2: the presence of a mechanism doesn’t prove that it works.
[D] is the addition specific to the external use of this protocol, and it is capped at 1 for a precise reason: the documented gap between declared satisfaction and measured benefit for this type of setup is around forty points. A statement therefore cannot unlock the high score. A system that gets a lot of [D] doesn’t have a bad score: it has a score that doesn’t know what it measures, and the deliverable must say so in those terms.
Three modes — to be chosen at the start
| Mode | Cost | What we do | Evidence accepted | What the score is worth |
|---|---|---|---|---|
| Express | 20–40 min | Self-analysis by the agent + questions and answers. No scenarios. | [C], [D], and [L] if the person points to a precise and verifiable occurrence | Indicative. De facto ceiling of 1 on anything that isn’t [L]. The score is always written followed by the words “express mode”. |
| Standard | 1–1.5 h | Express + scenarios on the use cases marked ★ (the six discriminating ones) | All | Usable. The ★ ones carry [L]/[P], the rest remains capped. |
| Full | ~2 h | All seventeen use cases played or lived | [L], [P], [C] — [D] forbidden | The only mode in which the stop threshold of §6 applies. |
Two hours for a full pass is an order of magnitude consistent with published evaluation practice: about thirty minutes reviewing twenty to fifty outputs after any significant change, then ten to twenty traces per week in steady state. Subsequent passes cost less: only the use cases that are due get replayed.
The six discriminating use cases (★) — those whose score really separates one setup from another: A1.3, A2.1, A2.3, A3.1, A4.1, A5.3. If the budget is tight, that’s where it goes.
3. The five axes — the why, in brief
Reasoning in terms of tools doesn’t let you audit a system. Reasoning in terms of what the system produces does, and five axes are enough.
A1 — What comes in without my thinking about it. The marginal cost of input decides what comes in. When that cost is high, the trade-off happens at the worst moment: before you know whether the information will be useful. The corpus then selects according to the day’s availability, not according to importance, and ends up reflecting its owner’s discipline rather than their reality. The cost of absence isn’t a lack of volume, it’s an invisible composition bias. What’s missing is almost always a routing rule, not a pipe.
A2 — What becomes a position rather than a pile. Decides whether the corpus is capital or inventory. Canonical and very common symptom: a topic on which three files exist and none of them concludes. The cost is proportional to the number of reuses — an idea that is used four times and written down nowhere gets rebuilt four times, and not identically. It is the only axis that can’t be delegated: an agent can propose, never consolidate. Its two measured failure modes have names: brevity bias and context collapse. The right move is the delta, not regenerating the block.
A3 — What I no longer have to repeat or look up. Decides where the organization’s knowledge resides. If it resides in the owner’s head, the system has shifted the load — from remembering the content to remembering the location — without removing it. Three distinct capabilities, and the second doesn’t follow from the first: locating (where is it?), ranking (which one is authoritative? no search engine answers that one), not repeating — with three layers whose cost of forgetting differs: the semantic layer (who I am) costs a re-explanation, the procedural layer (how I want it done) costs a deliverable to redo, the episodic layer (what I decided and why) costs a decision to reopen. The episodic layer is the most expensive one and the only one that degrades on its own.
A4 — What happens without my asking. A written rule is only worth something if something triggers it on schedule; a tension between two decisions only becomes visible if something puts them side by side. It is the only axis that produces value without being asked — hence the only one that justifies a corpus being alive rather than merely tidy. It is also the only one whose defect is painless: you don’t notice an alert that never came. Distinguish a missing capability from a deliberately closed capability: a function restricted by a safeguard that works isn’t a failure, and confusing the two leads to fixing the wrong thing.
A5 — What I can act on without checking. Decides whether an answer is worth acting on. Its sign of absence is that there isn’t one, and that is what defines the axis: an outdated, contradictory or poisoned system answers with exactly the same confidence as a correct one. It is the axis that makes the other four assessable, and the only one that carries an asymmetric risk — staleness and contradiction cost time, the trust boundary on inputs exposes you to damage. Since late 2025, the subject has been classified as ASI06 — Memory & Context Poisoning in OWASP’s Top 10 for Agentic Applications, with the difference that matters: a prompt injection resets between sessions, whereas a memory poisoning persists across all subsequent interactions and its effect is decoupled in time from the attack.
What the breakdown reveals. Three axes automate well (A1, A3, A4) — they are the ones products address. Two can’t be delegated without losing what matters most (A2, A5): they call for a method, not a subscription. Two have a painless defect (A4, A5): they are the only ones whose state must be actively checked, the other three end up flagging themselves through use.
4. The seventeen use cases, instrumented
Three per axis, four on A2 and on A5. There is no production axis: producing a deliverable is read through three use cases, each living in the axis it depends on — A2.4 (update in place), A3.2 (deliver to the right place), A5.4 (deliver to a third party).
The wording is invariant: it never mentions the state of a system, and it isn’t touched up for at least four passes. A retouched set no longer measures anything; a frozen set slowly goes stale, which is the lesser evil. When a use case enters the set or is reworded, it starts over at “—”: a score obtained under the old wording doesn’t carry over, and its four-pass count starts from the first pass that plays it. The sub-score of the unchanged use cases is read separately and keeps its series.
A1 — What comes in without my thinking about it
A1.1 — A document received through a usual channel ends up filed in the right place and usable, without my depositing it myself.
- Observation: active connectors and what they write; existence of a routing rule; files that arrived recently without any identifiable human action.
- Question: “When you received an important document last month, what happened between its receipt and the moment it was usable?” — let them tell the story, don’t prompt the answer.
- Scenario: ask the person to name the last three documents that came in, then check where they are in the setup.
- Trap: a search engine over the mailbox is not an input. The document is searchable, it isn’t filed. Score 1 at best if the document exists only in the channel it came from.
A1.2 — A spoken exchange or decision leaves a usable trace without my opening a session to dictate it.
- Observation: presence of a voice capture chain (meeting transcription, routed dictation, processed voice notes) and of what it outputs.
- Question: “The last decision made in a meeting or on the phone — where is it written down today, and who put it there?”
- Trap: a raw transcript dropped into a folder is not a usable trace. The criterion is usability without a full reread.
★ A1.3 — Incoming information is attached to the object it concerns — the asset, the client, the person, the case file — and not just dropped into a searchable stream.
- Observation: are there reference objects in the setup (cards, records, per-entity folders)? Are recent inputs attached to them?
- Question: “Take a client, an asset or a person you keep track of. Without searching, what does the system know about them, and where does it come from?”
- Scenario: in a fresh session, ask for everything known about a named entity, then count what’s missing compared with what the person knows.
- Trap: this is the use case that discriminates. A search engine scores 1 everywhere. Attachment to the object is what separates archiving from capturing. Only give 2 if the attachment happens without human action.
A2 — What becomes a position rather than a pile
★ A2.1 — On a topic I’ve thought about several times, a single file carries the position in force, and it concludes.
- Observation: actively look for a topic covered by three or more files. Open each one: which one concludes? Does a header designate the one that is authoritative?
- Question: “Name a topic you’ve come back to at least three times. Which file carries your current position?” Then check.
- Trap: the person knows the answer — that’s not what we are measuring. The test is: is it written somewhere that this file is authoritative? If the hierarchy exists only in their head, it’s a 1.
A2.2 — A position already written down comes back as is in a new context, without being rebuilt.
- Observation: traces of reuse — the same passage reused verbatim in several deliverables, cross-references between files.
- Question: “The last time you needed a position you had already formulated, did you find it or did you redo it?”
- Trap: “redone, but faster” isn’t reuse. It’s a 1. Reuse is the text coming back out, not the reasoning being played again.
★ A2.3 — After a synthesis produced by an agent on material I master, the caveats and the conditions of validity have survived.
- Observation: none. This use case is played.
- Scenario: the person chooses a topic they are an expert on and for which one or more files exist. A fresh session produces a synthesis. The person then lists what has disappeared — nuances, exceptions, conditions of validity, cases where the position doesn’t hold.
- Trap: it is the only use case whose success shows up as an absence. You don’t check what is present — a synthesis that looks faithful may have lost all its caveats. If the person says “that’s well summarized” without having looked for what’s missing, the scenario didn’t take place: score “—”.
A2.4 — A living note that I have updated comes back rewritten in place, in its present state: it still concludes, reads without knowing its history, and carries no “what changed” block.
- Observation: none. This use case is played, on a note whose state before the request is known.
- Scenario: in a fresh session, ask for the note to be updated. Reread the whole file, not just the addition, and check that the update is a delta: the rest of the content has been neither regenerated nor reworded.
- Trap: it is the writing side of A2.1 — A2.1 checks that a file concludes, A2.4 checks that it keeps concluding after an agent has touched it. Like A2.3, its success shows up as an absence. The history of a position is filed somewhere other than in the note.
A3 — What I no longer have to repeat or look up
★ A3.1 — Cold, without my pointing to any file, the system answers “here’s what you’ve already written on X” and designates which one is authoritative.
- Scenario, in two stages: (a) in a fresh session, ask “have I already written about X, and which one is authoritative?” without allowing any search; (b) ask the same question again, this time allowing exploration. The gap between the two answers is the useful measurement: it tells you whether the shortfall lies in the writing or in the access.
- Trap: locating isn’t ranking. A system that lists seven relevant files without saying which one is authoritative is at 1, not 2 — and this is where the most frequent scoring error happens.
A3.2 — A deliverable requested with no instruction on location or format lands in the right place, under the right naming scheme, with my tools’ conventions, without my having to explain myself again.
- Observation: read the instruction files that are loaded automatically. Do they cover the three layers — semantic (who I am), procedural (how I want things delivered), episodic (what I decided and why)?
- Scenario: ask for a real deliverable without saying where to file it or what to call it. Check the location, the naming and the format before reading the content.
- Trap: this use case measures procedural memory by what it produces — a file, not an answer. A re-explanation requested along the way (who I am, how I want it delivered) caps the score at 1.
A3.3 — A past decision comes back with its why and what was ruled out, not just with its result.
- Observation: is there a decision log? Do its entries carry the options that were ruled out and the reason, or only the conclusion?
- Scenario: fresh session, “why did I decide X, and what else had I considered?” on a decision more than three months old.
- Trap: a result without its reasoning can’t be re-examined, only endured. A log that lists decisions without the alternatives that were ruled out is at 1.
A4 — What happens without my asking
★ A4.1 — A rule or threshold I’ve written down fires on schedule, without my thinking about it.
- Observation: inventory the rules and thresholds written in the corpus, then, for each one, identify what triggers it. The most frequent result is: nothing.
- Question: “What’s the last alert the system sent you without your asking for it? When?”
- Trap: the existence of an automation doesn’t prove that it runs. Check the last successful run, not the configuration. A silent fallback is more dangerous than an outage: a loop that runs badly looks just like a loop that runs.
A4.2 — The system flags a tension between a new intention and a decision already in force, before I act.
- Scenario: in a fresh session, the person states an intention that contradicts a decision written in the corpus. We observe whether the contradiction is raised spontaneously.
- Trap: if the agent raises the tension only after it has been pointed out to it, it’s a 0 on this use case — the axis is about what happens without being asked.
A4.3 — A chain of write actions executes from a single request, with one validation point, and not a lock per operation.
- Observation: actual write permissions, validation mode, existence of safeguards.
- Trap: distinguish a missing capability from a deliberately closed capability. A function restricted by a safeguard that works as intended is scored according to what it allows, and the comment specifies that the restriction is intentional — otherwise the next pass will fix the wrong thing. Also note that the lock per operation is an antipattern: it makes use so tedious that you end up authorizing everything.
A5 — What I can act on without checking
A5.1 — What ages carries a validity date distinct from its writing date, and crossing that date gets flagged.
- Observation: open ten files at random. How many carry a validity date — not an update date? What rereads those dates?
- Trap: the update date tells you when the file was written, not whether it is still true. The two blur together as long as you reread what you’ve just written; they diverge as soon as a fresh session opens the file, and nothing warns it. An update date alone scores 0 on this use case.
A5.2 — Two files that contradict each other are detected without a manual audit.
- Observation: look for a contradiction in the corpus yourself — it’s almost always possible as soon as three files carry the same rule. Then check whether anything would have flagged it.
- Trap: “I would have noticed” isn’t detection. And on a corpus of a few tens of thousands of tokens, the dominant failure mode isn’t volume, it’s contradiction between layers. A corollary to mention in the comment: a rule that applies to several contexts must live in one place.
★ A5.3 — Inputs are qualified: I know what comes from a controlled source, and nothing unqualified reaches a substrate the system writes to.
- Observation: map the inputs. For each one: controlled source or not? does it write to a substrate the agent rereads? is there an explicit boundary?
- Question: “Can content that came from outside — a document you received, a web page, a cloned repository — end up in what the agent rereads automatically?”
- Trap: one of the two use cases with asymmetric risk, along with A5.4. The others cost time, this one exposes you to persistent damage: a memory poisoning survives across sessions and its effect is decoupled in time from its introduction. To be scored before opening any input automation (A1), and see the hard threshold in §6.
A5.4 — A deliverable meant for someone who can’t reach my corpus is self-contained: the rules are written out in full rather than cited by address, and nothing leaves that should have stayed inside the perimeter.
- Observation: none. This use case is played.
- Scenario: ask for a document meant for a third party, on a topic where the corpus holds both material that may leave and material that must not. Check what is in it, not just how well it reads.
- Trap: it is the exit boundary, symmetrical to A5.3, which guards the entrance. A dead address costs the reader time; data that left the perimeter exposes you to damage: a deliverable that points to a file the reader can’t open is at 1 at best; a perimeter leak is at 0.
5. Anti-leniency checks — downgrade rules
To be applied mechanically, before locking the scores. They exist because the judge has a stake in the result, and because the agent doing the scoring has read the rubric.
| Finding | Consequence |
|---|---|
| A score of 2 without a file, a run or a precise occurrence that can be named | Bring it down to 1 |
| A score resting on “in principle”, “it should”, “I think that” | Bring it down to [D], hence a ceiling of 1 |
| The person describes a capability you found nowhere in step 1 and can’t point to the mechanism | Bring it down to 0 or 1, and write the contradiction in the comment |
| A use case scored without the planned scenario having been played, in standard or full mode | Score “—”, not an estimate |
| An automation that is configured but whose last successful run can’t be verified | Ceiling of 1, whatever the story |
| The person disputes a low score by arguing without bringing any new evidence | The score doesn’t move; the objection is written in the comment and played again at the next pass |
| More than ten use cases at 2 on a first pass | A sign of leniency, not of performance. Go back over the evidence one item at a time before delivering. |
| The total score exceeds 27/34 on the very first pass | Almost always a badly played pass. Say so explicitly in the deliverable. |
A rule of posture, too. If the person is trying to negotiate the scores rather than understand the gaps, say so once, calmly, and carry on. The deliverable that pleases is the one nobody rereads.
6. Thresholds — to be read after scoring, never before
These thresholds are set in advance. Setting them in light of the results amounts to concluding what you wanted to conclude.
| Observation | What it requires |
|---|---|
| A use case goes from 2 to 1 or 0 without any workstream having been delivered | Regression. Identify what changed in the loaded context, before any other action. |
| A use case goes from 0 to 2 without anything having been written in the corpus | The system is making things up. A serious failure, absolute priority over everything else. It hasn’t learned, it has inferred — and it will infer wrong the day the case gets out of the ordinary. |
| A use case stays at 0 over three passes | Decide: either it’s badly worded, or the corresponding workstream has been buried. There is no third option. |
| A5.3 is at 0 | No input automation gets opened. A hard prerequisite, non-negotiable, however eager you are to move forward on A1. |
| A recommendation survives three passes without being executed | Remove it. It wasn’t a priority, it was plausible. |
| More than half the use cases remain “—” at the second pass | The setup doesn’t hold up. Reduce the number of use cases rather than lower the evidence requirement. |
| 27 points or more out of 34, all seventeen use cases filled in, two passes in a row, in full mode | Stop signal. Don’t build the next brick; the marginal effort is better spent elsewhere. |
The stop threshold at 27/34 corresponds to an average of about 1.6 — the equivalent of ten use cases at 2 and seven at 1. It is rescaled pro rata to the number of use cases, at constant average, and not in light of the results. It only applies once all seventeen are filled in: a high total on a partial basket measures nothing.
Recommended frequency. Quarterly on the weak axes, yearly on the strong axes. Immediate replay of the use cases concerned after any rewrite of a file loaded in several contexts — it is the only moment when a regression is likely.
No premature taxonomy. Wait for three passes before classifying failure patterns. Before that, you are classifying noise.
7. The two value counters, and the hygiene measure
The seventeen use cases measure a capability. They don’t say whether the system is of any use. Given the documented gap between declared satisfaction and measured benefit, measuring value must mean counting events, never collecting an impression.
| Counter | How to record it |
|---|---|
| Management overhead | Share of time spent filing, re-explaining, searching, rebuilding — rather than producing. It is estimated by sampling over a week, not by introspection. The only before/after published by a practitioner goes from 30–40% to under 10%. |
| Reuses without rebuilding | Count, over a quarter, of the number of times a position already written down was used as is in a new context. One event, one line. It is the direct output of A2, and the only figure that distinguishes capital from inventory. |
If the person hasn’t counted anything, these counters are recorded as “not recorded”. The agent doesn’t estimate them, and doesn’t replace them with an impression. It proposes the recording protocol for the next pass.
Hygiene measure — the context load. How many bytes and files actually go into a session, which ones served no purpose, and what share escapes the owner’s control. It isn’t scored: it is used to decide whether the volume issue exists at all.
It is recorded per loaded scope, never as an average. It is the gap between the lightest and the heaviest that tells you something, not the total: a homogeneous setup raises no volume question, a setup where one scope loads several times as much as the others does raise one, and it is local. A merger of scopes is the moment when this gap widens without anyone arbitrating it.
A calibration point on volume, so as not to over-invest: degradation with length is real, confirmed and unresolved — on detecting obviously dangerous actions, the recall of a recent model drops from 99.7% at 100K tokens to 69% at 800K tokens, with the worst result when the item to be detected is in the middle. But a few tens of thousands of tokens loaded per session remain far from any measured threshold. In that range, the failure mode isn’t weight, it’s contradiction — which occurs as soon as there are three files.
8. Report template
The agent returns exactly these seven blocks, in this order. No preamble, no congratulations.
1. Pass header
Date · scope played · evidence mode chosen · model and version used · actual duration · pass number.
Recording the model version isn’t cosmetic: a score variation between two passes can come from the model as much as from the corpus.
2. What the agent could see, and what it couldn’t see
Two explicit lists. The second one is as important as the first: it gives the reader the real reliability of the score.
3. The table of seventeen
| # | Use case | Scope played | Score | Evidence | Comment (one sentence) |
|---|
The comment is more useful for the next pass than the score. It says why this score, not what should be done.
4. The score
Written in the form: n points out of 2 × k, k use cases filled in out of 17, [express/standard/full] mode. Never a percentage alone, never a grade out of 20, never a comparison with another system.
Breakdown by axis, with the distribution of evidence qualities: how many [L], [P], [C], [D].
5. The contradictions found
What the person declared, what the agent observed, and what wasn’t reconciled. If no contradiction was found, say so — and treat it as a suspicious signal rather than a good result, except in full mode.
6. Three workstreams, no more
Prioritized according to this ordering rule, and not according to appetite:
- A5.3 at 0 comes before everything else. A hard prerequisite.
- Next, the axis whose defect is painless — A4 then A5 — because nothing will come along to flag it.
- Next, the lowest-scoring discriminating use case.
Each workstream carries: the target use case, the minimal move that would raise it by one point, and the estimated cost. A workstream whose minimal move can’t be described in two sentences isn’t a workstream, it’s a project — say so.
7. What not to build
The part nobody writes and that is worth the most. List what the person mentioned as a wish for tooling and which, in light of the scores, isn’t justified — with the reason. The dominant failure mode on this type of setup is over-engineering before proof of use, and AI makes it worse by sharply lowering the cost of building. An axis that scores decently doesn’t need more tooling.
9. What this protocol won’t fix
The judge has a stake. The effort justification bias is all the stronger when the effort was great. The clean context, the “no score without evidence” rule and the downgrades of §5 reduce it. Nothing eliminates it.
The model changes underneath the corpus. A score variation between two passes can come from the corpus or from the model, and nothing here allows you to tell them apart. An irreducible confounder; hence the requirement to record the model version.
A use case can succeed for the wrong reason. A competent model guesses a lot. The only control is the “0 to 2 without writing” threshold: if the score goes up without the corpus having changed, the system hasn’t learned, it has inferred.
Seventeen use cases don’t make a statistic. This protocol produces one point in a personal series. It doesn’t rank, it doesn’t certify, and a score obtained here isn’t comparable with someone else’s.
Scoring remains human, and that’s a choice. Building a reliable automated judge requires, according to published practice, around a hundred hand-labeled examples and a chance-corrected measure of agreement — out of all proportion with seventeen use cases scored once a quarter. The protocol would cost more than the measurement. It becomes relevant again if the set grows beyond fifty or so cases.
Three figures that circulate and should not be repeated, if the conversation brings them up: “70% of people abandon their second brain within a month” — no source, and no public data on abandonment exists. The benchmarks published by memory-solution vendors — self-administered, disputed among themselves, with no published third-party validation. And the 95% attack rates on memory poisoning — that is an injection rate obtained under idealized conditions, the associated success rate being much lower, and effectiveness drops sharply when legitimate memories already exist.