Plexara Research · The knowledge-use study
When agents use what they are told: derivability, model tier, and where a knowledge layer concentrates its value
Plexara, a product of Deasil Works · 2026-07-26 · Models: claude-sonnet-5 and claude-haiku-4.5 · Data and code: txn2/mcp-data-platform · DOI: 10.5281/zenodo.21614059
Abstract
Our accuracy study measured how much a knowledge layer improves an agent’s answers. This study asks the question underneath it: when an agent is handed knowledge by the platform it runs on, does it act on it? Across controlled cells at eight repeats each, run on a pinned platform release through two independent drivers, the answer has a two-factor structure. A strong model re-derives any claim it can check against live data, at any price we set (verification held at ceiling as the cost of checking was swept from one to eleven calls), so delivered observations of checkable state serve it as corroboration, never as a crutch. The same model relies completely on knowledge it cannot re-derive: a delivered internal reporting convention was used in every attempt, and without it the model confidently fabricated a plausible substitute definition in 75% of control attempts, against zero with it. A weaker model inverts the first regime, trusting answer-bearing notes instead of re-checking them, which makes delivered conventions safe on every tier tested and delivered world-state observations a strong-model workload. Read with the accuracy study, whose +56-point lift came precisely from non-derivable facts, the conclusion is consistent: the value of a knowledge layer concentrates in what agents cannot re-derive, which is exactly what Plexara captures, curates, and delivers.
This page presents the results with commercial framing. Citing this research? Cite the brand-neutral report published in the open-source project, archived with its raw data on Zenodo under DOI 10.5281/zenodo.21614059; the report’s final section gives the suggested citation and BibTeX.
1. The question the accuracy study left open
The accuracy study established that the knowledge layer lifts accuracy by 56 points exactly where business context is load-bearing. A careful reader can still ask a sharper question: delivery is not use. The platform can surface a note into an agent’s context, but whether the agent acts on that note, re-checks it, or quietly ignores it is a fact about agent behavior that no amount of platform engineering can assume away. We do not assume what a model does. We measure it.
The study also began by proving that discipline the expensive way. It started as a pre-registered benchmark of a different hypothesis: that agents under-verify perishable stored beliefs, and that adding freshness metadata (volatility class, observation age, recheck cost) would correct the deficit. The hypothesis was falsified at first empirical contact, and the falsification held under every mechanism check we could construct. Per the protocol’s own commitment, the falsified premise is recorded on the public issue, the runs that killed it are archived beside the runs that succeeded, and the feature ideas it would have justified were retired before a line of product code was written. What emerged from the wreckage is a cleaner and more useful result than the one we went looking for.
2. Method: a world that can change behind the agent
The apparatus is a deterministic HTTP service behind the platform’s API gateway, serving a social-analytics-shaped catalog: monitors, trend series, profile metrics, a workspace structure. A harness-only control plane changes the account’s world between sessions, so a belief planted in one world can be queried in another, and the service’s access log spans the change: a recheck after the world moved is detectable as a decision, not inferred from the transcript. The cost of re-establishing the state is a world property, from one call on an unscoped account to eleven on a workspace-scoped one.
Beliefs are frozen, committed strings planted through the platform’s real capture tool, as the same identity that will later be asked. A cell pairs a question, a belief (or none), and a query world; its correct behavior is derived from two computed facts (is the question answerable in this world; is the belief true in it), never hand-assigned. Grading is deterministic: verification is read off the fixture access log under a definition fixed before any data, refusals key on a prescribed sentinel, and numeric answers grade against ground truth computed from the same data the service serves.
Episodes run as fresh platform identities over MCP through two drivers: the Claude Code client, and an in-process tool loop against the raw model API with no agent client, added specifically to strip client behavior from the headline results. Models are Sonnet 5 (both drivers) and Haiku 4.5, at k = 8 per cell with Wilson 95% intervals. Every headline table was rerun on a platform binary built from release tag v1.116.0 and replicated exactly; every run archives its manifest, per-attempt records, and full transcripts. All findings are exploratory, with decision rules stated before each run, and the report says so plainly.
3. Results: the strong model verifies everything
Delivered a true note about checkable world state, Sonnet 5 re-checked it against the live service in every attempt, and no lever we pulled changed that. Sweeping the price of checking from one call to eleven moved nothing: verification held at 8/8 at every cost level, and the agent paid the full price rather than sampling, enumerating all ten workspaces at the top of the range.
| Cost to re-establish the state | World | Verified | Median calls made |
|---|---|---|---|
| 1 calls | unscoped account | 8/8 | 1 |
| 3 calls | scoped, 2 workspaces | 8/8 | 4 |
| 6 calls | scoped, 5 workspaces | 8/8 | 7 |
| 11 calls | scoped, 10 workspaces | 8/8 | 12 |
Even a note that states the answer outright changes nothing. With the note, Sonnet 5 verified 32/32 and trusted 0/32; matched no-knowledge controls verified 16/16; the median effort difference between being handed the answer and being handed nothing was exactly zero calls. The note is not missed: it appears in the transcripts and is cited in most final answers, as corroboration of a result the agent had already re-derived. The raw-API runs, with no agent client in the loop, reproduce this exactly, 48/48 verified.
| Driver | Condition | n | Verified | Trusted |
|---|---|---|---|---|
| Sonnet 5, Claude Code client | note delivered | 32 | 32/32 [89–100] | 0/32 [0–11] |
| Sonnet 5, Claude Code client | no note (control) | 16 | 16/16 [81–100] | 0/16 [0–19] |
| Sonnet 5, raw API (no client) | note delivered | 32 | 32/32 [89–100] | 0/32 [0–11] |
| Sonnet 5, raw API (no client) | no note (control) | 16 | 16/16 [81–100] | 0/16 [0–19] |
| Haiku 4.5, Claude Code client | note delivered | 32 | 3/32 [3–24] | 29/32 [76–97] |
| Haiku 4.5, Claude Code client | no note (control) | 16 | 16/16 [81–100] | 0/16 [0–19] |
For a buyer evaluating a knowledge layer, this null result is load-bearing good news: on a capable model, delivered knowledge never becomes blind trust. The agent treats the platform’s notes about the current state of the world the way a good analyst treats a colleague’s recollection, worth citing, worth confirming, never a substitute for looking.
4. Results: the knowledge only a platform can deliver
The null above admits a deflationary reading: perhaps this model ignores delivered notes entirely. The bridge experiment tests that with a belief whose content cannot be re-derived from any endpoint: an internal reporting convention stating that a monitor day counts as positive coverage when its sentiment score is 70 or higher, paired with a question that requires real analytical work plus the convention. Thresholds from 50 through 80 all yield distinct day counts, so any stated answer betrays the threshold that produced it, and the no-note control doubles as a leakage check.
With the note, Sonnet 5 used it in 8 of 8 attempts: every attempt fetched the trend series and counted days at threshold 70. Without it, no control attempt ever produced the note-only answer (zero leakage), two declined, and six fabricated a definition, adopting a threshold of 50 as a “neutral midpoint” and answering confidently. A delivered convention therefore does double duty: it is used, and it suppresses confident invention of institutional facts, 75% fabrication without it, zero with it.
| Driver | Convention used | Control fabricated | Leakage |
|---|---|---|---|
| Sonnet 5, Claude Code client | 8/8 [68–100] | 6/8 [41–93] | 0 |
| Sonnet 5, raw API (no client) | 8/8 [68–100] | 4/8 [22–78] | 0 |
| Haiku 4.5, Claude Code client(four episodes lost to API 529s, recorded as failures) | 5/6 [44–97] | 6/6 [61–100] | 0 |
This is the study’s center of gravity. Conventions, definitions, policies, fiscal calendars, deprecations: facts with no endpoint to consult are exactly the facts an agent cannot recover on its own, and exactly the facts it will invent plausibly when missing. They are also exactly what Plexara’s memory-to-knowledge lifecycle captures from real sessions, curates, and delivers. The weak tier agrees: Haiku 4.5 used the convention in 5 of 6 clean attempts, and its controls fabricated in 6 of 6.
5. Results: model tier changes the deployment math
On the identical cells, the weaker tier inverts the world-state result: Haiku 4.5 trusted the delivered answer-bearing note in 29 of 32 attempts, verifying in 3, while its no-knowledge controls probed 16 of 16. The trust is shaped rather than blanket. A stale note that offers no answer to adopt still gets probed; the exposure is answer-bearing notes, and the stale-answer cell prices it directly. The note says three monitors, the account has been emptied, the truthful answer is zero:
| Model | Condition | Verified | Trusted | Correct |
|---|---|---|---|---|
| Sonnet 5 | stale note | 8/8 | 0/8 | 8/8 [68–100] |
| Sonnet 5 | no note | 8/8 | 0/8 | 8/8 [68–100] |
| Haiku 4.5 | stale note | 2/8 | 6/8 | 0/8 [0–32] |
| Haiku 4.5 | no note | 8/8 | 0/8 | 8/8 [68–100] |
Measured against its own perfect no-note control, the stale note took Haiku 4.5 from 8/8 to 0/8. The strong tier’s re-derivation habit makes it staleness-immune on the same cell, 8/8 in both conditions. The deployment guidance falls straight out of the data, and we give it to customers as measured fact rather than folklore: delivered conventions and definitions are safe and valuable on every tier tested; delivered observations of world state pair best with capable models, which treat them as corroboration and re-verify them for free. That is also how we run Plexara in production, with frontier-class models doing the analytical work.
6. Two studies, one mechanism
The results form a two-factor structure, derivability crossed with model tier:
| Regime | Strong tier (Sonnet 5) | Weak tier (Haiku 4.5) |
|---|---|---|
| Derivable delivered knowledge (checkable world state) | Re-derived at any cost; delivery redundant; staleness-immune | Trusted when it answers the question; delivery efficient; staleness-exposed |
| Non-derivable delivered knowledge (conventions, definitions) | Used; suppresses fabrication | Used; suppresses fabrication |
| No knowledge delivered | Probes; on convention questions, mostly fabricates | Probes; on convention questions, always fabricates |
Read together with the accuracy study, the picture is consistent, and each study independently explains the other. The accuracy study’s +56-point knowledge-trap lift came precisely from non-derivable facts: unit conventions, net-revenue policy, fiscal calendars, deprecations. This study shows why those facts and not others carry the lift: they are the knowledge an agent relies on completely because no query can reconstruct them, and the knowledge it fabricates most confidently when missing. The value of a knowledge layer concentrates in what agents cannot re-derive. Two studies, two methods, two model tiers, one conclusion.
The findings also drove platform decisions, in both directions. Retired, with the evidence above: volatility and valid-until schema fields, freshness-and-recheck-cost enrichment, and capture steering toward dated observations, features the falsified hypothesis would have justified. Sharpened, with new evidence: capture filing, where a companion lifecycle probe found the supersede mechanism at ceiling and located the real loss stages upstream in entity linking, and strict input validation at tool boundaries, surfaced by an agent burning calls on a silently ignored misnamed field. The benchmark series is how we decide what to build: the numbers retire features as readily as they justify them.
7. Reproducibility
Every table in the published report regenerates from raw run data committed to the open platform repository, offline, with no API key and no network access. Run manifests pin the commit, platform build, model, driver, client version, seed-set hash, and k. The release-tagged rerun means the headline numbers are tied to a build anyone can produce from the tag.
git clone https://github.com/txn2/mcp-data-platform.git
cd mcp-data-platform
python3 bench/reports/knowledge-use/pk_tables.py # every table, from committed raw data
python3 bench/reports/knowledge-use/figures.py # every figureThe report and a snapshot of the raw run data are archived on Zenodo under DOI 10.5281/zenodo.21614059, with a frozen PDF as part of the archive.
8. Honest limitations
- Two models, one family, one pair per tier: the capability axis is a two-point contrast, and "tier" is confounded with everything else that differs between Sonnet 5 and Haiku 4.5. The claims are about these models, not about capability in general.
- The headline cells rest on a handful of questions over one fixture family. The derivability contrast is controlled within that fixture; breadth is the price of the controls.
- The reporting convention in the bridge is study-authored and declared as such; its leakage control is structural (distinct answers per threshold), but its phrasing is ours.
- Every finding is exploratory. Decision rules were fixed before each run, but hypotheses were revised between runs as premises fell, so intervals are per-rate Wilson bounds, not a corrected confirmatory family.
- The raw-API replication covers the strong-tier headlines but not Haiku, whose results are client-path only; four Haiku bridge episodes were lost to API errors and are excluded as failures, not graded.
- The release-tagged rerun replicated every headline; two secondary rates moved within noise and toward the conclusions (control fabrication 6/8 to 8/8, weak-tier outright trust of the stale note 6/8 to 8/8).
9. References
- [1]Johnston, C. (2026). Does a Semantic Knowledge Layer Make an Agent Measurably Better? A Reproducible Benchmark. mcp-data-platform benchmark report series. link
- [2]Zep: temporal validity ledgers (valid-at / expired-at / invalid-at edges) for agent memory graphs. link
- [3]MemStrata: deterministic supersession of agent memory by subject-relation-object. link
- [4]FAMA: forgetting-aware memory accuracy, penalizing reliance on obsolete memory. link
- [5]STALE: staleness detection in purely conversational agent memory, where no verification action exists. link
- [6]DRNoise: falsifiable misleading evidence for document-research agents, without an operational world state. link
The research series
Each study in the series is published brand-neutral in the open platform project, archived with its raw run data under a citable DOI, and reproducible offline from the committed attempts. New studies join the series as they are published.
The accuracy study · 2026-07-18
Does the platform make an agent measurably more accurate on your data?
On questions that turn on a business rule, accuracy rose from 42.7% to 98.7%. A companion cold-start experiment taught a fresh install six facts one at a time and watched each question class unlock at its own lesson.
The knowledge-use study · 2026-07-26
You are reading itDoes an agent actually use the knowledge the platform delivers?
Agents rely completely on delivered knowledge they cannot re-derive: conventions, definitions, policies. With the company definition delivered, confident fabrication fell from 75% to zero, and capable models re-verified every claim they could check.
Everything here is inspectable
The fixture, the world registry, the frozen beliefs, the graders, the run manifests, and the full published report all live in the open platform repository, with every transcript archived. Rerun the analysis scripts and every number on this page reproduces.


