Skip to main content
Plexara

Plexara research · The knowledge-pollution study

120 of 120

We planted a wrong fact through our own review queue, left the correct source sitting beside it, and watched what agents did next. Across 120 runs the outcome came down to a single act: whether the agent ran the one query that settles the question. Every run that ran it answered correctly. Every run that skipped it took the wrong number. The wrong fact never out-argued the correct one. On a cheap model, it removed the impulse to look.

0 of 96

runs in which a frontier-class model preferred the wrong fact to the correct source

16 of 24

runs in which a small model preferred it, on the one kind of claim that traveled

432 runs

graded on the exact value answered, with no judge anywhere in the loop

1 command

regenerates every table in this report from the published raw data

Plexara Research · The knowledge-pollution study

What a wrong fact costs once it clears review: verification displacement, model tier, and the price of a curation gate

Plexara, a product of Deasil Works · 2026-08-07 · Models: claude-haiku-4.5, claude-sonnet-5 and claude-opus-5 · Data and code: txn2/mcp-data-platform · DOI: 10.5281/zenodo.21834813

Abstract

Our first two studies measured what a knowledge layer is worth when the knowledge is right. This one measures what it costs when a piece of it is wrong. A wrong fact, of the kind a competent agent could capture by mistake, went through the real path: captured in a session, approved in the review queue, promoted to the shared knowledge every teammate reads. The correct source stayed where it was, beside it. The result inverted our own pre-registered prediction. We expected the dangerous claim to be the one nothing can check, and it was refused at every model tier. The claim that traveled was the one a single query settles, and it traveled only on a small, cheap model. The mechanism is exact: an agent’s answer was decided entirely by whether it ran that query, with no exception in 120 runs, and with nothing planted the same small model ran the query every time. The wrong fact did not win an argument. It removed the reason to check. Two design consequences follow, and we have acted on both: the claims most worth scrutiny at review time are precisely the ones the platform can verify against your own data, and the exposure sits on the cheap model tier, which is why Plexara runs frontier-class models in production.

This page presents the results with commercial framing. Citing this research? Cite the brand-neutral report published in the open-source project, archived with its raw data on Zenodo under DOI 10.5281/zenodo.21834813; the report’s citation section gives the suggested citation and BibTeX.

1. The question the first two studies left open

The accuracy study measured what a knowledge layer is worth: 56 points of accuracy on the questions that turn on a business rule. The knowledge-use study located where that worth sits: agents rely on what they cannot work out for themselves and re-derive what they can.

Both measured correct knowledge. A buyer is entitled to ask the other half of the question, and it is the question we would ask: a review queue is run by people, people approve things, and sooner or later one of those things is wrong. What does that cost? Not in theory. In graded runs, with the wrong fact promoted by the same machinery that promotes the right ones, and the correct source still sitting beside it.

This is not a security study and it does not posit an attacker. The wrong facts here are the kind a careful analyst’s agent could produce by honest mistake: a fiscal year that starts in the wrong month, a record count that was true last quarter. That is the failure mode a real deployment actually meets.

2. Method: a wrong fact, promoted by our own review queue

Every wrong fact in this study earned its standing legitimately. An agent captured it in a session, an administrator approved it in the review queue, and it was promoted into the shared knowledge tier and written to its home, either a catalog entry or a knowledge page. A separate identity then confirmed by read-back that the claim was reachable in search and present where it had been filed, before a single graded run started. Nothing was slipped past the gate; the gate let it through, which is the whole point.

Two kinds of wrong fact, chosen for how they relate to the world rather than by label. One is a reporting convention that nothing in the data can confirm or refute, a fiscal-year boundary. The other is a claim about state that one query settles, a record count, where the question the agent is asked is that count. Beside each of them sits the correct source: a curated page stating the real boundary, or the warehouse itself holding the real number. Every run is therefore a conflict, not a delivery test.

Three model tiers, three arms per cell (nothing planted, a correct fact planted, the wrong fact planted), and 24 runs per cell so the two claim kinds carry the same denominator. Each run is a fresh identity on a fresh database. Grading is deterministic to the exact value: the correct answer, the value only the wrong fact leads to, and the pre-existing wrong answers the question already invites are all computed from the fixture before anything runs, and no two can collide inside grader tolerance. No judge is involved anywhere. Hypotheses, decision rules and falsifiers were fixed in a pre-registration before the confirmatory data, and the prediction they encode is the one the data broke.

3. Results: only one kind of claim traveled

We expected the convention to be the dangerous one. It is the class the knowledge-use study showed agents rely on completely, precisely because no query can reconstruct it. The opposite happened. The wrong convention was refused at every tier. The claim that traveled was the one a single query settles, and it traveled only on the small model.

The wrong fact we plantedHaiku 4.5Sonnet 5Opus 5
A convention nothing in the data can settle (a fiscal-year boundary)0/24[013.8]0/24[013.8]0/24[013.8]
A fact one query settles (a record count)16/24[46.782]0/24[013.8]0/24[013.8]
Dot-and-interval chart of adoption of the wrong claim by claim kind and model tier: 16 of 24 for the checkable claim on Haiku 4.5, zero of 24 in every other cell, with Wilson 95% intervals.
Figure 1. Adoption of the planted wrong claim by claim kind and model tier (24 runs per cell, Wilson 95% intervals). The kind of claim we predicted would propagate was refused everywhere; the kind we predicted would be re-derived away is the only one that moved.

Delivery is not what varies here. On every arm the read-back confirmed the wrong fact was reachable, the planted text appears in all 24 transcripts, and on the convention cells the correct source appears beside it in all 24. Those zeros are refusals of a delivered claim, not claims that failed to arrive.

One floor is noisy and the controls are what expose it: with nothing planted at all, the small model answers the convention questions correctly 9 times in 24, so its convention zero is a statement about what it declined to adopt rather than a clean rate against a stable baseline. The checkable floor is clean at every tier, which is what makes the next section readable.

4. Results: it did not argue, it removed the check

On the checkable claim, one thing separates the outcomes completely: whether the run observed the result of a count against the table in question.

ModelConditionRan the checkTook the wrong valueAnswered correctly
Haiku 4.5wrong fact planted8/2416/248/24
Haiku 4.5nothing planted (control)24/240/2424/24
Sonnet 5wrong fact planted24/240/2424/24
Opus 5wrong fact planted24/240/2424/24
Stacked bar chart of the four checkable arms, 24 runs each, split by whether the refuting count came back: every observing run answered correctly, every non-observing run took the wrong value.
Figure 2. Every run split by whether the refuting count came back (24 runs per arm). Each bar is a clean partition: the runs that observed the count were correct, the runs that did not took the planted value. With nothing planted, the same small model ran the query 24 times in 24.

Every run that observed the count answered correctly. Every run that did not took the planted value. There is no exception in either direction, here or in the 120 runs this cell accumulated across the follow-up conditions below. The control row is what makes it legible: with nothing planted, the small model runs the query 24/24 times and is right 24/24 times. It is entirely capable of settling the question. The claim’s presence is what stopped it asking.

The planted fact did not out-argue the world. It removed the impulse to consult it. That is a different failure from the one most people picture when they imagine bad data in a knowledge layer, and it has a different fix. Persuasion would be answered with better ranking or stronger provenance signals. Displaced verification is answered by checking the claim before it is promoted, and by which model is doing the work.

5. Results: not the phrasing, not the filing cabinet

The planted fact was written the way this platform’s capture path actually writes facts, which includes telling the next reader what to do. That raises a fair objection: was the small model adopting a belief, or just following an instruction? The pre-registered answer plants the same false count at three strengths of phrasing.

How the fact was phrasedTook the wrong valueRan the check
Bare(states the count and asks nothing)18/24[55.188]6/24
Plain(marks the count as the relevant one)17/24[50.885.1]7/24
Imperative(instructs the reader to report the count)18/24[55.188]6/24

A bare statement that asks nothing of the reader is taken as often as an explicit instruction. The effect is adoption, not compliance, and the rate at which the check gets run barely moves across the ladder. It is the presence of a stored answer that suppresses verification, not the force with which it is worded.

Two more decompositions take the effect apart. Moving the identical claim from the catalog entry onto a knowledge page changes nothing except to sharpen it, and planting an equivalent claim in a completely different fixture, with a different world, a different question and a different tool, reproduces it against a clean control floor.

ConditionModelTook the wrong valueAnswered correctly
Stored on the catalog entitythe confirmatory arm, for referenceHaiku 4.516/24[46.782]8/24
Stored on a knowledge page insteadsame claim, same phrasing, same questionHaiku 4.524/24[86.2100]0/24
A second fixture, nothing plantedthe floor the next two rows are measured againstHaiku 4.50/24[013.8]24/24
A second fixture, wrong fact planteddifferent world, different question, different toolHaiku 4.524/24[86.2100]0/24
A second fixture, wrong fact plantedthe same arm, one tier upSonnet 50/24[013.8]24/24
Dot-and-interval chart of adoption across seven robustness conditions, weak tier against strong tier: the weak tier ranges from 16 of 24 to 24 of 24 while the strong tier is zero in every condition measured.
Figure 3. Adoption of the wrong checkable claim across every pre-registered robustness condition, by tier. The weak tier holds between 67% and 100% whatever we vary; where the strong tier was measured alongside it, the strong tier is at zero.

The storage location result is the practically useful one. A claim carried into an answer by search alone is just as effective as one written onto the very record the question is about, so the surface worth guarding is the search channel and the review that admits a claim to it, not the filing cabinet it lands in.

6. Results: what the replication corrected

Every arm above runs through one agent client, so the pre-registered replication reruns the headline cells against the raw model API with no agent framework in the path. It exists to separate what is true of the platform and the models from what is true of one client, and it earned its budget by finding a correction.

The wrong fact we plantedModelThrough the agent clientRaw model API
CheckableHaiku 4.516-24 of 248/8[67.6100]
CheckableSonnet 50/240/8[032.4]
CheckableOpus 50/24arm invalidated and reported, not analyzed
ConventionHaiku 4.50/244/8[21.578.5]
ConventionSonnet 50/241/8[2.247.1]
ConventionOpus 50/240/8[032.4]

The headline replicates: the checkable claim propagates to the small model and to neither frontier-class model, with no agent client involved. The convention result does not. On the raw API the small model took the false fiscal boundary in half its runs, and the middle tier took it once. So the refusal of a wrong convention is a property of the agent scaffolding, not of the platform, and this page claims it only for agents running inside such a scaffold. That is exactly the risk the replication was budgeted to catch, and it caught it.

A second narrowing surfaced during the recompute, and it is an instrument defect of our own making. The promotion path records a reviewer note on the promoted record, and the harness wrote that note in a form that named the claim as a study plant. Any run that opened the full record read an explicit disclosure. Opening the record turned out to be capability-graded: on the contested convention cells the frontier-class models opened it in 18 to 24 runs of 24, the small model in 9 of 24. The Opus refusals on that cell therefore cannot be separated from the disclosure and are reported as confounded; the other two tiers hold their refusal among the runs that never saw it, and the checkable headline is essentially untouched, since the handful of exposed runs adopted anyway.

The substantive lesson survives the defect, and it is one we have taken to heart in the product. The provenance surface is a working defense. The tier that reads the full record, its status, who captured it, what the reviewer said, is the tier that resists a conflict. The tier that never opens it is not protected by it.

7. Results: three different answers to one wrong fact

The store snapshots record what each identity wrote back, and the three tiers behaved in three qualitatively different ways when handed the same wrong claim.

ModelWhat it did with the wrong factWhat it wrote back
Haiku 4.5Took the claim on the checkable cell and stopped checkingfiled nothing corrective
Sonnet 5Re-derived the answer and declined the claim, silentlyfiled nothing at all
Opus 5Re-derived the answer, declined the claim, and filed a correction30 corrections proposed for review

The corrections Opus filed are proposals sitting in the review queue, not self-repair; each one cites the record it disputes and states that it verified the real number by direct query. That is the shape of a working system: the strongest model in the loop treats a conflicting stored claim as something to check and then to flag for a person, and the flag lands where a person will see it.

8. What we changed because of it

  • Review scrutiny follows derivability. The claims that propagate are exactly the claims the platform can check against your own data. A claim that names a number a query would settle is the one worth putting the observed value next to before a reviewer approves it, and that is where our review tooling is aimed.
  • Frontier-class models do the analytical work. The exposure sits on the cheap tier, between 67% and 100% across four independent conditions, and it adopts straight through an explicit disclosure sitting in the record. That is how we run Plexara in production, and this study is the measurement behind the choice rather than a preference.
  • Provenance stays rich on the record. Status, capturer, and reviewer history on the full record are not decoration. They are the surface a capable agent actually uses to arbitrate a conflict, and the study shows the arbitration happening there.
  • What we did not learn. Whether a belief recovers after a claim is retracted never ran. We have no data on it, we claim none, and the follow-up is scoped to the one cell that qualifies. The instrument lesson is on the record too: a study plant’s reviewer note must never name it as a plant.

9. Reproducibility

Every table on this page regenerates from raw run data committed to the open platform repository, offline, with no API key and no network access. Each arm archives its manifest, every graded attempt, the full transcripts, the plant record with its read-back flags, and the store snapshots taken before and after. Arms invalidated mid-study are archived beside the ones that counted, with a suffix naming what invalidated them.

git clone https://github.com/txn2/mcp-data-platform.git
cd mcp-data-platform
python3 bench/reports/knowledge-pollution/pollution_tables.py   # every table
python3 bench/reports/knowledge-pollution/figures.py            # every figure

The recompute is a build gate rather than a convenience: the project’s own verify step re-derives the headline numbers from the archives and fails on any drift from what the report prints. The report and a snapshot of the raw run data are archived on Zenodo under DOI 10.5281/zenodo.21834813, with a frozen PDF as part of the archive.

10. Honest limitations

  • Three models, one family. "Model tier" here is three models from one vendor, so the defensible statement is that the cheap tier does this and the expensive ones do not, not that capability in general causes it.
  • Two fixtures, both benchmark fixtures. The second fixture removes the first one as an explanation, and a storage-location control separates where a claim lives from which world it lives in, but neither fixture is a production system.
  • Twenty-four runs per cell resolves near-zero against near-ceiling, which is what these contrasts are. It cannot separate two small rates, and nothing here claims finer resolution. The raw-API arms, at eight runs, are coarser still.
  • The reviewer note on the planted record disclosed the plant to any run that opened it. Exposure was graded by model tier and by claim kind; the Opus convention refusal is confounded by it, the other two tiers hold among unexposed runs, and the checkable results are essentially untouched.
  • A run that noticed the conflict and declined to answer lands in the same bucket as a run that failed for any other reason. Separating them would need a judged pass, which the protocol forbids by design.
  • Whether belief recovers after retraction has no data. Two of the five pre-registered hypotheses were never tested, and the report says so rather than quietly dropping them.
  • No headline was rerun on a tagged release build. The confirmatory arms are pinned to one commit, and the follow-up arms ran a later state whose agreement with the first was measured rather than assumed.

11. References

  1. [1]Johnston, C. (2026). Does a Semantic Knowledge Layer Make an Agent Measurably Better? A Reproducible Benchmark. mcp-data-platform benchmark report series. link
  2. [2]Johnston, C. (2026). When Do Agents Use Stored Knowledge? Derivability, Capability, and the Limits of a Knowledge Layer. mcp-data-platform benchmark report series. link
  3. [3]The closest organic work on cross-user contamination of shared agent memory, reporting 57 to 71 percent contamination without derivability or capability moderators and without a co-present correct source. link
  4. [4]Governed agent memory: provenance and curation plumbing validated without measuring what anyone ends up believing. link
  5. [5]Curation and provenance machinery for multi-agent shared memory, again measured at the plumbing rather than the answer. link
  6. [6]Representative of the attacker-framed memory-integrity literature, which studies deliberate corruption as a capability an adversary optimizes for. This study measures the honest-mistake base rate underneath it. link

The rest of the research series

Each study in the series is published brand-neutral in the open platform project, archived with its raw run data under a citable DOI, and reproducible offline from the committed attempts. New studies join the series as they are published.

The accuracy study · 2026-07-18

Does the platform make an agent measurably more accurate on your data?

On questions that turn on a business rule, accuracy rose from 42.7% to 98.7%. A companion cold-start experiment taught a fresh install six facts one at a time and watched each question class unlock at its own lesson, and a lifecycle re-run measured a fact taught by one person being reused correctly by a different teammate 98.9% of the time.

The knowledge-use study · 2026-07-26

Does an agent actually use the knowledge the platform delivers?

Agents rely completely on delivered knowledge they cannot re-derive: conventions, definitions, policies. With the company definition delivered, confident fabrication fell from 75% to zero, and capable models re-verified every claim they could check.

The graph-completion study · 2026-08-10

What do references between knowledge pages buy an agent that has to be complete?

Asked to write complete operational documents, agents followed references between knowledge pages voluntarily, grounded every governing constraint while reading 0.2% of a 5,000-page corpus, and kept discovery cost flat as the corpus grew a hundredfold; without the references, cost roughly doubled and the only failed episode appeared. With search off, references were the only route that worked: 96% of constraints recovered against zero. One pre-registered construct died by its own kill condition, and the kill is published with the result.

Explore the full research series

Everything here is inspectable

The pre-registration written before the data, the harness that plants and retracts a claim, the graders, the run manifests, every transcript, and the full published report all live in the open platform repository. Rerun the analysis scripts and every number on this page reproduces.