Skip to main content

Plexara research

We test Plexara in public

The performance claims on this site come from controlled studies, and we publish the whole study: the method, the raw run data, and the code that turns one into the other. If a number looks too good, don’t take our word for it. Download the runs and check. When the data kills one of our own ideas, we publish that too.

What we publish with every study

A benchmark you cannot rerun is an ad. Each of ours ships as a working repository, and these four things come with it:

The report
A brand-neutral technical report in the open-source project, with a frozen PDF and a snapshot of its data archived on Zenodo under a citable DOI. Other vendors and researchers are welcome to cite it, argue with it, or build on it.
The raw runs
Every graded attempt, full transcript, and run manifest, committed to the repository. If you want to know why one answer was marked wrong at 2:14 on a Tuesday, the transcript is in the tree.
The code
The harness, the deterministic fixtures and seed generators, the graders, and the analysis scripts. Every table and figure regenerates offline from the committed data. No API key, no network access.
The protocol
Decision rules written down before the data comes in, and headline runs pinned to tagged releases so the exact build behind a number stays buildable. When a pre-registered hypothesis fails, the failure stays on the public record next to the runs that killed it.

Why we benchmark

Plexara runs in production for real clients, answering real questions every day, and that field record tells us the approach works. Field experience does not tell you by how much, on which kinds of question, or where to aim the next round of engineering. For that you need a number you did not choose: a controlled study that holds everything constant except the thing under test.

The studies also decide what we build. When the knowledge-use study falsified our pre-registered hypothesis about staleness metadata, we dropped the features that hypothesis would have justified before writing a line of product code, and the same data pushed capture filing and input validation up the roadmap instead.

The research series

Each study in the series is published brand-neutral in the open platform project, archived with its raw run data under a citable DOI, and reproducible offline from the committed attempts. New studies join the series as they are published.

The accuracy study · 2026-07-18

Does the platform make an agent measurably more accurate on your data?

On questions that turn on a business rule, accuracy rose from 42.7% to 98.7%. A companion cold-start experiment taught a fresh install six facts one at a time and watched each question class unlock at its own lesson.

The knowledge-use study · 2026-07-26

Does an agent actually use the knowledge the platform delivers?

Agents rely completely on delivered knowledge they cannot re-derive: conventions, definitions, policies. With the company definition delivered, confident fabrication fell from 75% to zero, and capable models re-verified every claim they could check.

Start from the archives

Each Zenodo record holds the frozen report PDF and a snapshot of the raw run data behind it, hosted independently of us and of this site. The open repository holds the harness and every run since.