Every article shows its work.
Open any article and you see what it was written from and what was checked - and if something could not be put right, it says so, with what it would take to fix.
I built this panel before anyone asked for it. I was the one signing the articles, and a green tick I could not check was worth nothing to me. Every article carries the sources it used, the checks it passed, and what could not be put right.
Several checks, each on the record
No single pass decides. Independent checks run against every draft, and each one records what it found - so the article arrives with its own account of itself rather than a verdict you have to trust.
Nothing ships silently. A check that fails sends the article back to be written once more, and if it still falls short it arrives saying so - never quietly passed off as clean.
Sources vetted before a word is written
Before drafting starts, an AI judge checks each candidate source against the specific story - the same event, the right date window, real substance behind the claim.
A topic that can't be sourced is dropped, not padded with a thin source just to hit a count.
An authenticity trace on every article
Every published article carries a trace: the sources it used, the checks it passed, and the named editor of record.
Refined across 20,000+ published articles
The gates weren't tuned once and left alone. They've been refined and tested across more than 20,000 published articles - the bar they hold today is earned, not assumed.
Two measured verdicts, and neither needs a model to judge
Every article is measured on two things no model has to have an opinion about: how hard it is to read against the reading level you set, and how much of it appears word for word in one of its own sources. Readability is graded with Flesch-Kincaid in English and with LIX in the continental languages.
Each measure has three answers rather than two: a tick when it passed, amber when it failed, and not recorded when it could not run at all - too little text, a language neither formula is calibrated for, or no sources to compare against. A measure that could not run never draws a tick.
A recorded failure at either one, or at the editor-in-chief, sends the article back to be written again, once. When that happens the panel says so and names what sent it back, and the discarded draft's verdicts are kept - so a rewrite cannot be quietly presented as the whole story.
A floor you set, and a run that misses it says so
You set the minimum number of sources an article should be built on. The default is two. It tries for it, and the article is written and delivered either way, because we do not refuse you the work.
What changed is that the shortfall is recorded and shown. The panel reads 'Only 1 source of the 2 you asked for' in amber and the Sources step carries no tick, where it used to draw a green line about having researched the piece.
The reason that mattered: measured before the fix, 3 of 9 real runs delivered under the floor, most often on evergreen topics. An article written before it shipped has no floor recorded at all, so the panel states the count with no tick rather than guessing what the bar had been.
What we could not fix
The desks had been writing down the problems they could not repair for a long time before anything rendered it. One recorded run carried seven open problems, and the article read as clean.
The authenticity panel now carries them: each entry in the desk's own words, with the repair it would take, attributed to the desk that raised it - fact-check, SEO, or editor-in-chief - and a count on the panel header, since the panel is collapsed by default. An entry appears exactly when a check withheld its pass, so this is the detail behind the amber row, never a second verdict.
A desk that did not do the work is not shown as having done it
The trace lists the desks that ran. It used to list every desk the engine called, which is not the same thing: a desk can be called and do nothing, either because it caught its own failure or because there was nothing to work on - no article to fact-check, no sources to check it against.
Both used to arrive as a completed step, so a run where three desks never touched the article still showed nine of nine. Now a desk that failed or stood down appears on the timeline without a tick, the desk strip lights only the desks that worked, and the count is the real one.