RealEstateAI
Workshop

A newsroom that checks itself (and knows when to stop)

How we write reader articles with AI: drafts only from the facts in the brief, four checks before publishing, and one article a week that goes out without human review only if every hard claim is confirmed.

2 October 2026 · 7 min

A study corner with a desk by a window onto greenery (frame from an AI-animated video).AI · video animato

On our websites we keep a small library of articles for home buyers: how the notarial deed works, what to check in a land registry extract, how the market is moving. We write them with AI. Since September, on one of our sites, one article a week goes out without anyone having read it first. That sounds like the opposite of caution. In fact it's the part of the project where we put the most checks. Here's how it works, and where it broke.

The starting rule: only the facts in the brief

Every article starts as an idea with a brief and, ideally, some sources. The writing engine has one non-negotiable instruction: use only facts that appear in the brief and the sources. It must never invent numbers, percentages, regulations, deadlines, names or examples. If a section is missing a piece of data, the engine doesn't fill the gap. It writes the section without it and flags in the internal notes which section is missing which fact.

An article with neither brief nor sources isn't generated at all: it would come out invented. We paid for this in a different way. The first ideas in the queue had a brief but no sources, and the articles came out thin. That told us the problem was upstream, not in the model.

There's a less visible precaution too. A brief can contain text pasted in from elsewhere, and pasted text can hide instructions in disguise. The article's values are neutralised before they enter the prompt, so they can't close or reopen the structure of the request. The structure is ours. The data is only data.

One language at a time

We write in several languages. The first version of the generator produced everything in one enormous call. For almost two weeks this summer it produced no drafts at all, with no error anywhere: the call simply timed out. Today the engine writes one language at a time, the main language first, then the others as rewrites of it. Without the main version, the run stops rather than reinventing the facts in another language.

Two more rules come from the same failure. Time is checked ahead: a pass doesn't start if it can't finish in the time left, because a process killed halfway writes nothing and the tokens are paid for anyway. And the engine only writes into empty fields: whatever a person has written always wins.

That last rule had an instructive side effect. An article went live with its title still the scribbled note the idea had started from, all lower case. The engine had followed the rule, since the title had been "written by a person", and translated everything else. Nobody had told anyone. Now a title that looks like a note is a blocker, not a detail.

Four checks, and any one can stop everything

Before an article goes out without a person, it passes four checks in a row.

  1. Readiness. The basics: a title that looks like a note, a body that's too short, missing languages.
  2. Reserved names. The same guard we use on outgoing mail: the internal names of properties and owners never go out, not even by accident. If the guard isn't armed (for instance, it can't load the list of names), nothing is published. It fails closed, not open.
  3. Claim extraction. A model reads the main text and lists every verifiable statement, sorting it into "hard" or "soft". Hard claims are rates, amounts, rules, deadlines and obligations: if they're wrong, they cause harm. Soft claims are descriptive.
  4. Web verification. A second pass, with real web search, checks each claim against sources: first the ones the article cites, then official ones. Each claim comes back confirmed, contradicted or unverifiable, with the address of the page actually consulted.

The final gate is a simple function that can be tested without a model and without a network: publish only if every hard claim is confirmed and none is contradicted. A single unverifiable hard claim stops everything. Better an explicit gap than an unproven figure. The sources the check consulted are merged into the article's sources, without duplicates.

The rhythm: one a week, in total

On Monday the system prepares a batch of drafts. Tuesday is the window for people: a reminder goes out, and anyone who wants to can read and schedule. On Wednesday, if nothing has already gone out that week by human hand, the system takes the oldest complete draft and puts it through verification.

  • If verification passes, the article goes out. Whoever follows the editorial work gets a reminder with the link and the number of claims checked, plus an invitation to read it.
  • If it fails, the draft gets a block marker and the week stays empty. Only a person can remove the marker, after fixing the draft.
  • If there's a technical fault, it tries again the next day.

The cycle has an end date written into its configuration. After that date it switches itself off. It's an experiment, and it behaves like one.

What the check is not

Automated verification is not a human signature, and it doesn't pretend to be. Our system has a field that means "a person has read this". For articles that went out on their own, that field stays empty. When a person reads and confirms, they fill it in themselves. The verification result lives in the internal notes, where anyone can look it up.

And readers know. Every article carries a note at the top saying it was written with AI, and a box at the end explaining this, with a reference to Article 50 of the EU AI regulation. Without human review there's no alternative to the label, and that's as it should be.

Where it broke, and what's still open

Not everything is fixed, and we keep track.

  • The topic explorer dies of timeouts. The function that looks for new ideas on the web runs past the maximum execution time. Until we put a cap on each call, the idea queue is refilled by hand. If the queue is empty, the system says so: "nothing goes out this week".
  • One language longer than the rest. In one of our languages the same text uses many more tokens. The translation came out truncated, was discarded and retried on every run, forever. The higher cap now applies only to that language.
  • Languages not read by native speakers. Two of our languages remain marked "to be reviewed", and the warning only clears when a person records that they've read them. The system can't do that for them.

We don't think an AI should write unchecked. We think that if it writes, every important claim should be checked more thoroughly than a person in a hurry would check it. And when the check isn't enough, the system should know how to stay quiet.

#editorial#fact-checking#llm#transparency

Keep reading