The Checks We Run Before AI Writing Ships

A Gate Is Worth Building Only If It Runs in Seconds
The short answer is that an automated gate earns its place through speed, not rigor. Ours runs five checks in about 1.5 seconds including interpreter startup, with roughly 170 milliseconds of that spent on the checking itself. At that price it runs on every edit and across the entire queue whenever a rule changes, which is what makes it real rather than ceremonial.
Some scale, so you know what this is drawn from. The gate guards nine blogs publishing one post per day each, 528 posts so far, with 199 drafts queued behind them. Its rules live in two files totaling 867 lines.
What follows is what it checks, what it deliberately refuses to block, and the two things it cannot see at all.
Five Checks and What Each One Actually Blocks
- ● Naturalness - sentences over 28 words
- ● Mobile readability - paragraphs over four sentences
- ● Advertising standards - length and substance
The gate is five stages, and a draft has to clear all of them before anything can publish it. Each stage owns one question.
Naturalness looks for the tells of unedited generated prose. Sentences past 28 words, paragraphs past four sentences, and a heading that runs into its first sentence past 30 words combined.
Mobile readability checks the same shape from the reader’s side, since a paragraph that scans fine on a laptop becomes a wall on a phone.
The advertising-standards stage checks the things a review would check anyway: length, description length, and whether the page has enough substance to sit next to an ad.
Topic fit compares the title and keyword against the cluster the file claims to belong to. Differentiation is the last stage and the strictest, because it compares the draft against everything already published on that blog.
What the Differentiation Stage Looks For
Uniformity was the failure that started all of this. Individually decent posts, all built to the same skeleton, read as one template repeated.
So this stage refuses generic section headings outright. Introduction, Feature Comparison, How to Choose, and Conclusion are all rejected, which forces every heading to be a sentence about that specific post.
It also compares titles against published siblings. A subtitle already used by two or more posts blocks, and a year-tail title blocks once 30 percent of that blog already ends the same way.
The last piece is a required verdict section, whose heading must contain a phrase like which fits or who should. A post without an opinion is the exact thing that reads as filler.
Why Some Rules Warn Instead of Blocking
- ● Blocked drafts become held and leave the rotation
- ● Title repetition resolves itself once another post ships
- ● Time dependent rules warn instead of blocking
The design rule we arrived at is narrow and worth stating plainly. A gate blocks only defects that live inside the file, and everything else warns.
The reason is mechanical. A blocked draft becomes held, and held means removed from the publishing rotation, so a rule that depends on what shipped yesterday can silently drain the queue.
Title format repetition is the clean example. Three posts in a row sharing a format is worth flagging, but it resolves itself as soon as a different post publishes. Blocking on it would hold a draft for a reason that expires on its own.
Figure count works the same way. Too few concrete numbers in a commercial post is a real weakness, so it warns and someone adds the numbers, rather than freezing the draft over a judgment call.
The Rule We Measured Before We Enabled It
- ● 174 of 189 planned drafts would have failed
- ● Enabling it unmeasured would have held nine blogs
- ● Drafts were fixed first, the rule switched on after
The newest rule requires a verdict and a real number inside the first 120 words. No preamble, no context-setting, just the answer.
We wrote it, and then instead of enabling it we ran it across the queue as a report. Of 189 planned drafts, 174 failed.
Turning that on unmeasured would have held nine blogs the same night. Every individual rejection would have been correct, and the pipeline would have stopped anyway.
We fixed the drafts first and enabled the rule second. A fresh count today shows all 199 queued drafts carrying a verdict inside the opening, which is the same destination reached in the order that did not break anything.
The Review Flag That Expires by Itself
One stage cannot be automated, and pretending otherwise was our most expensive mistake in this whole system. Naturalness needs someone to actually read the draft, so the file carries a flag saying the reading happened.
A boolean flag is a promise, and promises rot. A generator rewrote a body, a bulk edit touched a paragraph, and the flag stayed true, so the file claimed a review that no longer described its contents.
The fix was to make the flag depend on the text instead of on trust. Alongside the flag we store a fingerprint of the body at review time, and the gate recomputes that fingerprint on every run.
stored = post.get("naturalness_reviewed_hash")
if stored != body_fingerprint(post.content):
fail("body changed since review — review again")
What we like about this is where the enforcement lives. We could have added a reset to every script that writes a body, and we would have missed one, then missed another when an agent wrote the next script.
Putting the check in the reader instead means it holds no matter who edited the file or how. A human edit, a generator, a bulk repair pass, an agent working from a different directory, all of them invalidate the review automatically because all of them change the text.
What This Gate Cannot See
It reads text and structure. That means it can confirm a number is present and can never confirm the number is right.
It also cannot see a page. Every check runs against the source file, so anything that only exists after rendering, such as a broken image or a heading colliding with a table, passes silently.
There is a third gap that took longer to admit. The rule set only knows about failures we have already had, and every check in it exists because something specific went out wrong.
That makes the gate a record of past mistakes rather than a definition of quality. It is genuinely useful and genuinely backward-looking, and treating a green result as proof of a good post is the mistake it invites.
Cost and Coverage of Each Stage
| Stage | Blocks on | Typical catch | Needs published history | Can it be wrong |
|---|---|---|---|---|
| Naturalness | Sentence and paragraph limits | A 38-word sentence | No | Rarely, limits are mechanical |
| Mobile readability | Paragraph shape and length | A five-sentence block | No | Rarely |
| Advertising standards | Word count, description length | A 174-character description | No | No, both are counts |
| Topic fit | Cluster and keyword mismatch | A draft filed under the wrong cluster | No | Yes, it warns on borderline titles |
| Differentiation | Generic headings, repeated subtitles, missing verdict | A reused subtitle | Yes | Yes, hence the warn-only rules inside it |
The fourth column explains the design. Only one stage needs to know what already shipped, and that is precisely the stage where most of the warn-only rules live.
Which Version of This Fits Your Setup
If you review everything an AI writes: you need the mechanical half only. Length limits and structure checks save reading time, and your judgment already covers the rest.
If you publish on a schedule with occasional review: add the differentiation stage. Sameness across posts is invisible one file at a time and obvious across thirty, which is exactly the comparison a person skimming will not make.
If nothing is reviewed before it ships: measure every new rule as a report before enabling it, and count how many items sit in your rejected state every morning. A blocking rule you have not measured is a scheduled outage.
If accuracy matters more than consistency: this shape of gate is the wrong tool. Verifying claims needs source checking, and no amount of text analysis substitutes for opening the page the number came from.
What We Would Build First Next Time
Start with the two counters, not the rules. A daily count by state, and a count of how many items each rule would reject if it were on, tell you more than any individual check.
Then add rules one at a time, each one traceable to something that actually went wrong. A rule with no incident behind it is a guess, and guesses are what make gates slow and resented.
If you want the specifics of what generated the incidents behind ours, the Blogger API reference covers the platform half. For the content standards these rules are trying to approximate, Google’s helpful content guidance is the document worth reading before writing any rule at all.
Frequently Asked Questions
How fast does an automated quality gate need to be? Ours runs in about 1.5 seconds end to end, including interpreter startup, and the checking itself takes roughly 170 milliseconds. That cost is the whole reason it runs on every edit instead of once before shipping.
Which rules should block and which should only warn? Block only defects that live inside the file itself, such as a missing section or a sentence over the length limit. Anything that depends on what shipped yesterday should warn, because blocking on it can empty your queue overnight.
What can a text-based gate not catch? It reads text and structure, so it verifies that a number is present and never that the number is correct. Screenshots, rendered pages, and factual accuracy all need a different kind of check.
How do you introduce a new rule without breaking a running pipeline? Run it as a report across everything you already have. When we tested one new rule that way, 174 of 189 queued drafts failed it, so we fixed the drafts first and turned the rule on second.
Does a quality gate go stale? Yes, in the sense that the rule set only knows about failures you have already seen. Every check in ours exists because something specific went out wrong, which also means the gate is silent about whatever has not happened yet.
FAQ
How fast does an automated quality gate need to be?
Ours runs in about 1.5 seconds end to end, including interpreter startup, and the checking itself takes roughly 170 milliseconds. That cost is the whole reason it runs on every edit instead of once before shipping.
Which rules should block and which should only warn?
Block only defects that live inside the file itself, such as a missing section or a sentence over the length limit. Anything that depends on what shipped yesterday should warn, because blocking on it can empty your queue overnight.
What can a text-based gate not catch?
It reads text and structure, so it verifies that a number is present and never that the number is correct. Screenshots, rendered pages, and factual accuracy all need a different kind of check.
How do you introduce a new rule without breaking a running pipeline?
Run it as a report across everything you already have. When we tested one new rule that way, 174 of 189 queued drafts failed it, so we fixed the drafts first and turned the rule on second.
Does a quality gate go stale?
Yes, in the sense that the rule set only knows about failures you have already seen. Every check in ours exists because something specific went out wrong, which also means the gate is silent about whatever has not happened yet.
Some links may be affiliate links. We may earn a commission at no extra cost to you.
This article was written with AI assistance. It is researched and fact-checked, not based on personal hands-on testing unless explicitly stated.
Comments
Post a Comment