What Breaks When an AI Agent Publishes Unattended

What Breaks When an AI Agent Publishes Unattended

Unattended Failures Return Zero and Keep Going

The short answer from two months of unattended operation is that nothing crashed. Four separate failures each returned success, wrote something plausible, and kept the schedule. One of them cleared the labels on 210 of 502 live posts before anyone noticed, and another would have held 174 of 189 queued drafts the moment we enabled a new rule.

Some scale, so you know what this is drawn from. The pipeline publishes to nine blogs, one post per blog per day, with nobody watching. It runs on 44 Python scripts totaling 5,481 lines, and 528 posts have gone out through it.

If you are handing an AI agent a schedule and walking away, exit codes will not tell you what happened. State will.

A Status Value That Only Half the Pipeline Knew

One Writer, One Reader, Two Vocabularies
  • ● A fourth status value was added for retired posts
  • ● The candidate picker never learned it
  • ● The gate then overwrote retired with held

Each post carries a status in its frontmatter: planned, published, or held. We added a fourth value, retired, for posts we consolidated and redirected away.

We taught the writer about it. We did not teach the reader. The candidate picker skipped published and held, so a retired post looked exactly like an unpublished one and came back into the daily rotation.

The second half is worse than the first. The gate rejected those posts, correctly, and wrote held into the status field. That overwrote retired, so the marker recording the decision was gone by the next morning.

Four posts on one blog went through that loop before we caught it. The fix was three characters wide:

SKIP_STATUS = {"published", "held", "retired"}

The lesson is not about a set literal. A status value is a contract between everything that writes it and everything that reads it. An agent adding a new value updates the writer in front of it, not the four readers it never opened.

The Update Call That Cleared Fields We Never Sent

A Full Replacement Pretending to Be a Merge
  • ● Absent fields are cleared on the live post
  • ● 210 of 502 posts ended up with no labels
  • ● The repair belongs in the shared helper

Blogger exposes an update method for changing a published post. We used it whenever we added images, repaired an internal link, or stripped a dead one.

That method is a full replacement, not a merge. Every field absent from the request body is cleared on the live post, and labels were absent from ours.

By late August, 210 of 502 live posts had no labels at all. Nothing failed. Four separate scripts had all been quietly deleting the same field for weeks, because they all called the same helper.

The repair lives in the helper rather than in each caller:

if labels is None:
    cur = service.posts().get(blogId=blog_id, postId=post_id).execute()
    labels = cur.get("labels")

Read the live values and send them back inside the same call. For the cleanup of already-damaged posts we used the patch method instead, which updates only the fields you name and leaves the body untouched.

This generalizes past one API. Any method named update deserves one question before an agent is allowed to call it in a loop: is this a merge or a replacement?

The Background Worker That Left No Trace

Long jobs are the obvious candidate for a background worker, so we tried it. Queue the work, let it run, collect the output later.

A session restart removes that worker and everything it held in memory. There is no error, no partial output, and no entry in any log, because the process that would have written the log is the process that disappeared.

We changed the shape of the work rather than trying to make the worker survive. Each step now writes one file and finishes, so an interrupted run costs a few seconds of repetition instead of an hour of invisible loss.

That restructuring paid twice. It also made every run resumable, which matters more than speed when the thing running is not being watched.

Why We Measured the Queue Before Turning On a Rule

Run the Rule as a Report First
  • ● 174 of 189 drafts would have been held
  • ● Nine blogs would have stopped the same night
  • ● Every individual rejection would have been correct

We wanted a new blocking rule: the first 120 words of every post must carry an actual verdict and a real number, not throat-clearing. It is a good rule and we still believe in it.

Before enabling it, we ran it across the queue as a report rather than a gate. Of 189 planned drafts, 174 would have failed, and every one of those becomes held, which means removed from the publishing rotation.

Turning that on unmeasured would have stopped nine blogs the same night, quietly, with each individual decision defensible. The order we used instead was to fix the drafts first and enable the rule second.

The queue passes it now. A fresh count across 199 planned drafts shows 199 carrying a verdict inside the opening, which is the state we wanted, reached in the order that did not break anything.

The Rejection Pile Nobody Was Reading

A gate that blocks a post has to put it somewhere. Ours moves it to held, which is honest and also a place things go to be forgotten.

Twelve drafts had accumulated there. Some were genuinely finished work that failed on a rule we later relaxed, and some were duplicates that deserved to stay blocked, but nothing distinguished the two piles from the outside.

Working through them took an afternoon and split cleanly. Seven went back into the queue after small fixes, and five were retired for good, which is a healthy ratio and an unhealthy delay.

The gap was never a missing check. It was a missing habit, because held has no schedule attached to it and planned does. Any state your automation can move things into needs a standing reason for someone to look at it, or it becomes storage.

We now count held every morning next to the other statuses. Two sit there today, both recent, which is what a working rejection pile is supposed to look like.

What Each Failure Cost and What Caught It

Failure Silent for Blast radius What surfaced it Fix location
Retired status not read by the picker About a day 4 posts, retirement marker erased Status counts stopped matching One constant in the picker
Update call clearing labels Several weeks 210 of 502 live posts Label audit across the network The shared API helper
Background worker vanishing Immediate, repeatedly Whole jobs, no output Expected files never appeared Job shape, not the worker
Strict rule enabled unmeasured Would have been one night 174 of 189 drafts held Dry run before enabling Enable order
Encoding failure inside a log line Minutes to hours Everything after the print Loop stopped partway Shared module, one line

The middle column is the one worth reading twice. Time-to-notice tracks how far the damage sits from anything a person looks at, not how serious the bug is.

Which Safeguard Fits Your Setup

If you run one repository and read every diff: you need almost none of this. Review catches contract changes such as a new status value, which is exactly the failure mode that survives longest without review.

If an agent runs on a schedule and you check in daily: count rows by state every morning and compare against yesterday. A cheap counter that moves when nothing shipped beats an elaborate alert you never tune.

If an agent writes to a live external service: audit one field you never intentionally change, such as labels or tags. Fields nobody edits are the ones that disappear without complaint.

If you are adding a rule to an automated queue: run it as a report first and read the count. Any rule that removes items from a queue can empty the queue, and defensible individual decisions add up to a stopped pipeline.

If the work takes longer than a session: make each step persist and finish. Resumability is worth more than throughput once nobody is watching the terminal.

The Pattern Underneath All Four

None of these were model failures. The agent wrote correct code against the file in front of it every single time, and the damage came from everything that file touched.

That is the honest summary of unattended AI work as we have experienced it. The generation step is not the fragile part. The fragile part is the set of assumptions nobody restated when one of them changed.

We now write down two things before an agent touches a pipeline: which fields this call replaces, and which readers depend on this value. Both questions are boring, and both would have caught more than half of what is on this page.

The layer underneath this one has its own set of quiet failures, and we wrote those up separately as six Windows traps from the same pipeline. If you automate against this particular platform, ten minutes on the reference pages also pays for itself. Read posts.update and posts.patch side by side, and note which one says the request body replaces the resource.

Frequently Asked Questions

How do you notice a failure that returns success? Watch state, not exit codes. Every failure here returned zero and wrote something plausible. The cheap version is a daily count of rows per status, because a number that moves when nothing shipped is the earliest honest signal you get.

Is adding a new status value to a pipeline safe? Only if every reader of that field learns the new value at the same time. We added a retired status to the writer and forgot the candidate picker, so retired posts came back as publish candidates and had their marker overwritten within a day.

Does the Blogger update call merge the fields you send? No. Blogger posts.update is a full replacement, so any field you leave out gets cleared. Labels vanished on 210 of 502 live posts before we found it, and the fix was to read the live values and send them back inside the same call.

What happens to a background agent when the session restarts? Nothing survives, and nothing is logged. A session restart removes the worker along with whatever it held in memory, so unattended work belongs in steps that persist after each file rather than in a long-lived background job.

What is the risk of adding a strict rule to an automated pipeline? Measure the queue first. Turning on one new blocking rule would have held 174 of 189 queued drafts, which stops publishing entirely, so we fixed the drafts and enabled the rule afterward.

FAQ

How do you notice a failure that returns success?

Watch state, not exit codes. Every failure here returned zero and wrote something plausible. The cheap version is a daily count of rows per status, because a number that moves when nothing shipped is the earliest honest signal you get.

Is adding a new status value to a pipeline safe?

Only if every reader of that field learns the new value at the same time. We added a retired status to the writer and forgot the candidate picker, so retired posts came back as publish candidates and had their marker overwritten within a day.

Does the Blogger update call merge the fields you send?

No. Blogger posts.update is a full replacement, so any field you leave out gets cleared. Labels vanished on 210 of 502 live posts before we found it, and the fix was to read the live values and send them back inside the same call.

What happens to a background agent when the session restarts?

Nothing survives, and nothing is logged. A session restart removes the worker along with whatever it held in memory, so unattended work belongs in steps that persist after each file rather than in a long-lived background job.

What is the risk of adding a strict rule to an automated pipeline?

Measure the queue first. Turning on one new blocking rule would have held 174 of 189 queued drafts, which stops publishing entirely, so we fixed the drafts and enabled the rule afterward.


Some links may be affiliate links. We may earn a commission at no extra cost to you.

This article was written with AI assistance. It is researched and fact-checked, not based on personal hands-on testing unless explicitly stated.

Comments