Check this page with an assistantOpens a chat asking it to summarise this article and name the evidence behind each claim.
Claude opens with the prompt on your clipboard: Anthropic does not support prefilled prompts on the web, and we would rather copy it than ship a button that drops it.
The Failure Mode Is Confidence, Not Quality
An AI draft usually reads well. That is the problem. The sentence containing a price that changed eighteen months ago is written with the same fluency as the sentence containing something true, so there is nothing in the prose telling an editor where to concentrate. Fluency has been decoupled from accuracy.
This inverts normal editing. With a human draft, uncertainty leaves traces: hedged phrasing, a thinner paragraph, a missing citation where the writer could not find one. Those traces are how editors have always allocated attention. A model produces no such traces, because it is not experiencing uncertainty in a way that reaches the output.
The practical consequence is that editing an AI draft cannot be reading-led. You cannot skim for the weak parts, because there are no weak parts to spot. It has to be claim-led: enumerate every factual assertion and check each one, regardless of how confident the surrounding sentence sounds.
The Model Cannot Check Itself
Asking a model whether its own draft is accurate frequently returns reassurance, because the same process that produced the claim is being asked to evaluate it. Verification has to reach outside to a source you open, or the loop never closes.
This is the single most common failure in AI content workflows, and it is attractive precisely because it feels like a control. A review step exists, a checklist gets ticked, output is produced saying everything checks out, and nobody opened a source. The workflow has the shape of verification with none of the substance.
Using a second model as the checker is better and still not sufficient. It removes the self-consistency problem and leaves the shared-training problem: two models can be confidently wrong about the same widely repeated inaccuracy. Automated checks are useful for catching internal contradictions and missing citations. They are not a substitute for opening the page you are citing.
The Review Order
Work outside in: structure first, then claims, then voice, then the things only a human knows. Doing it in the reverse order, which is the natural instinct, means polishing sentences that are about to be deleted.
- Cut what should not exist.
Models pad. Sections restating the introduction, paragraphs explaining why the topic matters, and conclusions summarising what you just read can usually go entirely. Do this before any other work, because it removes a third of what you would otherwise verify.
- Enumerate every factual claim.
Numbers, dates, prices, names, quotes, and any statement about what a company or standard says. Put them in a list. This is mechanical and it is the step that determines whether the piece is publishable.
- Verify each against a primary source.
Open the source. Confirm it says what the draft claims, note the URL and the date you checked. Claims that cannot be traced get cut, not softened, because a hedged version of an unverifiable claim is still unverifiable.
- Add what only you know.
The specific case, the thing that surprised you, the mistake you watched a client make. A model cannot produce this and it is the only part of the piece a competitor cannot generate in thirty seconds.
- Fix the tells.
Hedged non-commitment, tricolon lists, section openings that restate the heading, and the habit of describing something as multifaceted rather than saying what it is. These read as machine output to anyone who reads a lot, which increasingly includes your audience.
- Name someone accountable.
A real person who reviewed it and would answer for it. That is the reason bylines exist, and it is why an AI byline is not a disclosure, as covered in author entities and E-E-A-T.
Ignore the Percentage Benchmarks
You will see claims that the industry benchmark for human revision is some specific band, with anything below it indicating insufficient oversight. These numbers have no published methodology and no plausible way of having been measured. Treating them as a target produces rewriting for its own sake.
Think about what the metric would even mean. A draft that is factually correct and needs light editing has been well-produced; a draft requiring a fifty percent rewrite was a poor draft. Under a percentage target, the first gets flagged as insufficiently reviewed and the second looks like diligence. The measure rewards bad inputs.
The honest quality gate is binary and per-claim: is every factual assertion verified against something you opened, and is every sentence one you would defend if challenged? A piece passing that test is publishable whether you changed five percent or eighty. A piece failing it is not publishable at any revision ratio.
Where the Policy Line Sits
Google's position is that it rewards quality regardless of how content is produced, while scaled content abuse, generating pages at volume primarily to manipulate rankings rather than to help readers, is a policy violation. The line is purpose and quality, not the tool.
That framing is more useful than it first appears, because it tells you what actually creates risk. It is not using a model to draft. It is publishing at a volume that could not possibly have been reviewed, on topics chosen because they exist rather than because anyone asked, with no accountable human anywhere in the process. Those are production decisions, and they are visible in the output.
It also means the editing process described here is the risk control. A piece that has been fact-checked against primary sources, carries something only you could contribute, and has a named human behind it is not the thing the policy describes, whatever drafted the first version. The wider policy question is worked through in does Google penalise AI content, and the pipeline version in automating blog publishing.
Questions People Ask About Editing AI Content
- What is wrong with AI drafts, specifically?
Not the prose, which is usually competent. The problem is confident wrong facts: outdated prices, invented statistics, misattributed quotes, and sources that do not say what the draft claims. A model states these in exactly the same tone as things it has right, so there is no signal in the text telling you where to look.
- Can I ask the AI to fact-check its own draft?
Not as your only check. Asking a model to verify its own output frequently produces confident confirmation, because the same process that generated the claim is being asked whether the claim is true. Verification has to reach outside the model to a primary source you open yourself, or it is not verification.
- How much of an AI draft should I rewrite?
As much as the draft is wrong, which varies. Benchmarks circulating that specify a percentage of a draft requiring revision are invented figures with no methodology behind them, and treating them as a target produces rewriting for its own sake. The useful test is whether every claim is verified and every sentence is one you would defend, not whether you have hit a quota.
- Does Google penalise AI-written content?
Not for being AI-written. Google's position is that it rewards high-quality content regardless of how it is produced, while scaled content abuse, mass-producing pages to manipulate rankings without serving readers, is a policy violation. The distinction is purpose and quality, not authorship, and it is the reason editing matters more than concealment.
- Should I disclose that AI was involved?
Be transparent in a way that helps the reader rather than performing a ritual. Google has said that giving AI an author byline is not the right way to disclose involvement, since a byline is an accountability claim. A named human who reviewed and stands behind the work, plus a plain note on how content is produced, is the version that holds up.

