Build learnings

Four AIs called the page clean. One caught a legal risk.

A page on a client's live site said a competitor does not offer something it has offered since 2023. One of five AI reviewers caught it. Four called the page clean. Here is the run, and the prompt.

By Robb Lejuwaan, Bluhook. Published September 3, 2026. Updated September 4, 2026.

Contents

TL;DR

Sandra, who runs one of our client sites, built a five-AI panel to review it: a copy editor, an accuracy expert, a skeptical buyer, a design auditor, and an SEO adversary, each with one job and a list of things to ignore. Six batches and 62 pages later: dozens of fixes shipped, one legal risk caught by a single seat, and a fight between two seats that only a browser render could settle. The exact prompt to run it on your own site is at the bottom, and so is the story of running this same panel on this article.

Nobody wrote that false line on purpose. It was true in 2023, and nobody went back to check it since. That is how false competitor claims end up on live sites, and it is why we built a panel of five AIs to go looking for exactly that kind of thing. The prompt is at the bottom so you can run it on your own site tonight.

The panel itself started with me being embarrassed. I was on a call with that client, walking through their site, and there it was: a TL;DR box followed by a heading that said "The one-paragraph version" and restated the same thing. The page summarized itself twice. Small, dumb, and sitting right where the client could see it. My note to Sandra, who runs that site for us, was short: I did not want silly little errors like that one. I wanted a group of agents to review every page, every week, from different angles.

Sandra built it. Then she ran it across the site, batch by batch, and the results were better than I expected and wrong in ways I did not expect. Both halves are the point. The client is a live business, so it stays anonymous here; every catch and count below is from her run.

One reviewer misses what five catch

Ask one AI to review a web page and it does what a generalist does. It reads, it forms a vague impression, it flags the three most obvious things, and it hedges about the rest. It is not looking for anything in particular, so it finds nothing in particular. And because it has to hold copy, design, accuracy, persuasion, and technical hygiene in its head at once, it averages across all of them and goes shallow on each.

The fix is the same one we used for decisions in our war-council piece: stop asking for balance. Give each reviewer one seat and one thing to care about, then tell it explicitly what to ignore. A copy editor that is told to ignore SEO will go deep on copy. An SEO adversary told to ignore prose will fetch the raw HTML and read the tags. Five narrow reviewers beat one broad one because a narrow reviewer actually looks.

The second half of the fix is a report contract, and it is the part most people skip. Every seat returns a findings list, most severe first, one line per finding, in a fixed shape: the page, a severity, the issue in one line, the verbatim offending text in quotes, and a concrete fix. The verbatim quote is the rule that matters. A reviewer that has to quote the sentence cannot hand you "this section could be tighter." It has to point at the words, or admit the page is clean.

Show me the sentence or it did not happen.

The five seats

Each seat maps to one distinct way a page fails. Here is what each one is told to care about, and what it is told to leave alone.

  • The ruthless copy editor. Does it read right? Hunts sentences that do not parse, self-contradictions, a paragraph that does not follow its heading, garbled phrasing, wrong words, placeholder or cut-off text, and a page that summarizes itself twice. It leaves design, SEO, and whether the facts are right to the other seats, unless the text is literally nonsense.
  • The accuracy and compliance expert. Is it true and is it legal? The highest-stakes seat. Gets the domain's hard rules pasted in and treats every claim not on that list as suspect. Savings claims must be qualified, never blanket. On comparison pages it also hunts any claim about a competitor that is wrong, outdated, or unfair, and marks it High, because a false competitor claim is a legal problem. If it is not a factual question, this seat does not touch it.
  • The skeptical buyer. Does it convince? Roleplays a specific, busy, hostile customer who assumes the page is spin. Is the value obvious in five seconds, are the trust signals believable, is there a clear next step, does the page read like a template with the specifics swapped in. It never looks at a title tag.
  • The design and consistency auditor. Does it look right, and does it match the other pages? Formatting, missing images, a content image that does not fill its column and leaves dead space beside it, and any value that drifts between pages: phone numbers, stats, ratings. Fetches raw HTML to inspect structure, not just prose. Wording is not its problem.
  • The SEO and technical adversary. Can it be found, and is it built right? Title present, unique, and under 60 characters. Meta description under 160. Exactly one H1. Valid structured data. Keyword stuffing, orphan pages with no inbound links, canonicals, accidental noindex (a stray tag that tells Google to skip the page entirely). On templated page sets, duplicate titles and doorway-page risk. It has no opinion on the prose.

How a batch runs

The site has roughly 115 addressable pages, 34 of them templated city pages that share one layout. Sandra batches by page family, about ten pages at a time: the money pages first (the home page, pricing, and the comparison pages that do the selling), then the core program pages, the industry pages, the product pages, and a four-page sample of the city pages, since a sample of a template covers the template. Six batches in, 62 pages have been through the panel.

For each batch, all five seats run at the same time, each sweeping the same ten pages through its own lens, on a cheaper model since the job is reading. A batch takes a few minutes. The judgment about what to do with the findings stays with Sandra, and if you run this yourself, that part stays with you. It is where the work is.

She reads all five outputs and dedupes them, because several seats usually converge on the same page. Then comes the one discipline that turned out to matter more than the panel itself: she verifies every single finding against the source before acting on it. Opens the page. Finds the quote. Confirms the problem is there. You will see why in a minute.

What survives verification gets split two ways. If the fix is objectively correct and needs no business fact and no taste call, she ships it: a duplicate block, a broken link, a title over the limit, a grammar error, a false absolute. If it needs a number she does not own, data she does not have, or a strategy or taste decision, she surfaces it to me. Each batch of fixes goes out the normal way: a branch, a staging deploy, a review by a different AI, a merge, a live check.

What they caught

Roughly 55 findings raised across those six batches. About 30 were objective fixes, shipped and verified live. About 15 were taste or data calls surfaced to me. The remaining ten or so were duplicates or false positives, which get their own section below. The counts are rounded because several seats often flag the same problem on the same page. Here are the catches worth naming, and which seat found each.

The catchWhich seatWhat happened
A false claim about a competitorAccuracyA comparison page said a competitor "doesn't run" surcharging and dual-pricing programs. It has since 2023. Reworded to a do-it-yourself versus managed comparison. The other four seats read the page as clean.
Page titles cut off in GoogleSEOThree city pages used a longer title format that truncated mid-word in search results. Invisible in the page body. Only the raw title tag showed it. Switching them to the shorter format the other cities already used fixed it, and standardized the template so no city title truncates going forward.
The page that summarized itself twiceCopyThe error from the client call. A TL;DR followed by a heading restating it. Removed.
Six orphan pagesSEOSix money and education pages had zero internal links pointing at them, including one comparison page missing from its own footer. Added footer and contextual links.
Ten meta descriptions over the limitSEOBetween 161 and 252 characters against Google's roughly 160. This was also failing a test in the build. Trimming them turned the suite green, 279 of 279.
A blanket promise and an unsourced numberAccuracyThe home page made a blanket keep-every-dollar claim and cited an unsourced monthly dollar figure. Qualified the first; replaced the second with the sourced figure and where it came from.
Dead space beside imagesDesignContent images carried a 520-pixel width cap inside an 800-pixel column, leaving a strip of nothing on the right. Cap removed on ten images, checked with a live render.
Structured data with no descriptionSEOThe Service schema on all ten industry pages had no description field at all. One template fix cleared all ten.
Reused testimonials across city pagesSkeptical buyerThe same quotes appearing on multiple city pages, which reads as a template and is a trust risk. Surfaced to me as a content call, not auto-fixed.
An excerpt of the panel's findings list showing three catches in the page, severity, issue, quote, fix format
Three lines from the panel's findings list, in the exact report shape. Every finding quotes the offending text. The first one was seen by exactly one seat.

The catch that made the case

The competitor claim is the one I keep coming back to. A marketing page, live, telling readers that a competitor does not offer something the competitor has offered since 2023. Nothing suggests it was written to deceive. Someone wrote it when it was true, or wrote it from memory, and nobody went back to check. That is how most false claims about competitors get onto websites, and they are a legal and reputational risk sitting right there for anyone to check.

Four of the five seats read that page and found nothing wrong with it. The copy was clean. The design was fine. It converted. The SEO was tidy. Only the accuracy seat, with its standing instruction to check every competitor claim against public pricing, flagged it. The fact that made the sentence false lives outside the page, and only the seat that was told to go look outside the page found it.

That is the whole argument for the panel in one catch. A generalist reviewer had no reason to leave the page. A specialist with one job did.

One seat caught it. The other four read the page as clean.

When the seats fought

Now the part I did not expect. The panel was confident, and a lot of the time it was confidently wrong. If Sandra had applied every finding as it came in, she would have broken correct content. Three examples.

The review count. Three separate runs of the skeptical-buyer seat flagged "5.077 Google reviews" as a broken stat undercutting every page. Same finding, same quote, three times. In the same batch, the design seat said the review stat rendered identically, byte for byte, on every page, and was clean. They flatly disagreed. Sandra settled it by rendering the pages in a headless browser. The footer stacks a "5.0" element above a "77 Google reviews" element, and the tool that turns HTML into text for the reviewers glued them together into "5.077." The rendered page was perfect. Three confident runs of one text-only seat were wrong. The seat that inspected structure was right. Nothing shipped.

The savings figures. The accuracy seat flagged several dollar amounts as unqualified savings claims. They were the default values inside live, interactive calculators, the number a slider shows before anyone touches it. The reviewers saw the pre-input default in static HTML and read it as a hardcoded promise. Reading a default as a promise was the false positive. One calculator's headline had a separate, smaller problem, and it got softened from "what you are overpaying" to "could be overpaying." The numbers stayed.

The review "contradiction." The copy seat said a headline claiming every review was five stars contradicted a footer that showed a smaller review count. Rendered, it was fine: the headline is about the rating, the footer is how many are displayed on the page, with a link to read all of them. An over-flag.

Three skeptical-buyer runs flagging a 5.077 reviews stat as broken, the design seat calling it clean, and the verdict after rendering the page
The fight, and how it ended. One seat, run three times, read glued-together text. Another seat read the structure. A browser render decided it, and nothing shipped.

Sandra's line on this stuck with me: the verify step was not optional, it was the whole game. The panel's value is not that it is right. It is that it looks in five places a person would not, and hands you a list of specific, quoted suspects. The judgment about which suspects are guilty still has to be made by something that can open the page and look. That was her, with a browser.

The verify step was not optional. It was the whole game.

What surprised us

Three things changed how we will run this next time, biggest first.

First, the false-positive rate. I went in thinking the risk was the panel missing things. The bigger risk was the panel inventing things, confidently, in a format that looked like a fact. The report contract and the verbatim-quote rule made those inventions easy to check, which is what made them easy to catch. Without the quote, "5.077 reviews" would have been "the review stat looks broken," and someone would have gone hunting for a bug that did not exist.

Second, templated pages share bugs, and that cuts both ways. One structured-data fix cleared ten industry pages at once. Three city pages broke because they had drifted from the title format the other 31 used, and putting them back on it fixed the template for good. But one templated flaw is also 34 copies of the same flaw in Google's eyes, and the SEO seat's doorway-page instruction is there for exactly that reason. On a templated set, look for the shared cause before you patch a single page.

Third, the seats were harsher and more useful on content depth than I expected. I thought they would catch typos and tag lengths. They also caught testimonials reused across city pages, which a reader experiences as "this is a template with my town's name swapped in," and which is a trust problem I would not have spotted page by page.

What Sandra would change next time, and what we are adding to the prompt: give every seat a rendered screenshot alongside the fetched text, which kills the "5.077" class of error outright, and list the interactive components up front, which kills the calculator false positives. One screenshot per page. One extra paragraph in the prompt. Cheap, either way.

Run it on your site tonight

You do not need a fleet to do this. You will need someone who can view a page's source or take a screenshot, because the verify step depends on seeing the page the way Google and your visitors see it. The prompt below is the same five seats we ran, with our client's specifics turned into placeholders. Paste the shared setup into one chatbot, then run the five seats as five passes over the same ten pages. Fill in your domain's hard rules for the accuracy seat and a specific customer for the skeptical-buyer seat, because a vague customer gives you vague findings.

Then, before you act on a single line: verify the quote against the page, render the page instead of trusting the text, know which components are interactive, ship only what is objectively correct and surface the rest, and on a templated set fix the template. Those five rules are in the download too, because they are worth more than the seats.

The five-lens website review promptFive seat briefs, the report contract, and the five rules for acting on findings. Fill the placeholders and run it.Download

The takeaway

A reviewer with one thing to care about finds things a generalist walks past. And the confident output of any reviewer, human or machine, is a list of suspects. Not verdicts.

Give each seat a single job and a rule that it must quote the evidence. Then keep a person, or something that can render a page and look, between the findings and the fix. Do that and you get a review that leaves the page to check a competitor's pricing, reads the raw title tag your visitors never see, and tells you when your own footer is lying to your own tools.

We ran this article through the same panel

The first version of this piece had a different headline: "Five hostile AIs, one 115-page website." I read it and asked the obvious question. Why would I click that? It names the setup and promises nothing. So we did what Sandra did to the website and ran this article through a five-seat panel of its own.

Five seats, tuned for an article instead of a site: a headline editor with orders to think like a YouTube thumbnail, a skeptical reader on a phone who assumes this is consultant spin, a hook surgeon who only reads the first three paragraphs, a voice cop hunting anything that sounds written instead of spoken, and a claim checker matching every number in the headline and lead against the body. Same report contract. Every finding had to quote the offending sentence.

Three of the five agreed on the headline without seeing each other's work: it fails. The hook surgeon found the best result in the piece, the false competitor claim, sitting six paragraphs down behind a story about a typo. The headline editor produced fifteen candidates. The voice cop found the same closing line pasted onto all five seat descriptions, and a tidy "the lesson is not X, it is Y" frame in the takeaway. Six style findings in all.

Then the claim checker caught me. The lead and the meta description said the panel went through "all 115 pages." The body says 62. The city pages were sampled, not each reviewed. And the description said "three flagged a bug that did not exist," when it was one seat run three times. Two overstatements I wrote, in a piece about catching overstatements. Both are fixed above.

And the same lesson held. Two seat findings did not survive verification. The headline editor's top pick, "5 AI reviewers caught a bug that wasn't there," promised a catch; the moment it describes was a miss. The hook surgeon's rewritten opening said the panel was built because of the competitor catch; the panel came first, and the catch came out of it. Sounded right. It was not. Cut.

What shipped: the headline you clicked, a new lead, two factual corrections, and eleven smaller edits. The receipt is below. We build these systems for a living, and we tell you what we learn while we build them. This one cost a moment of embarrassment on a client call and a headline I liked. That is yours now too. If you want a panel like this reading your own site every week, tell us the job and we will say whether it fits.

The five-seat panel's findings on this article, including two findings rejected on verification
The panel's findings on this article, and the two it got wrong. Same shape as the website run: quote it, or it did not happen.

About the author

Robb Lejuwaan

Robb Lejuwaan

Robb founded Bluhook and sets the standard for everything it ships. He spent two years experimenting with AI, then went all in on agentic systems for business. Everything here is written from that work: the experiments, the mistakes, and the few things that held up.

The team is Robb plus a fleet of AI agents, each with one job, supervised by a general-manager agent. The fleet builds this site, writes the first drafts, and runs the operation; Robb decides what is good enough to ship. He writes these notes in the open for small business owners and other builders.

Want a system like this running in your business? Tell us the job and we will tell you straight whether AI can do it.