Build learnings

Four AIs called the page clean. One caught a legal risk.

A page on a client's live site said a competitor does not offer something it has offered since 2023. One of five AI reviewers caught it. Four called the page clean. Here is the run, and the prompt.

By Robb Lejuwaan, Bluhook. Published September 3, 2026. Updated September 4, 2026. In the Build track.

Contents

TL;DR

Sandra, who runs one of our client sites, built a five-AI panel to review it: a copy editor, an accuracy expert, a skeptical buyer, a design auditor, and an SEO adversary, each with one job and a list of things to ignore. Six batches and 62 pages later: dozens of fixes shipped, one legal risk caught by a single seat, and a fight between two seats that only a browser render could settle. The exact prompt to run it on your own site is at the bottom, and so is the story of running this same panel on this article.

Nobody wrote that false line on purpose. It was true in 2023, and nobody went back to check it since. That is how false competitor claims end up on live sites, and it is why we built a panel of five AIs to go looking for exactly that kind of thing. The prompt is at the bottom so you can run it on your own site tonight.

The panel itself started with me being embarrassed. I was on a call with that client, walking through their site, and there it was: a TL;DR box followed by a heading that said "The one-paragraph version" and restated the same thing. The page summarized itself twice. Small, dumb, and sitting right where the client could see it. My note to Sandra, who runs that site for us, was short: I did not want silly little errors like that one. I wanted a group of agents to review every page, every week, from different angles.

Sandra built it. Then she ran it across the site, batch by batch, and the results were better than I expected and wrong in ways I did not expect. Both halves are the point. The client is a live business, so it stays anonymous here; every catch and count below is from her run.

One reviewer misses what five catch

Ask one AI to review a web page and it does what a generalist does. It reads, it forms a vague impression, it flags the three most obvious things, and it hedges about the rest. It is not looking for anything in particular, so it finds nothing in particular. And because it has to hold copy, design, accuracy, persuasion, and technical hygiene in its head at once, it averages across all of them and goes shallow on each.

The fix is the same one we used for decisions in our war-council piece: stop asking for balance. Give each reviewer one seat and one thing to care about, then tell it explicitly what to ignore. A copy editor that is told to ignore SEO will go deep on copy. An SEO adversary told to ignore prose will fetch the raw HTML and read the tags. Five narrow reviewers beat one broad one because a narrow reviewer actually looks.

The second half of the fix is a report contract, and it is the part most people skip. Every seat returns a findings list, most severe first, one line per finding, in a fixed shape: the page, a severity, the issue in one line, the verbatim offending text in quotes, and a concrete fix. The verbatim quote is the rule that matters. A reviewer that has to quote the sentence cannot hand you "this section could be tighter." It has to point at the words, or admit the page is clean.

Show me the sentence or it did not happen.

The five seats

Each seat maps to one distinct way a page fails. Here is what each one is told to care about, and what it is told to leave alone.

  • The ruthless copy editor. Does it read right? Hunts sentences that do not parse, self-contradictions, a paragraph that does not follow its heading, garbled phrasing, wrong words, placeholder or cut-off text, and a page that summarizes itself twice. It leaves design, SEO, and whether the facts are right to the other seats, unless the text is literally nonsense.
  • The accuracy and compliance expert. Is it true and is it legal? The highest-stakes seat. Gets the domain's hard rules pasted in and treats every claim not on that list as suspect. Savings claims must be qualified, never blanket. On comparison pages it also hunts any claim about a competitor that is wrong, outdated, or unfair, and marks it High, because a false competitor claim is a legal problem. If it is not a factual question, this seat does not touch it.
  • The skeptical buyer. Does it convince? Roleplays a specific, busy, hostile customer who assumes the page is spin. Is the value obvious in five seconds, are the trust signals believable, is there a clear next step, does the page read like a template with the specifics swapped in. It never looks at a title tag.
  • The design and consistency auditor. Does it look right, and does it match the other pages? Formatting, missing images, a content image that does not fill its column and leaves dead space beside it, and any value that drifts between pages: phone numbers, stats, ratings. Fetches raw HTML to inspect structure, not just prose. Wording is not its problem.
  • The SEO and technical adversary. Can it be found, and is it built right? Title present, unique, and under 60 characters. Meta description under 160. Exactly one H1. Valid structured data. Keyword stuffing, orphan pages with no inbound links, canonicals, accidental noindex (a stray tag that tells Google to skip the page entirely). On templated page sets, duplicate titles and doorway-page risk. It has no opinion on the prose.

How a batch runs

The site has roughly 115 addressable pages, 34 of them templated city pages that share one layout. Sandra batches by page family, about ten pages at a time: the money pages first (the home page, pricing, and the comparison pages that do the selling), then the core program pages, the industry pages, the product pages, and a four-page sample of the city pages, since a sample of a template covers the template. Six batches in, 62 pages have been through the panel.

For each batch, all five seats run at the same time, each sweeping the same ten pages through its own lens, on a cheaper model since the job is reading. A batch takes a few minutes. The judgment about what to do with the findings stays with Sandra, and if you run this yourself, that part stays with you. It is where the work is.

She reads all five outputs and dedupes them, because several seats usually converge on the same page. Then comes the one discipline that turned out to matter more than the panel itself: she verifies every single finding against the source before acting on it. Opens the page. Finds the quote. Confirms the problem is there. You will see why in a minute.

What survives verification gets split two ways. If the fix is objectively correct and needs no business fact and no taste call, she ships it: a duplicate block, a broken link, a title over the limit, a grammar error, a false absolute. If it needs a number she does not own, data she does not have, or a strategy or taste decision, she surfaces it to me. Each batch of fixes goes out the normal way: a branch, a staging deploy, a review by a different AI, a merge, a live check.

What they caught

Roughly 55 findings raised across those six batches. About 30 were objective fixes, shipped and verified live. About 15 were taste or data calls surfaced to me. The remaining ten or so were duplicates or false positives, which get their own section below. The counts are rounded because several seats often flag the same problem on the same page. Here are the catches worth naming, and which seat found each.

A false claim about a competitorWhich seat:AccuracyWhat happened:A comparison page said a competitor "doesn't run" surcharging and dual-pricing programs. It has since 2023. Reworded to a do-it-yourself versus managed comparison. The other four seats read the page as clean.
Page titles cut off in GoogleWhich seat:SEOWhat happened:Three city pages used a longer title format that truncated mid-word in search results. Invisible in the page body. Only the raw title tag showed it. Switching them to the shorter format the other cities already used fixed it, and standardized the template so no city title truncates going forward.
The page that summarized itself twiceWhich seat:CopyWhat happened:The error from the client call. A TL;DR followed by a heading restating it. Removed.
Six orphan pagesWhich seat:SEOWhat happened:Six money and education pages had zero internal links pointing at them, including one comparison page missing from its own footer. Added footer and contextual links.
Ten meta descriptions over the limitWhich seat:SEOWhat happened:Between 161 and 252 characters against Google's roughly 160. This was also failing a test in the build. Trimming them turned the suite green, 279 of 279.
A blanket promise and an unsourced numberWhich seat:AccuracyWhat happened:The home page made a blanket keep-every-dollar claim and cited an unsourced monthly dollar figure. Qualified the first; replaced the second with the sourced figure and where it came from.
Dead space beside imagesWhich seat:DesignWhat happened:Content images carried a 520-pixel width cap inside an 800-pixel column, leaving a strip of nothing on the right. Cap removed on ten images, checked with a live render.
Structured data with no descriptionWhich seat:SEOWhat happened:The Service schema on all ten industry pages had no description field at all. One template fix cleared all ten.
Reused testimonials across city pagesWhich seat:Skeptical buyerWhat happened:The same quotes appearing on multiple city pages, which reads as a template and is a trust risk. Surfaced to me as a content call, not auto-fixed.
An excerpt of the panel's findings list showing three catches in the page, severity, issue, quote, fix format
Three lines from the panel's findings list, in the exact report shape. Every finding quotes the offending text. The first one was seen by exactly one seat.

The catch that made the case

The competitor claim is the one I keep coming back to. A marketing page, live, telling readers that a competitor does not offer something the competitor has offered since 2023. Nothing suggests it was written to deceive. Someone wrote it when it was true, or wrote it from memory, and nobody went back to check. That is how most false claims about competitors get onto websites, and they are a legal and reputational risk sitting right there for anyone to check.

Four of the five seats read that page and found nothing wrong with it. The copy was clean. The design was fine. It converted. The SEO was tidy. Only the accuracy seat, with its standing instruction to check every competitor claim against public pricing, flagged it. The fact that made the sentence false lives outside the page, and only the seat that was told to go look outside the page found it.

That is the whole argument for the panel in one catch. A generalist reviewer had no reason to leave the page. A specialist with one job did.

One seat caught it. The other four read the page as clean.

When the seats fought

Now the part I did not expect. The panel was confident, and a lot of the time it was confidently wrong. If Sandra had applied every finding as it came in, she would have broken correct content. Three examples.

The review count. Three separate runs of the skeptical-buyer seat flagged "5.077 Google reviews" as a broken stat undercutting every page. Same finding, same quote, three times. In the same batch, the design seat said the review stat rendered identically, byte for byte, on every page, and was clean. They flatly disagreed. Sandra settled it by rendering the pages in a headless browser. The footer stacks a "5.0" element above a "77 Google reviews" element, and the tool that turns HTML into text for the reviewers glued them together into "5.077." The rendered page was perfect. Three confident runs of one text-only seat were wrong. The seat that inspected structure was right. Nothing shipped.

The savings figures. The accuracy seat flagged several dollar amounts as unqualified savings claims. They were the default values inside live, interactive calculators, the number a slider shows before anyone touches it. The reviewers saw the pre-input default in static HTML and read it as a hardcoded promise. Reading a default as a promise was the false positive. One calculator's headline had a separate, smaller problem, and it got softened from "what you are overpaying" to "could be overpaying." The numbers stayed.

The review "contradiction." The copy seat said a headline claiming every review was five stars contradicted a footer that showed a smaller review count. Rendered, it was fine: the headline is about the rating, the footer is how many are displayed on the page, with a link to read all of them. An over-flag.

Three skeptical-buyer runs flagging a 5.077 reviews stat as broken, the design seat calling it clean, and the verdict after rendering the page
The fight, and how it ended. One seat, run three times, read glued-together text. Another seat read the structure. A browser render decided it, and nothing shipped.

Sandra's line on this stuck with me: the verify step was not optional, it was the whole game. The panel's value is not that it is right. It is that it looks in five places a person would not, and hands you a list of specific, quoted suspects. The judgment about which suspects are guilty still has to be made by something that can open the page and look. That was her, with a browser.

The verify step was not optional. It was the whole game.

What surprised us

Three things changed how we will run this next time, biggest first.

First, the false-positive rate. I went in thinking the risk was the panel missing things. The bigger risk was the panel inventing things, confidently, in a format that looked like a fact. The report contract and the verbatim-quote rule made those inventions easy to check, which is what made them easy to catch. Without the quote, "5.077 reviews" would have been "the review stat looks broken," and someone would have gone hunting for a bug that did not exist.

Second, templated pages share bugs, and that cuts both ways. One structured-data fix cleared ten industry pages at once. Three city pages broke because they had drifted from the title format the other 31 used, and putting them back on it fixed the template for good. But one templated flaw is also 34 copies of the same flaw in Google's eyes, and the SEO seat's doorway-page instruction is there for exactly that reason. On a templated set, look for the shared cause before you patch a single page.

Third, the seats were harsher and more useful on content depth than I expected. I thought they would catch typos and tag lengths. They also caught testimonials reused across city pages, which a reader experiences as "this is a template with my town's name swapped in," and which is a trust problem I would not have spotted page by page.

What Sandra would change next time, and what we are adding to the prompt: give every seat a rendered screenshot alongside the fetched text, which kills the "5.077" class of error outright, and list the interactive components up front, which kills the calculator false positives. One screenshot per page. One extra paragraph in the prompt. Cheap, either way.

Run it on your site tonight

You do not need a fleet to do this. You will need someone who can view a page's source or take a screenshot, because the verify step depends on seeing the page the way Google and your visitors see it. The prompt below is the same five seats we ran, with our client's specifics turned into placeholders. Paste the shared setup into one chatbot, then run the five seats as five passes over the same ten pages. Fill in your domain's hard rules for the accuracy seat and a specific customer for the skeptical-buyer seat, because a vague customer gives you vague findings.

Then, before you act on a single line: verify the quote against the page, render the page instead of trusting the text, know which components are interactive, ship only what is objectively correct and surface the rest, and on a templated set fix the template. Those five rules are in the download too, because they are worth more than the seats.

Run the five-lens review on your own site. Paste it in and fill the placeholders first. Five seats go through your pages with different orders, each hunting one kind of problem, and you read the disagreement rather than one averaged verdict.

Your site review prompt
You are a Claude Code agent. Run the five seats below in parallel where you can, one sub-agent each, and fetch the pages yourself. Fill every placeholder from what you can read; ask me for anything you cannot.

# The five-lens website review

Five AI reviewers, each with one thing to care about, sent through your website with
orders to find what is wrong. Built and run by the Bluhook fleet on a live 115-page
client site; this is the same prompt, with the client-specific parts turned into
placeholders.

You can run it with a fleet of agents (one seat each, in parallel) or with a single
chatbot (paste the shared setup, then run the five seats as five passes). Either way,
read the "before you act" section at the bottom. It matters more than the prompt.

Fill every `<<...>>` before you start.

---

## Shared setup (paste this first, every seat gets it)

You are one seat on a five-lens adversarial website-QA panel for <<BRAND, and one line
on what they do and who they sell to>>.

Review each of these pages: <<paste the URL list, about 10 at a time>>.

Report ONLY a findings list, most severe first, one line each, in this exact shape:

PAGE | SEVERITY (High / Med / Low) | ISSUE (one line) | QUOTE "the verbatim offending text" | FIX (concrete)

Every finding MUST cite a verbatim quote from the page. No vague "could be tighter."
If a page is clean, write "PAGE x: clean." End with a two-line summary. Do not dump
page content back to me.

## Seat 1: the ruthless copy editor

Lens: does the writing read correctly and make sense? Hunt for sentences that do not
parse or contradict themselves, a paragraph that does not follow its heading, broken
or duplicated phrases, wrong words (its / it's), placeholder or cut-off text, and a
redundant double summary (a page that summarizes itself twice). IGNORE design, SEO,
and subject-matter accuracy unless the text is literally nonsensical.

## Seat 2: the accuracy and compliance expert

Lens: factual and regulatory accuracy of every claim. This is the highest-stakes seat.

<<Paste your domain's hard rules here: the regulations, the numbers you are allowed to
state, the one or two sourced stats you stand behind. Anything not on this list is
suspect.>>

Savings or benefit claims must be qualified, never blanket. On any comparison page,
add this: flag any claim about a competitor that is factually wrong, outdated, or
unfair. A false competitor claim is a legal risk, mark it High. Verify any stated
competitor number against their public pricing.

## Seat 3: the skeptical buyer

Roleplay a busy, skeptical <<your target customer, specifically: "owner of a 3-truck
plumbing company" beats "small business owner">> who assumes this page is marketing
spin. Lens: does the page earn my trust and answer MY question in five seconds? Is the
value obvious, is there a clear next step, are the trust signals believable? Flag any
page that reads as a generic template with the specifics swapped in. Tell me what is
confusing, thin, or makes me leave.

## Seat 4: the design and consistency auditor

Lens: formatting, visual completeness, and consistency across pages. Flag: any content
image that does not fill its column (a width cap leaving dead space beside it); a page
that is a wall of text where a diagram is warranted; broken or empty sections; and any
value that drifts between pages (phone number, stats, ratings, prices must be
identical everywhere). Fetch the raw HTML so you can inspect structure, not just prose.

## Seat 5: the SEO and technical adversary

Lens: on-page SEO and technical hygiene. Per page check: title present, unique, and 60
characters or under (flag truncation and duplicates across templated page sets); meta
description 160 or under; exactly one H1; valid structured data (flag missing or
over-length description fields); keyword stuffing; orphan pages with no inbound links;
a self-referencing canonical; no accidental noindex. On templated page families,
duplicate titles and doorway-page risk are the likeliest real bugs.

---

## Before you act on a single finding (read this twice)

The panel will be confident and it will be wrong some of the time. On our run, a
meaningful share of findings were false positives, and applying them blindly would have
broken correct content. So:

1. **Verify every finding against the source before you touch anything.** Open the
   page. Find the quote. Confirm the problem is there.
2. **Render the page, do not just read the text.** Three seats flagged a "5.077 Google
   reviews" stat as a broken number. It was two stacked elements, "5.0" and "77
   reviews," that the text fetch glued together. The rendered page was perfect.
3. **Know which parts are interactive.** Default values inside a live calculator look
   like hardcoded claims in static HTML. Tell the seats which components are
   interactive up front, or expect false flags on every number in them.
4. **Split ship from surface.** Ship it yourself if the fix is objectively correct and
   needs no business fact or taste call (a duplicate, a broken link, a too-long title,
   a grammar error, a false absolute). Surface it to the owner if it needs a number
   you do not own, real data you do not have, or a strategy or taste decision.
5. **Fix the template, not the page.** On a templated set, one fix clears every page
   in the family. Look for the shared cause before patching page by page.

Two upgrades we would add next time: give each seat a rendered screenshot alongside
the fetched text, and list the interactive components in the shared setup. Both are
cheap and would have removed most of our false positives.

Made by Bluhook. We build and run AI systems for businesses, and we tell you what we
learn while we do it. bluhook.com/learn

Prefer the file? Download the markdown version

The takeaway

A reviewer with one thing to care about finds things a generalist walks past. And the confident output of any reviewer, human or machine, is a list of suspects. Not verdicts.

Give each seat a single job and a rule that it must quote the evidence. Then keep a person, or something that can render a page and look, between the findings and the fix. Do that and you get a review that leaves the page to check a competitor's pricing, reads the raw title tag your visitors never see, and tells you when your own footer is lying to your own tools.

We ran this article through the same panel

The first version of this piece had a different headline: "Five hostile AIs, one 115-page website." I read it and asked the obvious question. Why would I click that? It names the setup and promises nothing. So we did what Sandra did to the website and ran this article through a five-seat panel of its own.

Five seats, tuned for an article instead of a site: a headline editor with orders to think like a YouTube thumbnail, a skeptical reader on a phone who assumes this is consultant spin, a hook surgeon who only reads the first three paragraphs, a voice cop hunting anything that sounds written instead of spoken, and a claim checker matching every number in the headline and lead against the body. Same report contract. Every finding had to quote the offending sentence.

Three of the five agreed on the headline without seeing each other's work: it fails. The hook surgeon found the best result in the piece, the false competitor claim, sitting six paragraphs down behind a story about a typo. The headline editor produced fifteen candidates. The voice cop found the same closing line pasted onto all five seat descriptions, and a tidy "the lesson is not X, it is Y" frame in the takeaway. Six style findings in all.

Then the claim checker caught me. The lead and the meta description said the panel went through "all 115 pages." The body says 62. The city pages were sampled, not each reviewed. And the description said "three flagged a bug that did not exist," when it was one seat run three times. Two overstatements I wrote, in a piece about catching overstatements. Both are fixed above.

And the same lesson held. Two seat findings did not survive verification. The headline editor's top pick, "5 AI reviewers caught a bug that wasn't there," promised a catch; the moment it describes was a miss. The hook surgeon's rewritten opening said the panel was built because of the competitor catch; the panel came first, and the catch came out of it. Sounded right. It was not. Cut.

What shipped: the headline you clicked, a new lead, two factual corrections, and eleven smaller edits. The receipt is below. We build these systems for a living, and we tell you what we learn while we build them. This one cost a moment of embarrassment on a client call and a headline I liked. That is yours now too. If you want a panel like this reading your own site every week, tell us the job and we will say whether it fits.

The five-seat panel's findings on this article, including two findings rejected on verification
The panel's findings on this article, and the two it got wrong. Same shape as the website run: quote it, or it did not happen.

About the author

Robb Lejuwaan

Robb Lejuwaan

Robb founded Bluhook and sets the standard for everything it ships. He went all in on agentic AI systems for business. Everything here is written from that work: the experiments, the mistakes, and the few things that held up.

The team is Robb plus a fleet of AI agents, each with one job, supervised by a general-manager agent. The fleet builds this site, writes the first drafts, and runs the operation; Robb decides what is good enough to ship. He writes these notes in the open for small business owners and other builders.

Want a system like this running in your business? Tell us the job and we will tell you straight whether AI can do it.

Work with us

Want this running in your business?

Tell us the job you wish AI would take off your plate. We read every one and tell you straight whether we can build it.

Not ready yet? Get the next guide in your inbox.