HelpWithWebGet Help Now
← Back to Blog
AI Workflow45 min read

How to Make a 10/10 Website with AI

We used ChatGPT as an independent auditor and Replit Agent as the builder to take a client site from 8.4 to #1 against two named competitors in a controlled AI audit. The exact loop, the real prompt, the scorecards, measured PageSpeed numbers, a downloadable evidence pack, and every trap that cost us a day.

ByDino Bartolome
Newport Beach Skin Tightening homepage hero, rebuilt with AI

Most people use AI to build websites. That's the boring half.

The interesting half is using a second AI to grade the site — coldly, against real named competitors, on a 1–10 scale — and then feeding those findings back to the builder. Do that enough times and something strange happens: the score stops being a vibe and starts being a number you can actually move.

We ran this loop on a client site — newportbeachskintightening.com, a non-surgical skin tightening studio serving Newport Beach.

Two timelines, and it's worth separating them. Building and refining the site was a months-long project. The audit loop described in this article was a three-week sprint at the end of it — several cold audits a day, every day, until the findings dried up. Everything below is about those three weeks.

Going in, the site scored around 8.4. It finished ranked #1 against both benchmark competitors in the audit, with multiple categories at a perfect 10.0, and it stopped not because we declared victory but because the auditor ran out of things to fix.

One qualification up front, because the title oversells it: that #1 is a placing in a repeated head-to-head AI audit against two named competitors. It is not a Google ranking, not a market position, and not a claim that this is the best website in Newport Beach. What it means precisely — and what it doesn't — is the subject of the next section.

Here's exactly how it worked, including the parts that don't show up in the highlight reel.

Everything in this article is live and checkable: newportbeachskintightening.com. Open it alongside the scorecards below and judge the calls yourself — that's rather the point of publishing the rubric.


What's in here

It's a long read. Jump to what you need:


What the auditor concluded

Here's the final scorecard, from a cold audit run on 1 August 2026.

Read it for what it is. These aren't measurements. Nothing was instrumented, no users were tested, no stopwatch was involved. They're numbers a language model produced by reasoning about three websites and then expressing that reasoning as digits. A 9.8 doesn't mean the homepage is 98% of something. It means the model looked, thought, and landed there.

Final ChatGPT audit scorecard
CategoryNB Skin TighteningKevin SadatiForever Ageless
Homepage overall9.89.69.1
Hero / first screen9.99.28.8
Brand distinction10.09.89.0
Treatment clarity9.99.49.1
Before-and-after proof9.89.88.6
Reviews and authority9.710.09.2
Conversion / booking9.99.79.4
Local SEO relevance9.99.89.4
Technical structure9.89.49.2
Luxury presentation9.710.09.6
Overall9.8 — #19.6 — #29.1 — #3

And the segment score, which is the one that actually matters — more on why below:

Non-surgical facelift categoryScore
Newport Beach Skin Tightening9.9 — #1
Kevin Sadati8.8 — #2
Forever Ageless8.5 — #3

Overall, the client wins by 0.2. In its own category, it wins by 1.1. Hold that thought.

One caveat to carry with you from here: this is one rubric, from one session. A fresh cold audit a fortnight earlier scored a partly different set of categories entirely — four of these ten didn't exist in it. We show the side-by-side later, because it changes what these numbers can honestly be used for.

The auditor also disclaims its own scope, which is worth quoting because it's the model being appropriately modest about what a website score measures:

…quality scores—not medical outcome rankings.

So why bother with the numbers at all?

Because a made-up number can still be a useful one, provided you're clear about which part is load-bearing.

What's weak: the absolute value. 9.8 is not a fact about the website. Run the audit again in a fresh session and you might get 9.6, with a partly different rubric — we show exactly that happening later in this article.

What's stronger: the comparison within a single run. All three sites in that table were judged by the same reasoning, against the same criteria, in the same session, minutes apart. Whatever bias the model has, it applied it to all three. So "the client ranked above both competitors on 8 of 10 categories" survives a lot better than "the client scored 9.8."

What's strongest: the direction of movement over many runs. One score is an opinion. Forty scores trending upward, while you make specific changes the auditor asked for, is a signal — not because any individual number is right, but because it's hard to move a biased instrument consistently in one direction by accident.

That's the whole epistemic basis for this method, and it's worth being blunt about: we are not measuring website quality. We are using a model's judgment as a repeatable critic, and treating its consistency — not its accuracy — as the thing we can rely on.

Except where it genuinely is measuring

Not every dimension is judgment. A few are real numbers the auditor fetches rather than invents:

  • Page speed. ChatGPT can run PageSpeed from inside the chat session — you don't leave the audit to go check Lighthouse manually. Those are actual measurements.
  • Technical and indexing. The Search Console figures further down — 95 crawled-not-indexed, 32 not-found — are Google's numbers, not the model's.

So the rubric is a mix of measurement and judgment, and it's worth knowing which is which. "Mobile optimization: 96" is a fact. "Luxury presentation: 9.7" is a read. Both are useful; only one is checkable.

The measured dimensions also do quiet work anchoring the judged ones. If the auditor's page-speed number matches what PageSpeed actually reports, that's small evidence it's looking at your real, current site rather than a cached memory of it — which, as the next section explains at some length, is not a given.

About that title

"10/10" is hook bait. You clicked it, so it worked.

What the auditor actually landed on was 9.8 overall and 9.9 in the category, with individual dimensions at 10.0. And per the section above, even those digits are a model's judgment rather than a measurement — a different session would land somewhere slightly different, on a partly different rubric. There is no certifying body handing out 10s. There's a model with an opinion and a consistent enough one to be useful.

We're saying this on the way in rather than burying it, because the entire method below depends on treating the score as a measuring instrument rather than a trophy. If you'd rather have the honest version of the headline: how to make a website that beats its two best competitors on a cold, adversarial audit, for about a tenth of what that used to cost. Less catchy. Same article.


The two-AI setup

The whole method rests on a separation of duties:

RoleToolJob
The auditorChatGPT with live browsingCold deep-dive audits. Scores 1–10 across ~12 dimensions against named competitors. Never touches the code.
The builderReplit Agent (Power mode) — or Claude CodeTakes audit findings, writes the changes on isolated branches. Never grades its own work.

That last clause is the entire point. The thing that builds the site cannot be the thing that scores the site. An agent that just wrote a hero section will tell you the hero section is excellent — it has no distance from the work. A separate model, given a cold prompt and a list of competitors, has no idea who wrote what, and will happily tell you your homepage is a 7.

Asking the builder to grade its own hero section is like asking the compiler whether your code is beautiful. It compiled, didn't it? Ship it.

What "Power mode" meant

Replit Agent set to Power routed to frontier models — Opus, and Fable during the stretch when Fable was available. (If you were around for it: Fable existed, then didn't, then did again.) Use the best model your builder offers; this is not a place to economize.

To be clear: you don't need Fable. Opus did the overwhelming majority of this build and did it well. Fable happened to be in rotation for part of the project, and we'd struggle to point at a section of the site and tell you which one wrote it. If you're reading this and Fable is unavailable, or you simply prefer Opus, nothing about this method changes. Pick the strongest model your tool offers on the day and get on with it.

It's also a lesson in itself. Models churn — and not just on the builder side. We started these audits on GPT-5.5 and finished them on GPT-5.6 Sol, which shipped on 9 July, part-way through the project. We changed frontier models mid-run without changing a single thing about the method, and the scores kept climbing.

(OpenAI also cut GPT-5.6 pricing on 30 July, two days before our final audit. The audit half of this loop got cheaper while we were running it.)

The loop is the asset; the model is a commodity input. Anyone reproducing this a year from now will use different models and the same method.

The Replit plan that made it work

The $100/mo Replit plan was the real unlock, and not for the reason you'd guess. It allowed multiple concurrent tasks, each on its own git branch.

  1. Fire off several fixes at once, each as its own task.
  2. Each lands on its own branch.
  3. Review each branch independently.
  4. Merge the good ones into main.
  5. Publish from main.

Parallel agents, a human review gate, no cross-contamination. If task 3 produced garbage, you don't merge task 3 — everything else still ships.

There's a second-order effect worth naming: because a bad attempt costs nothing to discard, you attempt more aggressive changes than you otherwise would. Isolation doesn't just prevent merge conflicts, it raises your risk tolerance.


What it cost — and why that's the real story

Total AI spend across the engagement was about $2,500 — and it's worth breaking that down, because the split surprised us.

Essentially all of it was builder spend. The auditing — the half of this method that produces most of the value — ran on a $20/month ChatGPT plan. No API billing, no per-audit cost, no metering. We ran cold deep-dive audits several times a day for three weeks and the marginal cost of each one was zero.

Sit with that for a second. The expensive part was making the changes. The part that decided which changes were worth making cost twenty dollars a month.

Which means the cheap version of this is genuinely cheap. Claude Code on a subscription handles the build half fine and is natively branch-and-commit shaped, so the review-and-merge discipline survives intact. Pair it with the same $20 ChatGPT plan and you have the entire method for a couple of hundred a month rather than $2,500. The Replit premium bought parallelism and a task UI — not a better outcome.

Now put that against the benchmark. A site like Dr. Sadati's probably cost $20k+ the traditional way, and that's likely just the build — a site at that level implies years of ongoing spend behind it: SEO retainers, agency fees, professional photography and video, content production. He earned the #2 position; the money wasn't wasted.

The point isn't a head-to-head. It's this:

That level of website is now achievable for a fraction of the price.

The bar used to require a $20k+ build plus years of sustained investment. It doesn't anymore. That outcome has dropped by roughly an order of magnitude in cost, putting it within reach of practices and small businesses that could never have justified the traditional number.

What money bought him that AI can't buy you

Look back at the scorecard. The client lost exactly two categories: reviews and authority (9.7 vs 10.0) and luxury presentation (9.7 vs 10.0).

That's not a coincidence, and it's the honest boundary of this whole method. Sadati's standing isn't only web spend — it's hundreds of reviews and a deep before-and-after library, accumulated over years of practice.

  • AI closes the design, structure, copy, technical, and presentation gap fast and cheap.
  • AI does not manufacture hundreds of genuine reviews or a decade of real before-and-afters.
  • What it can do is ensure every piece of real proof you actually have is surfaced, organized, cross-linked, and working as hard as it possibly can.

The scorecard corroborates the caveat instead of contradicting it. A clean sweep would have been less believable.


The dimensions

These were never a fixed list. Every cold audit invented its own rubric, and we cover that at length further down — it's the biggest caveat in this whole method. What follows is the set of dimensions that recurred most often across many runs, not a checklist the auditor was working from:

  • Pricing consistency
  • Homepage overall
  • Hero
  • Brand distinction
  • Treatment clarity
  • Before-and-after proof
  • Reviews and authority
  • Conversion / booking
  • Local SEO relevance
  • Technical structure
  • Luxury presentation
  • Mobile optimization

Some audits scored nine of these. Some scored fourteen, including several that never appeared again. Treat the list as a description of what the model tends to care about, not a specification.

Retune for your vertical, but notice the shape: roughly one-third persuasion (hero, brand distinction, luxury presentation), one-third proof (before-and-after, reviews and authority, treatment clarity), one-third mechanics (technical structure, local SEO, mobile, booking, pricing consistency).

A site that's a 10 on mechanics and a 6 on proof ranks and doesn't convert. Scoring the three families separately stops you over-investing in whichever one you personally enjoy.

The entire homepage, shown as four columns
The entire homepage, sliced into four columns to fit on one screen. Twelve dimensions is a lot of rubric — but this is how much page there is to grade.

One gap, stated honestly: there's no accessibility dimension. For a medical-adjacent business that's both an ethical gap and real legal exposure — ADA and WCAG demand letters are common in this sector. If you're building your own rubric, make it thirteen.

Score against named competitors, not against the void

"Rate my website out of 10" produces a meaningless number. "Rate my website out of 10 against these two specific competitors, on these twelve dimensions" produces a number that moves when you do work. Here the benchmarks were Dr. Kevin Sadati and Forever Ageless — the sites actually competing for the same searches, in the same few-mile radius.

Run two scores when the comparison isn't apples to apples

Both benchmarks are physician-led. Our client is non-surgical. On a naive comparison that gap is unclosable — no amount of web design puts an MD behind a business.

So we tracked two numbers, and the data justifies it emphatically. Overall, the client leads by 0.2. In the non-surgical facelift category, the same audit separates it from Sadati by 1.1 points — 9.9 to 8.8.

The segment score isn't a consolation prize. It's where the advantage is actually visible. If your competitive set includes someone with a structural advantage you can't replicate, split the score — otherwise you'll spend weeks trying to fix a business fact with CSS.


The loop

Cold deep-dive audit  →  1–10 scores + specific findings
        ↓
Convert findings into build prompts
        ↓
Agent works each fix on its own branch
        ↓
Review → merge to main → publish
        ↓
Re-audit with NO CACHE + live browsing to verify it's live
        ↓
Repeat until the audit runs out of things to fix

What the prompt actually looked like

No prompt engineering. No persona, no XML tags, no seventeen-line system preamble. This is close to verbatim:

do a deep dive of a med spa newportbeachskintightening.com, dr sadati, and forever ageless, do a deep, deep dive, from 1-10 side by side

That's it. And when the session needed pushing:

…no cache
…use browser skill

Four things in that scruffy little prompt are doing real work, and they're worth separating out because they're the transferable part:

  • Naming all three sites in one request. This is what produces a single comparison judged by one rubric in one sitting — the only kind of comparison that survives, as covered above. Auditing your own site alone gives you a number with nothing to lean on.
  • "side by side." The load-bearing phrase. It forces a table rather than three paragraphs of prose, and a table forces the model to commit to a value in every cell instead of writing around the ones it's unsure about.
  • "from 1-10." Forces numbers. Without it you get adjectives, and adjectives don't trend.
  • "deep, deep dive." Saying it twice genuinely produces a longer, more thorough pass. Not elegant. Works anyway.

And the two afterthoughts — no cache, use browser skill — are what stand between an audit of your website and an audit of the model's memory of your website. Add them every time. Just don't mistake them for a guarantee: they improve the odds of a fresh retrieval rather than ensuring one, which is why the caching trap below insists you confirm with a canary change.

Dictated, not typed — and unfailingly polite

Worth knowing where that phrasing comes from: the client dictated these prompts with speech-to-text. That's why they read the way they do — run-on, lightly punctuated, the domain occasionally mangled by the transcriber. Nobody sat composing them.

And they were courteous to the machine. A typical prompt didn't stop at the instruction; it asked nicely and thanked it afterwards:

…please do a deep dive audit, please do a good job
thank you

We have no evidence the pleading improved a single score. There's some published work suggesting politeness nudges model output, and plenty suggesting the effect is noise. We weren't running a controlled trial on manners, so we can't tell you.

What we can tell you is the part that isn't superstition: the person who got the most out of this workflow was not a prompt engineer. He talked to it. He asked for a deep dive, asked it to do a good job, said thanks, and then — this is the bit that actually mattered — ran it again tomorrow. No syntax, no framework, no template library.

So if you've been putting this off because you assumed there was a craft to learn first: there isn't much. Say what you want, name your competitors, ask for numbers, and be relentless about repeating it. Being nice to the robot is optional, and cheap enough that you may as well.

The reason to show you something this unpolished is that the sophistication in this method isn't in the prompt. It's in running the prompt a hundred times, from fresh sessions, and being disciplined about what you do with the answers.

We ran this multiple times a day for three weeks straight — writing prompts, shipping fixes, re-auditing. Not weekly. Multiple times a day. Audits are cheap and the compounding is real. You can watch it in the timestamps on our own audit captures: 20 July the client trailed Sadati in most categories. 27 July it was at 9.3 overall. 1 August it was at 9.8 and ranked #1.

Audit scorecard from 27 July showing 9.3 overall
Five days before the final run: 9.3/10, with the auditor already naming the differentiated offer as the strongest asset. The last half point took the longest.

Nobody tells you how much of this is waiting

Here's the part the demo videos skip: this loop is slow.

A single agent task can sit and think for five minutes before it writes a line. Publishing can take another five. So the actual rhythm of a "fast AI workflow" is: write a careful prompt, wait, review a diff, merge, wait again to deploy, then ask a second AI to go look at it — which often takes a couple of minutes itself. An instant response is a warning sign, though freshness is only really settled by a canary change, as covered later.

The speed is inconsistent, which is its own problem. Sometimes a task lands in under a minute. Sometimes the identical kind of task takes ten. It depends on load, on model, on how much the agent decides to reason, and on nothing you can see or control. You cannot plan a day around an average when the variance is that wide — so you plan around the slow case and treat the fast ones as a bonus.

Call it fifteen-ish minutes per round trip for one fix, on a normal day.

Which is the actual reason this took three weeks rather than three days. Not because the work was hard — most individual fixes were small — but because the cycle time is long and there were a lot of cycles. The parallel-branch setup exists precisely to amortize that: if you're going to wait five minutes anyway, wait five minutes for six things at once.

Plan your day around it. The mistake is sitting there watching the spinner. The move is to keep four or five tasks in flight and treat each one's completion as an interrupt. AI made the work cheaper and better here; it did not make any individual step quick.

AI is fast the way a bus is fast: excellent top speed, and it stops constantly.

Keep a scorecard. Every run, every dimension, dated. That table makes the trend visible, makes an outlier run obvious, and makes a regression show up as a drop in one column instead of a vague sense that something got worse. It's also the single best client deliverable this process produces.


The traps

Everything above is the clean version. Here's what actually goes wrong.

1. Caching will silently invalidate your audits

The big one. Your audit session must use no-cache and live browsing. Otherwise it reads a cached copy and delivers a confident, detailed, obsolete audit — and you spend the day fixing things you fixed last week.

There are only two hard problems in computer science: cache invalidation, naming things, and persuading a language model that it is looking at a cached copy of your website.

Timing is a hint, not proof. A real fetch takes a noticeable amount of time, and an audit that comes back instantly probably didn't retrieve anything. That heuristic served us well and we still use it — but be clear about what it is. Saying no cache in a prompt is a request, not a guarantee: it doesn't prove a Cache-Control header went anywhere, and it certainly doesn't prove every intermediary cache and the model's own recent context were bypassed. Fast responses are suggestive. They aren't evidence.

So verify with a canary instead. Publish a change only you know about — a new heading, a changed price, a fresh testimonial — then ask the auditor to tell you what it says. If it can quote the thing you shipped twenty minutes ago, it genuinely fetched the current page. If it can't, nothing else it tells you is worth reading.

That's the reliable check, and it costs one question. Ask for live browsing, ask it not to rely on cached knowledge, watch the timing as a soft signal — and then confirm with something only a fresh retrieval could know.

Press it and it will often admit the lapse — "you're right, I didn't do a live check." Useful, though also worth noting that a model agreeing it made a mistake isn't independent confirmation either; it's an agreeable machine agreeing with you. The canary is the part that actually settles it.

If that feels like an excessive amount of suspicion to direct at your own tooling: this article spends ten thousand words arguing you shouldn't take an AI's scoring at face value. It would be strange to then take its browsing on faith.

2. Sessions started on mobile often can't browse

Some sessions don't have live browsing available, or can't honor no-cache. In our experience this hit sessions started on mobile especially. Start audits on desktop and verify browsing works before trusting the first score.

Auditing a website from a phone session with no browser is a bit like proofreading through a mail slot. Confident, thorough, and describing a different document entirely.

3. The auditor's memory works against you

ChatGPT carries context between sessions. It remembers what your site used to be — and it would score a dimension lower for an inconsistency already fixed, because the stale version was in memory. The audit was grading a ghost.

The fix is blunt: explicitly tell it to update its memory. Don't assume a live fetch overrides a stale memory. Say the words.

4. A warm 10 is not a cold 10

Two different instruments, and conflating them will convince you you're done when you aren't:

  • Warm re-check — "you flagged X, we fixed it, look again with no cache." Finds the fix, often hands out a 10. This verifies.
  • Cold deep-dive — fresh session, no history, full audit from scratch. This discovers, and will surface a new batch the warm re-check never looked for.

Warm says 10, cold says "here are six more things." Both correct. Use warm re-checks to confirm deploys, cold audits to decide whether you're finished. You're only done when a cold audit comes up empty.

5. Audits vary — until you've run enough of them

This one we can show you. Here's a run from 20 July:

Earlier audit with a different rubric

Note the categories: Owner biography, Case-study information, Video testimonials, Services and treatment pages, SEO and search depth, Authority and credentials. Now compare to the 1 August scorecard at the top of this article: Homepage overall, Hero / first screen, Brand distinction, Treatment clarity, Before-and-after proof, Reviews and authority, Conversion / booking, Local SEO relevance, Technical structure, Luxury presentation.

Different category names. Different counts. Different column order — Sadati is scored first in July, the client first in August. Same site, same auditor, same instruction.

Side by side, the drift is stark:

20 July rubric1 August rubric
HomepageHomepage overall
Hero sectionHero / first screen
Owner biography(folded into Reviews and authority)
Before-and-after galleryBefore-and-after proof
Case-study information(gone)
Video testimonials(gone)
Services and treatment pagesTreatment clarity
SEO and search depthLocal SEO relevance
Authority and credentialsReviews and authority
(not scored)Brand distinction
(not scored)Conversion / booking
(not scored)Technical structure
(not scored)Luxury presentation

This is the single most important caveat in the article, so we'll say it plainly: a fresh cold audit does not necessarily grade the same things as the last one.

Four categories that existed in July were gone by August. Four categories that mattered in August didn't exist in July. Two were renamed into near-synonyms. Nobody asked for any of this — we gave the same instruction both times.

The consequences are practical and they change how you're allowed to use the numbers:

  • You cannot diff two cold audits category by category. "Video testimonials went from 9.5 to — " has no ending, because the second audit never scored video testimonials.
  • A category appearing for the first time is not a regression. The first time "Technical structure" showed up we briefly thought something had broken. Nothing had broken; the rubric had simply grown a new opinion.
  • Only two things survive across runs: the within-run ranking (all three sites judged by the same rubric in the same session) and the overall trend across many runs. Everything else is comparing apples to a slightly different fruit each week.
  • If you want a stable rubric, you have to pin it yourself — paste the twelve dimensions into every fresh session rather than asking for "a deep dive audit" and accepting whatever categories come back. We eventually did this. It's the difference between a measurement and a conversation.

The upside: because the auditor keeps reinventing its own criteria, it keeps noticing things a fixed checklist would have stopped looking for. The drift is annoying and it is also where several of the best findings came from.

The drift settles down on its own

Here's the part that makes this liveable, and we only noticed it after weeks of runs: eventually the auditor starts reusing its own rubric.

Once enough context has accumulated — enough previous audits of the same three sites, in the same account, with the same recurring instruction — it stops inventing a fresh set of categories each time and starts scoring against roughly the same list it used last time. The rubric converges without being asked.

And you don't have to wait for it. You can simply tell it to use the same rubric as last time, and it will. Naming the previous categories, or just saying score it on the same dimensions as the last audit, is enough to lock the shape while leaving the judgment fresh.

Notice what that means: this is the same memory that was working against us three traps ago. When it remembered a stale version of the site, memory was the enemy and we had to explicitly tell it what had changed. When it remembers the rubric, memory is the whole point — it's what turns a series of unrelated opinions into something you can actually chart.

That gives you three usable positions, and they're genuinely different tools:

  • Fully cold, no pinning. Maximum discovery, minimum comparability. Best early, when you still want it finding categories you hadn't thought of.
  • Let context accumulate. The rubric stabilises by itself over repeated runs. This is where we ended up, and it's the least effort.
  • Explicitly pin the rubric. Paste the dimensions, or tell it to reuse the last set. Most comparable, and the right call once you're tracking a number rather than hunting for problems.

Early on you want drift. Later you want consistency. The mistake is wanting both at once and then being confused by your own scorecard.

Two responses:

  • Volume standardizes it. After enough runs the structure and scoring converge. Early scores are noisy; the trend is the signal.
  • Don't argue with a bad run. If a score contradicted an established trend, we opened a new session and ran it again. Cheaper than negotiating with a model about its own output, and it doesn't pollute the session with your objection.

6. Non-visual work needs explicit verification

Technical SEO — JSON-LD structured data especially — has no visual signal. A browsing model looking at a rendered page will not notice you added LocalBusiness schema, and will score "technical structure" off whatever it can see.

Instruct it explicitly: go find the structured data, tell me what types are present, confirm the fields. Otherwise your best technical work is invisible to the grader.

7. Content changes cause regressions

We had it near-perfect. Then the client changed the pricing packages and added a new one. It didn't propagate everywhere. Pricing said different things on different pages, and the score dropped.

One price achieved independence somewhere between the brief and production and started quoting its own numbers. We found it eventually. It had been living on the pricing page under an assumed value.

This is why pricing consistency is a standing dimension, not a one-time task. Any client edit — pricing, hours, services, staff — can regress a finished site. The loop isn't a project you complete. It's a monitor you keep running.

The pricing packages section
The packages section. When the client introduced a new package and it didn't propagate everywhere, pricing consistency dropped and took the overall score with it.

8. The auditor is biased toward finding something

The flip side of trap 4: a model asked to audit will produce findings whether or not real problems remain. That's the job you gave it. As you approach 10, a growing share of findings are manufactured nitpicks rather than genuine defects.

So "it ran out of things to fix" is a judgment call you make, not a state the model announces. The signal to watch: findings stop being specific and start being generic advice.

9. You need a human veto

The auditor will sometimes recommend things that are right for the rubric and wrong for the business — publish full pricing when the business deliberately gates it, soften the exact differentiator that's the wedge, add a section that dilutes a focused page. Implement findings selectively. The branch-per-task workflow exists partly to make rejecting a change free.

10. Push back — it is not always right

The auditor states everything with equal confidence, including the parts it's wrong about. If something doesn't make sense, challenge it.

Interrogate every score movement. When a number dropped or jumped, we pressed on why. Sometimes there was a real cause. Sometimes there wasn't, and it folded.

Be most suspicious of good news. There was a run where it handed out 10/10 immediately after we told it a fix had been made. That's the model being agreeable, not the site being perfect — a warm session over-crediting a change you just announced to it.

It's the same reflex as the caching check: if the answer arrives too easily, verify it. A warm session that congratulates you right after you announce a fix is not evidence. A cold session that can't find anything is.


The single most useful prompt

When a category scored low, we stopped reading the finding and started asking:

What would make this a 10?

This converts a score into a work order. Instead of "before-and-after proof: 8.5," you get a specific list of what's missing and what would close the gap. The number stops being a grade and becomes a spec.

Do it for every dimension below target. It's the highest-leverage habit in the whole method, and it's the reason the loop terminates at all — you're not guessing what the auditor wants, you're asking it to write the ticket.


What the auditor caught that we didn't expect

It worked at three quite different altitudes, which is unusual for a single tool.

Hard technical facts. It pulled real Google Search Console data and found broken and unindexed URLs:

Google Search Console indexing issues
- 95 pages crawled but not indexed - 8 pages discovered but not indexed - 32 not-found URLs - 7 URLs blocked by robots.txt - 2 noindex URLs - 2 redirects

That's the "technical structure" dimension with receipts, not vibes.

Domain relevance. It identified local SEO gaps — what was thin or missing for the geographic target.

Subjective composition. It flagged homepage layout proportions — whether the page was too long or too big. This is the one worth pausing on, because it's a judgment about pacing and density, not a rule check. A homepage can pass every technical audit and still be exhausting to scroll. Most tools do exactly one of these three things. It did all three.


What actually moved the score

The AI proposed the sections, not just the copy

We didn't hand it a wireframe. The agent came up with the homepage section ideas itself: comparison reviews, a "who this treatment is for" qualification section, and a dedicated reviews section.

Who the treatment is for, with candidacy criteria
A qualification section the AI proposed. Nobody briefed it — it identified that the homepage had no section answering "am I the right candidate."

That's a different and more valuable mode than "make this div prettier." It was identifying missing page architecture and proposing the section to fill the gap. If you're only asking AI for styling and copy, you're using maybe a third of it.

The seven technologies section
Another section the agent proposed: the seven technologies explained as one combined treatment rather than a list of machines. This is what moved "treatment clarity."

Marketing verbiage → compliant medical verbiage

It also rewrote marketing language into compliant medical language. In aesthetics this isn't a tone preference — language that overpromises outcomes, or makes claims a non-physician practice isn't permitted to make, is genuine liability.

You can see the result in the disclaimers throughout the finished site. Under the candidacy criteria above: "Women outside these age ranges may also be appropriate candidates. Treatment suitability and results vary. A consultation is required." Under the press mentions:

Local NAP block and press mentions with disclaimer
"These programs featured the technologies we use… They do not endorse any specific provider. Segments are shared for educational purposes." Compliance language, generated and placed by AI.

General lesson: in regulated verticals, put compliance language review into the loop deliberately. Cheap to ask for, expensive to miss.

Authority: give the person a background

We rebuilt the About section around credentials, certificates, and a background timeline — not a paragraph of adjectives, a verifiable trajectory.

About / founder section

This is worth more than the score movement suggests. Medical aesthetics is squarely YMYL territory, where Google leans hardest on E-E-A-T signals. The auditor's rubric and Google's quality guidelines are pointed at the same thing, which is a large part of why the rubric works at all.

Press and media mentions with compliance disclaimer
Media mentions, carefully worded. Authority signals are only worth having if the claims around them survive scrutiny — hence the disclaimer under the logos.

We added Google reviews, Yelp reviews, before-and-after galleries, and video testimonials — then cross-linked them to each other. This is the single best illustration of the idea:

Video testimonials cross-linked to Google, Yelp and Instagram
Each testimonial carries "Read her Google review," "Read her Yelp review," "View original on Instagram," "Read video transcript" — feeding into an aggregate 4.9 from 65 Google reviews.

The cross-linking mattered more than the individual additions. Proof sitting in isolated silos reads as decoration. Proof that references itself — a before-and-after linking to the video testimonial from the same patient, linking to their Google review, linking to the original Instagram post — reads as a verifiable trail. The audit noticed.

Google review cards on the homepage
The review wall those testimonials feed into — individual Google reviews surfaced on the page with the treatments each client had, not just an aggregate star rating.

Authenticity beats polish

The homepage originally led with glamour shots: beautiful, professional, and interchangeable with every other site in the category. We replaced them with the client's real, non-AI before-and-after videos.

Real before-and-after video section
"Real Clients. Real Results. No Surgery." Lower production value than glamour shots. Dramatically higher scores on before-and-after proof and brand distinction.

This is the counterintuitive lesson of the entire project. In a category where every competitor can now generate flawless imagery, flawless imagery is worthless and evidence is priceless. The AI-generated-everything era has quietly made "obviously real" a competitive advantage.


Page speed, and the fight it picked with everything else

We optimized performance on both mobile and desktop — not desktop only — and landed near 100 on both. It took several rounds: optimize, re-measure, feed the numbers back, go again.

Worth knowing: ChatGPT can run PageSpeed from inside the chat session. You don't have to leave the audit, run Lighthouse separately, and paste results back. That keeps the whole optimize → measure → optimize cycle in one place, which is most of why running it several times was practical rather than tedious.

Mobile homepage
The same hero at 390px. Mobile is the harder target and the one that matters.

There's a reason it wasn't a one-prompt fix, and it's the most interesting engineering problem in the project: we had just added a lot of video. Video testimonials, before-and-after videos, clips pulled from social. Video is the heaviest thing you can put on a page, and "near-100 on mobile" and "rich video proof everywhere" pull directly against each other.

This is the one place where two dimensions genuinely conflict, and resolving it is real work rather than prompting: lazy-load anything below the fold, poster images instead of autoplay, facade pattern for embeds, modern codecs, correct sizing per breakpoint, and never let a testimonial carousel block first paint.

The one place we can show you a real measurement

Everything above this point is a model's judgment. Page speed isn't. So here are all three sites run through Google PageSpeed Insights on mobile, the same day this article was written.

Mobile only, deliberately. Desktop isn't where the difficulty lives — everyone scores well on desktop, including both competitors, so a desktop table proves nothing about anyone. Mobile is the constrained test: slower simulated CPU, throttled network, smaller viewport, and it's where most of this client's traffic actually arrives. If you only publish one number, publish the mobile one. It's the one that was hard to earn.

Mobile (PSI)NB Skin TighteningDr. Kevin SadatiForever Ageless
Performance9964–7665–67
Accessibility9797–100100
Best Practices100100100
SEO100100100
PageSpeed Insights mobile — Newport Beach Skin Tightening
Client site, mobile: Performance 99, First Contentful Paint 1.1s, Largest Contentful Paint 2.1s — on a page carrying multiple videos.
PageSpeed Insights mobile — Dr. Kevin Sadati
Dr. Kevin Sadati, mobile.
PageSpeed Insights mobile — Forever Ageless
Forever Ageless, mobile.

A 99 against a 64 and a 65 is a bigger gap than anything on the ChatGPT scorecard — and unlike those numbers, this one is reproducible by anyone reading this article. Go run it yourself.

Note the ranges in that table, though. We ran the competitors twice and got 76 then 64, and 67 then 65. Lighthouse lab scores drift between runs — network conditions, ad and tracker loading, server variance. So even the "objective" measurement has a confidence interval, and the honest read is "~99 versus mid-60s-to-70s," not "99 versus 64."

Which is a useful corrective to the whole article: the AI scores vary between runs, and so does the real instrument. The difference is that PageSpeed tells you its methodology and anyone can re-run it. That's the standard the model scores don't meet — and it's why the trend, not the reading, is what you should trust in both cases.

And now the part that complicates the win

Look closely at the Sadati screenshot above. It has a section the other two don't:

Core Web Vitals Assessment: Passed — LCP 1.6s · INP 80ms · CLS 0 · FCP 1.4s

That's Chrome User Experience Report field data — measurements from real people on real devices over a 28-day window. His site scores 64 in the lab and passes Core Web Vitals with actual users.

Our client's report, and Forever Ageless's, both say "No Data" in that section.

There are two things worth sitting with here.

First, lab score is not user experience. A 99 in a simulated test on a datacenter connection is a good sign, not a result. Sadati's mid-60s lab score with passing field data arguably describes a faster real-world site than a 99 with no field data at all. We can't claim otherwise, so we won't.

Second, "No Data" is a function of age, not quality. CrUX only reports once a site has accumulated roughly 28 days of real traffic at sufficient volume. Sadati's site has been around for years and clears that bar easily. Our client's site is substantially newer and rebuilt very recently, so there is simply no sample yet — and there couldn't be, no matter how good the page is.

This is worth being precise about, because it's easy to garble: the lab score is the thing you control and we have it (99). The field data is the thing that accrues, and it can't be built, prompted, or bought — only waited for. A brand-new site showing "No Data" isn't failing a test. It hasn't been given the test yet.

It is, though, the same moat surfacing in a third independent place. Reviews, before-and-afters, and now real-user performance history: all three are time-denominated. That's the consistent shape of what AI couldn't close.

So the honest version of this comparison: we built a technically faster page, and he has a more established website. Those are different achievements, and only one of them was ever available to us in three weeks. Check back in a year and the field data column will exist — that one just needs the clock to run.

The duplicate-build incident

There was a second cause, and it's the best single argument for this entire method.

To handle mobile and desktop, the builder had quietly produced two separate versions of the page — a mobile one and a desktop one, both shipped in the markup. It looked fine. It worked.

The auditor caught it and wanted it consolidated into one responsive version with the DOM element count reduced.

Duplicating markup is the path of least resistance for an agent: it's the fastest way to make both viewports look right. But it doubles the DOM (excessive DOM size is a direct Lighthouse flag), tanks mobile performance, and leaves you maintaining two copies of every future content change — a consistency regression waiting to happen, on a site whose rubric grades consistency.

The agent's answer to responsive design was to build the site twice. Somewhere out there, a perfectly good media query is still sitting by the phone waiting for a call that never came.

Sit with the shape of that. The builder created the problem and would never have flagged it — its own work looked correct. The auditor, which knew nothing about how the page was built and saw only the result, caught it immediately.

That's the separation-of-duties argument in one concrete incident. Not a philosophical preference — a real defect that existed because one model built it, and got found because a different model looked at it.


SEO and local pages

Alongside general SEO we added more local SEO pages, expanding location and service coverage rather than trying to rank one page for everything. That's what moved "local SEO relevance," and it fits the business: this isn't a national play, it's a few-mile-radius play against two competitors in the same city.

One caution, because AI makes the wrong version effortless. Mass-produced local pages that differ only by city name are doorway pages, and Google penalizes them. The version that works has genuinely distinct content per page — local landmarks, location-specific pricing, proof from that area. It's trivially easy to generate hundreds of the bad kind in an afternoon. Don't.

The entire homepage on mobile, shown as seven columns
The same page on a phone, sliced into seven columns. One responsive build — not the two the agent originally shipped — and the version that had to score near 100.

The media work nobody expects an agent to do

The agent could look at raw video and identify the shots and who was in each one. From there we directed edits by chat prompt:

  • "Make the before and after eye-level and head position consistent."
  • "Move the person's head higher in the frame."

It also pulled videos from social media and reformatted them for web — aspect ratio, encoding, sizing. No video editor in the loop.

Transcribe every video — it's free SEO

The one piece of this we'd tell everyone to do regardless of the rest of the method: run local Whisper over every video on the site. Replit's agent can do it, Claude Code can do it, and it costs nothing — the model runs on your own machine, no API, no per-minute billing.

Why it matters: a video is invisible to a search engine. You can have the most persuasive patient testimonial in the county and, as far as Google is concerned, that page contains a <video> tag and some whitespace. A transcript turns every spoken sentence into indexable text — in the patient's own words, using the phrases real people actually say about the treatment, which is exactly the long-tail vocabulary you'd otherwise pay a copywriter to guess at.

Look back at the testimonials figure above and you'll see the result: every card carries a "Read video transcript" link. Those are Whisper output, lightly cleaned up.

Three things you get from one cheap step:

  • Indexable content where there was none, in natural patient language.
  • Accessibility — captions and transcripts are the single most-cited WCAG gap on video-heavy sites, and remember there's no accessibility dimension in our rubric to catch it.
  • VideoObject structured data with a real transcript field, which feeds the technical structure dimension.

It's the highest ratio of SEO benefit to effort anywhere in this project, and the reason it gets skipped is simply that transcribing video used to be tedious and expensive. It isn't either any more.

And then there's the benefit nobody expects. Once the video is text, the auditor can read it — which means video stops being a black box in your content and becomes something an AI can check like any other page element.

Specifically, it can tell you whether a video is actually aligned with the marketing message around it. That sounds abstract until you hit a real case, and they're common:

  • A testimonial praises a treatment the page isn't selling, or names an older service you renamed six months ago.
  • The page promises "no downtime" and the patient cheerfully mentions being red for two days.
  • A client says something warm and enthusiastic that, written down in plain text, is an outcome claim a non-physician practice isn't allowed to make — the exact compliance problem covered earlier, hiding in a video where no one thought to look for it.
  • The hero video's emotional register is completely different from the headline sitting next to it.

Every one of those is invisible while the content is locked inside an MP4. A human would have to sit and watch every video with the page copy open beside them. Nobody does that, which is why misaligned video sits on sites for years.

Transcribe first, and the auditor does that pass for free — on every video, every time you run it, forever.

Same story on stills. It cleaned up scanned certificates that came in skewed, yellowed and low-contrast — because having the credential is worthless if the scan looks like a fax — and removed an unwanted person from a photo that was otherwise usable — full Back to the Future treatment, except the erasure took about nine seconds and nobody's hand started fading out. One prompt, and a bystander who'd been standing in the shot since 2019 simply stopped having been there. The photograph now shows a past that never happened, which is either a triumph of accessible image editing or a small crime against the historical record, depending how strongly you feel about a stranger's elbow.

The broader point: the asset pipeline is part of the loop, not a prerequisite to it. The usual blocker on "add credentials and real proof" is that the raw material is a shoebox of bad scans and awkward photos, and cleaning it up traditionally means a designer and a budget. That blocker is mostly gone.

Between the video edits, the photo retouching, the copywriting and the front-end build, this project quietly did the work of about four specialists. Everyone asks whether AI is coming for our jobs. On this evidence: yes, obviously — except for the one job it still can't do, which is sitting there at 2am telling it that no, that's still wrong, do it again. Job security, of a sort. Not the sort anyone asked for.

One line worth holding, though: retouching presentation — deskew, contrast, remove a bystander — is fine. Retouching before-and-after results is not. It destroys the authenticity advantage the site was built on, and in medical aesthetics it's a regulatory problem. Fix the photo, never fix the result.


The part of this that isn't about AI at all

Here's the caveat that decides whether any of the above will work for you, and it has nothing to do with models, prompts or rubrics.

We got to the top of that scorecard because the client did the work.

He had real before-and-after photography. He had video of actual patients willing to appear on camera and say their names. He had certificates, a training history, a coherent professional timeline. When we asked for raw material, it arrived. When we asked him to film something, he filmed it. When we said the glamour shots had to go and be replaced with unretouched results, he agreed — which is not a small thing to agree to.

Most clients won't. In our experience the two most common walls are:

  • "I don't think that content is good enough to put on the site." People are strange about their own proof. A practitioner will sit on a folder of genuinely persuasive before-and-afters because the lighting isn't perfect, then ask why the site doesn't convert. The best material is almost always already on someone's phone.
  • Simply not doing it. Gathering assets is homework, and homework competes with running a business. A project can stall for a month waiting on twelve photos.

And then there's the review problem

Separately, and more often than you'd think: the client may not have review collection properly set up at all. No claimed Google Business Profile, no Yelp page under their control, no process for asking a happy client to leave a review, reviews scattered across a personal profile and a business one. Sometimes there's a Google listing nobody has ever logged into.

You cannot surface social proof that doesn't exist, and you cannot cross-link reviews that were never collected. Look back at the final scorecard: the client scored 9.7 on reviews and authority — a strong number, built on a real 4.9 from 65 Google reviews. With no review infrastructure that isn't a 9.7. It's a 5, and no amount of prompting moves it.

What this means practically

Before you quote this method to anyone, audit the client, not just the site:

  • Do they have real before-and-afters, case studies, or results they're willing to publish?
  • Will anyone go on camera?
  • Is the Google Business Profile claimed, correct, and collecting reviews?
  • Is Yelp claimed? Is anyone asking for reviews at all?
  • Are there credentials, certificates, or a history worth documenting?

Those answers set the ceiling before you write a single prompt. AI can take a site with good raw material and make it excellent. It cannot conjure the raw material. A client with nothing to show gets a beautifully built, technically flawless, entirely unconvincing website — which will score well on mechanics and lose on everything that makes someone book.

The honest framing for the whole method: this is a multiplier on what a business already has. Multipliers are wonderful and they do not work on zero.

The playbook, condensed

  1. Separate the auditor from the builder. Non-negotiable.
  2. Pick 8–13 scoring dimensions for your vertical. Persuasion, proof, mechanics — and accessibility.
  3. Name your real competitors. Score against them, not against nothing.
  4. Split the score if a competitor has a structural advantage you can't replicate.
  5. Always request live browsing and a fresh retrieval. A fast response is a warning sign, not proof — verify freshness with a recent canary change.
  6. Start sessions on desktop. Verify browsing works before trusting a score.
  7. Tell the auditor to update its memory whenever something significant changes.
  8. Warm re-checks verify deploys. Cold audits decide if you're done.
  9. Explicitly ask it to verify non-visual work — schema, meta, structured data.
  10. One fix per branch. Review, merge, publish.
  11. Re-run a bad audit in a fresh session instead of arguing with it.
  12. Keep a dated scorecard of every run across every dimension.
  13. Apply a human veto. Rubric-right and business-right aren't the same thing.
  14. Keep running it after you hit the top. Client edits regress.

Expect a non-linear curve

8.4 → ~9.5 is the fun part: add the missing sections, fix the hero, ship the schema, stack the proof. The last half point is almost entirely consistency and polish — pricing matching everywhere, framing consistent across before-and-afters, compliant verbiage on every page. Slower, less satisfying, and the stretch where most people quit.

This is a retainer, not a project

The pricing regression proves it. Any client content edit can regress a finished site. That makes this a recurring service — scheduled cold audit, scorecard delivered, regressions fixed. For an agency, that's the actual business model hiding in this post.


What a "10" actually is

We flagged at the top that the title is hook bait. Here's the fuller version of why that matters, now that you've seen the method.

A 10/10 website is just what the AI thinks a 10 is. It's an opinion, formed from the model's training and knowledge — not an objective measurement, not a law of physics. No external authority certifies it. We invented the rubric; the model applied it. And as the two rubrics side by side in this article show, it doesn't even apply it identically twice.

But it does a genuinely good job rationalizing the scores. It explains why a dimension is a 9.7 and what specifically would take it to a 10. The reasoning is legible, specific, and actionable — which is exactly what makes the loop work at all.

The score is an opinion. The reasoning behind the score is the deliverable.

You're not buying a certified grade. You're buying a structured, consistent, competitor-aware critique that tells you what to do next — plus a number whose value is that it moves in the right direction when you do the work.

That's also why the earlier advice holds: volume standardizes it, because enough runs make the opinion converge. And cold-vs-warm matters, because an opinion is only ever as good as the information it was formed on.


How we knew we were done

We kept asking the same question after every audit: what would make this a 10?

For three weeks, the answer was a list of fixes.

Now the answer contains no fixes at all. Nothing about the site. The only things left are:

  • More reviews
  • More before-and-afters
  • Standardized before-and-afters — consistent framing, angle, lighting and eye level across the whole library, so the set reads as a controlled series rather than assorted snapshots

That's the real ending, and it closes the loop on the caveat from the top of this article. AI closed every gap it could close. What's left is precisely the stuff it cannot manufacture: accumulated real reviews, and a deep library of real results.

It also matches the scorecard exactly. The only two categories the client didn't win were reviews and authority and luxury presentation — the two things you earn over years rather than build in three weeks. The auditor and the endpoint agree, which is the most convincing thing about either of them.

Note that one of those three items isn't an accumulation problem. Standardizing the framing on the before-and-afters you already have is craft work — exactly the tedious, high-impact, low-creativity task the agents turned out to be good at. So the remaining list is two things the business has to earn, and one thing we can still go do.

The closing qualification section
The last thing a visitor reads: a direct invitation to find out whether they're a candidate. Conversion scored 9.9 — the one category where being smaller than a surgical practice is an advantage.

The website is finished. The business isn't, and never will be. The score isn't the product. The loop is. Ours is still running.


Reproduce it yourself

We'd rather you checked this than took our word for it, so here is the evidence, including the parts that are incomplete.

What we're publishing

Artefact
Final scorecard, 1 Aug 2026scorecard-2026-08-01.csv
Earlier rubric, 20 Jul 2026scorecard-2026-07-20.csv
PageSpeed runs, 1 Aug 2026pagespeed-2026-08-01.csv
Raw audit screenshotsthe four unretouched captures shown throughout this article
Site screenshotsevery figure here is an unedited capture from the live site, 1 Aug 2026

The audit prompt, verbatim, dictated by the client:

do a deep dive of a med spa newportbeachskintightening.com, dr sadati, and forever ageless, do a deep, deep dive, from 1-10 side by side

…with no cache and use browser skill appended when the session needed pushing.

The subjects

Clientnewportbeachskintightening.com
Benchmark 1drkevinsadati.com — physician-led
Benchmark 2foreveragelessinc.com — physician-led
Audit dates shown20 Jul, 24 Jul, 27 Jul, 1 Aug 2026
PageSpeed date1 Aug 2026, mobile strategy, two runs per competitor
AuditorChatGPT, $20/mo plan — GPT-5.5 at the start, GPT-5.6 Sol by the end
BuilderReplit Agent, Power mode (Opus; Fable in rotation for part of it)

What we can't give you, and why

Being straight about the gaps, because a reproducibility section that only lists strengths isn't one:

  • No raw chat exports. The audit outputs exist as photographs of the screen, not text dumps. We've transcribed them faithfully and published the screenshots unretouched, but you're trusting our transcription. Set up your own audits with exports if you want cleaner evidence than that.
  • No "before" screenshots. We didn't capture the pre-project site, which is the single most annoying omission here — the glamour-shots-to-real-video swap is described but not shown. If you're running this, screenshot everything on day one.
  • No per-change revision log. Branch names and merge order weren't kept in a form worth publishing.
  • The 8.4 starting score comes from an early audit we didn't screenshot. Treat it as reported, not evidenced. The scored trajectory we can show starts at 9.3 on 27 July.
  • PageSpeed is lab data. Explained at length above: the client site has no Chrome UX Report field data because it's too new to have accumulated any, while Sadati's site does and passes.

The part that is fully reproducible right now: open PageSpeed Insights, run those three URLs on mobile, and compare. That's Google's instrument, not ours, and it will disagree with us by a few points because lab scores drift.


Go and try it on your own site

Genuinely — do. Open a fresh session, name two competitors you actually lose business to, ask for a side-by-side from 1 to 10, and read what comes back. It costs you twenty minutes and nothing else, and you'll learn more about your own site in that one audit than in a month of staring at analytics.

Fair warning: it's addictive. We didn't expect that part.

Something happens when a fuzzy worry — the homepage feels a bit flat — turns into "Hero: 9.2". Suddenly it's a number, and numbers want to go up. You fix the hero, re-run it, and it says 9.6. So you fix one more thing. Then you find yourself at eleven at night thinking the before-and-after section is probably a 9.4, and I know exactly what would make it a 9.7.

That's the loop's real trick. Not the intelligence — the scoreboard. Website work is usually a fog of opinions with no feedback until a quarter of traffic data eventually arrives. Give someone a number that moves the same day they do the work and the whole thing turns into a game they want to keep playing. Three weeks of daily audits didn't take discipline. It took the opposite: knowing when to stop.

Which is the one caution worth repeating. A scoreboard will happily keep you busy after the work stops paying. There's a difference between this fix wins bookings and this fix wins a tenth of a point, and around 9.7 they stop being the same thing. We kept going to the end because the remaining findings were still real. When they turned into polish for its own sake, we stopped — and the honest answer to "what would make this a 10" had become more reviews, which no amount of clicking refresh will produce.

Chase the score while it's pointing at real problems. Notice when it stops.


Want one of these?

If you've read this far, you already know the method isn't secret — it's written out above, and you're welcome to run it yourself. Most of it costs a $20 subscription and a lot of patience.

What it actually takes is three weeks of running the loop several times a day, the judgment to know which findings to ignore, and someone willing to keep pushing back on a confident machine at 2am. If you'd rather not do that part, that's what we're for.

If you want your own 10/10 website — get in touch. Tell us the two or three competitors you want to be measured against, and we'll run the first cold audit on your current site so you can see exactly where you stand before committing to anything.

You'll get the same thing this article is built on: a scorecard, a list of specific fixes, and an honest read on which of them are worth paying for.

Need Help With Your Website?

I fix these problems every day. Send me a message and I'll take a look.

Get Help Now
CallTextMessage