How to Make a 10/10 Website with AI
We used ChatGPT as an independent auditor and Replit Agent as the builder to take a client site from 8.4 to #1 against two named competitors in a controlled AI audit. The exact loop, the real prompt, the scorecards, measured PageSpeed numbers, a downloadable evidence pack, and every trap that cost us a day.

Most people use AI to build websites. That's the boring half.
The interesting half is using a second AI to grade the site — coldly, against real named competitors, on a 1–10 scale — and then feeding those findings back to the builder. Do that enough times and something strange happens: the score stops being a vibe and starts being a number you can actually move.
We ran this loop on a client site — newportbeachskintightening.com, a non-surgical skin tightening studio serving Newport Beach.
Two timelines, and it's worth separating them. Building and refining the site was a months-long project. The audit loop described in this article was a three-week sprint at the end of it — several cold audits a day, every day, until the findings dried up. Everything below is about those three weeks.
Going in, the site scored around 8.4. It finished ranked #1 against both benchmark competitors in the audit, with multiple categories at a perfect 10.0, and it stopped not because we declared victory but because the auditor ran out of things to fix.
One qualification up front, because the title oversells it: that #1 is a placing in a repeated head-to-head AI audit against two named competitors. It is not a Google ranking, not a market position, and not a claim that this is the best website in Newport Beach. What it means precisely — and what it doesn't — is the subject of the next section.
Here's exactly how it worked, including the parts that don't show up in the highlight reel.
Everything in this article is live and checkable: newportbeachskintightening.com. Open it alongside the scorecards below and judge the calls yourself — that's rather the point of publishing the rubric.
What's in here
It's a long read. Jump to what you need:
- The final scorecard
- The two-AI setup
- What it cost
- The dimensions
- The loop
- The traps
- The single most useful prompt
- What actually moved the score
- Page speed vs. video
- Video, images and transcripts
- What depends on the client
- The playbook, condensed
- What a "10" actually is
- How we knew we were done
- Reproduce it yourself
- Try it on your own site
What the auditor concluded
Here's the final scorecard, from a cold audit run on 1 August 2026.
Read it for what it is. These aren't measurements. Nothing was instrumented, no users were tested, no stopwatch was involved. They're numbers a language model produced by reasoning about three websites and then expressing that reasoning as digits. A 9.8 doesn't mean the homepage is 98% of something. It means the model looked, thought, and landed there.

| Category | NB Skin Tightening | Kevin Sadati | Forever Ageless |
|---|---|---|---|
| Homepage overall | 9.8 | 9.6 | 9.1 |
| Hero / first screen | 9.9 | 9.2 | 8.8 |
| Brand distinction | 10.0 | 9.8 | 9.0 |
| Treatment clarity | 9.9 | 9.4 | 9.1 |
| Before-and-after proof | 9.8 | 9.8 | 8.6 |
| Reviews and authority | 9.7 | 10.0 | 9.2 |
| Conversion / booking | 9.9 | 9.7 | 9.4 |
| Local SEO relevance | 9.9 | 9.8 | 9.4 |
| Technical structure | 9.8 | 9.4 | 9.2 |
| Luxury presentation | 9.7 | 10.0 | 9.6 |
| Overall | 9.8 — #1 | 9.6 — #2 | 9.1 — #3 |
And the segment score, which is the one that actually matters — more on why below:
| Non-surgical facelift category | Score |
|---|---|
| Newport Beach Skin Tightening | 9.9 — #1 |
| Kevin Sadati | 8.8 — #2 |
| Forever Ageless | 8.5 — #3 |
Overall, the client wins by 0.2. In its own category, it wins by 1.1. Hold that thought.
One caveat to carry with you from here: this is one rubric, from one session. A fresh cold audit a fortnight earlier scored a partly different set of categories entirely — four of these ten didn't exist in it. We show the side-by-side later, because it changes what these numbers can honestly be used for.
The auditor also disclaims its own scope, which is worth quoting because it's the model being appropriately modest about what a website score measures:
…quality scores—not medical outcome rankings.
So why bother with the numbers at all?
Because a made-up number can still be a useful one, provided you're clear about which part is load-bearing.
What's weak: the absolute value. 9.8 is not a fact about the website. Run the audit again in a fresh session and you might get 9.6, with a partly different rubric — we show exactly that happening later in this article.
What's stronger: the comparison within a single run. All three sites in that table were judged by the same reasoning, against the same criteria, in the same session, minutes apart. Whatever bias the model has, it applied it to all three. So "the client ranked above both competitors on 8 of 10 categories" survives a lot better than "the client scored 9.8."
What's strongest: the direction of movement over many runs. One score is an opinion. Forty scores trending upward, while you make specific changes the auditor asked for, is a signal — not because any individual number is right, but because it's hard to move a biased instrument consistently in one direction by accident.
That's the whole epistemic basis for this method, and it's worth being blunt about: we are not measuring website quality. We are using a model's judgment as a repeatable critic, and treating its consistency — not its accuracy — as the thing we can rely on.
Except where it genuinely is measuring
Not every dimension is judgment. A few are real numbers the auditor fetches rather than invents:
- Page speed. ChatGPT can run PageSpeed from inside the chat session — you don't leave the audit to go check Lighthouse manually. Those are actual measurements.
- Technical and indexing. The Search Console figures further down — 95 crawled-not-indexed, 32 not-found — are Google's numbers, not the model's.
So the rubric is a mix of measurement and judgment, and it's worth knowing which is which. "Mobile optimization: 96" is a fact. "Luxury presentation: 9.7" is a read. Both are useful; only one is checkable.
The measured dimensions also do quiet work anchoring the judged ones. If the auditor's page-speed number matches what PageSpeed actually reports, that's small evidence it's looking at your real, current site rather than a cached memory of it — which, as the next section explains at some length, is not a given.
About that title
"10/10" is hook bait. You clicked it, so it worked.
What the auditor actually landed on was 9.8 overall and 9.9 in the category, with individual dimensions at 10.0. And per the section above, even those digits are a model's judgment rather than a measurement — a different session would land somewhere slightly different, on a partly different rubric. There is no certifying body handing out 10s. There's a model with an opinion and a consistent enough one to be useful.
We're saying this on the way in rather than burying it, because the entire method below depends on treating the score as a measuring instrument rather than a trophy. If you'd rather have the honest version of the headline: how to make a website that beats its two best competitors on a cold, adversarial audit, for about a tenth of what that used to cost. Less catchy. Same article.
The two-AI setup
The whole method rests on a separation of duties:
| Role | Tool | Job |
|---|---|---|
| The auditor | ChatGPT with live browsing | Cold deep-dive audits. Scores 1–10 across ~12 dimensions against named competitors. Never touches the code. |
| The builder | Replit Agent (Power mode) — or Claude Code | Takes audit findings, writes the changes on isolated branches. Never grades its own work. |
That last clause is the entire point. The thing that builds the site cannot be the thing that scores the site. An agent that just wrote a hero section will tell you the hero section is excellent — it has no distance from the work. A separate model, given a cold prompt and a list of competitors, has no idea who wrote what, and will happily tell you your homepage is a 7.
Asking the builder to grade its own hero section is like asking the compiler whether your code is beautiful. It compiled, didn't it? Ship it.
What "Power mode" meant
Replit Agent set to Power routed to frontier models — Opus, and Fable during the stretch when Fable was available. (If you were around for it: Fable existed, then didn't, then did again.) Use the best model your builder offers; this is not a place to economize.
To be clear: you don't need Fable. Opus did the overwhelming majority of this build and did it well. Fable happened to be in rotation for part of the project, and we'd struggle to point at a section of the site and tell you which one wrote it. If you're reading this and Fable is unavailable, or you simply prefer Opus, nothing about this method changes. Pick the strongest model your tool offers on the day and get on with it.
It's also a lesson in itself. Models churn — and not just on the builder side. We started these audits on GPT-5.5 and finished them on GPT-5.6 Sol, which shipped on 9 July, part-way through the project. We changed frontier models mid-run without changing a single thing about the method, and the scores kept climbing.
(OpenAI also cut GPT-5.6 pricing on 30 July, two days before our final audit. The audit half of this loop got cheaper while we were running it.)
The loop is the asset; the model is a commodity input. Anyone reproducing this a year from now will use different models and the same method.
The Replit plan that made it work
The $100/mo Replit plan was the real unlock, and not for the reason you'd guess. It allowed multiple concurrent tasks, each on its own git branch.
- Fire off several fixes at once, each as its own task.
- Each lands on its own branch.
- Review each branch independently.
- Merge the good ones into
main. - Publish from
main.
Parallel agents, a human review gate, no cross-contamination. If task 3 produced garbage, you don't merge task 3 — everything else still ships.
There's a second-order effect worth naming: because a bad attempt costs nothing to discard, you attempt more aggressive changes than you otherwise would. Isolation doesn't just prevent merge conflicts, it raises your risk tolerance.
What it cost — and why that's the real story
Total AI spend across the engagement was about $2,500 — and it's worth breaking that down, because the split surprised us.
Essentially all of it was builder spend. The auditing — the half of this method that produces most of the value — ran on a $20/month ChatGPT plan. No API billing, no per-audit cost, no metering. We ran cold deep-dive audits several times a day for three weeks and the marginal cost of each one was zero.
Sit with that for a second. The expensive part was making the changes. The part that decided which changes were worth making cost twenty dollars a month.
Which means the cheap version of this is genuinely cheap. Claude Code on a subscription handles the build half fine and is natively branch-and-commit shaped, so the review-and-merge discipline survives intact. Pair it with the same $20 ChatGPT plan and you have the entire method for a couple of hundred a month rather than $2,500. The Replit premium bought parallelism and a task UI — not a better outcome.
Now put that against the benchmark. A site like Dr. Sadati's probably cost $20k+ the traditional way, and that's likely just the build — a site at that level implies years of ongoing spend behind it: SEO retainers, agency fees, professional photography and video, content production. He earned the #2 position; the money wasn't wasted.
The point isn't a head-to-head. It's this:
That level of website is now achievable for a fraction of the price.
The bar used to require a $20k+ build plus years of sustained investment. It doesn't anymore. That outcome has dropped by roughly an order of magnitude in cost, putting it within reach of practices and small businesses that could never have justified the traditional number.
What money bought him that AI can't buy you
Look back at the scorecard. The client lost exactly two categories: reviews and authority (9.7 vs 10.0) and luxury presentation (9.7 vs 10.0).
That's not a coincidence, and it's the honest boundary of this whole method. Sadati's standing isn't only web spend — it's hundreds of reviews and a deep before-and-after library, accumulated over years of practice.
- AI closes the design, structure, copy, technical, and presentation gap fast and cheap.
- AI does not manufacture hundreds of genuine reviews or a decade of real before-and-afters.
- What it can do is ensure every piece of real proof you actually have is surfaced, organized, cross-linked, and working as hard as it possibly can.
The scorecard corroborates the caveat instead of contradicting it. A clean sweep would have been less believable.
The dimensions
These were never a fixed list. Every cold audit invented its own rubric, and we cover that at length further down — it's the biggest caveat in this whole method. What follows is the set of dimensions that recurred most often across many runs, not a checklist the auditor was working from:
- Pricing consistency
- Homepage overall
- Hero
- Brand distinction
- Treatment clarity
- Before-and-after proof
- Reviews and authority
- Conversion / booking
- Local SEO relevance
- Technical structure
- Luxury presentation
- Mobile optimization
Some audits scored nine of these. Some scored fourteen, including several that never appeared again. Treat the list as a description of what the model tends to care about, not a specification.
Retune for your vertical, but notice the shape: roughly one-third persuasion (hero, brand distinction, luxury presentation), one-third proof (before-and-after, reviews and authority, treatment clarity), one-third mechanics (technical structure, local SEO, mobile, booking, pricing consistency).
A site that's a 10 on mechanics and a 6 on proof ranks and doesn't convert. Scoring the three families separately stops you over-investing in whichever one you personally enjoy.

One gap, stated honestly: there's no accessibility dimension. For a medical-adjacent business that's both an ethical gap and real legal exposure — ADA and WCAG demand letters are common in this sector. If you're building your own rubric, make it thirteen.
Score against named competitors, not against the void
"Rate my website out of 10" produces a meaningless number. "Rate my website out of 10 against these two specific competitors, on these twelve dimensions" produces a number that moves when you do work. Here the benchmarks were Dr. Kevin Sadati and Forever Ageless — the sites actually competing for the same searches, in the same few-mile radius.
Run two scores when the comparison isn't apples to apples
Both benchmarks are physician-led. Our client is non-surgical. On a naive comparison that gap is unclosable — no amount of web design puts an MD behind a business.
So we tracked two numbers, and the data justifies it emphatically. Overall, the client leads by 0.2. In the non-surgical facelift category, the same audit separates it from Sadati by 1.1 points — 9.9 to 8.8.
The segment score isn't a consolation prize. It's where the advantage is actually visible. If your competitive set includes someone with a structural advantage you can't replicate, split the score — otherwise you'll spend weeks trying to fix a business fact with CSS.
The loop
Cold deep-dive audit → 1–10 scores + specific findings
↓
Convert findings into build prompts
↓
Agent works each fix on its own branch
↓
Review → merge to main → publish
↓
Re-audit with NO CACHE + live browsing to verify it's live
↓
Repeat until the audit runs out of things to fixWhat the prompt actually looked like
No prompt engineering. No persona, no XML tags, no seventeen-line system preamble. This is close to verbatim:
do a deep dive of a med spa newportbeachskintightening.com, dr sadati, and forever ageless, do a deep, deep dive, from 1-10 side by side
That's it. And when the session needed pushing:
…no cache
…use browser skill
Four things in that scruffy little prompt are doing real work, and they're worth separating out because they're the transferable part:
- Naming all three sites in one request. This is what produces a single comparison judged by one rubric in one sitting — the only kind of comparison that survives, as covered above. Auditing your own site alone gives you a number with nothing to lean on.
- "side by side." The load-bearing phrase. It forces a table rather than three paragraphs of prose, and a table forces the model to commit to a value in every cell instead of writing around the ones it's unsure about.
- "from 1-10." Forces numbers. Without it you get adjectives, and adjectives don't trend.
- "deep, deep dive." Saying it twice genuinely produces a longer, more thorough pass. Not elegant. Works anyway.
And the two afterthoughts — no cache, use browser skill — are what stand between an audit of your website and an audit of the model's memory of your website. Add them every time. Just don't mistake them for a guarantee: they improve the odds of a fresh retrieval rather than ensuring one, which is why the caching trap below insists you confirm with a canary change.
Dictated, not typed — and unfailingly polite
Worth knowing where that phrasing comes from: the client dictated these prompts with speech-to-text. That's why they read the way they do — run-on, lightly punctuated, the domain occasionally mangled by the transcriber. Nobody sat composing them.
And they were courteous to the machine. A typical prompt didn't stop at the instruction; it asked nicely and thanked it afterwards:
…please do a deep dive audit, please do a good job
thank you
We have no evidence the pleading improved a single score. There's some published work suggesting politeness nudges model output, and plenty suggesting the effect is noise. We weren't running a controlled trial on manners, so we can't tell you.
What we can tell you is the part that isn't superstition: the person who got the most out of this workflow was not a prompt engineer. He talked to it. He asked for a deep dive, asked it to do a good job, said thanks, and then — this is the bit that actually mattered — ran it again tomorrow. No syntax, no framework, no template library.
So if you've been putting this off because you assumed there was a craft to learn first: there isn't much. Say what you want, name your competitors, ask for numbers, and be relentless about repeating it. Being nice to the robot is optional, and cheap enough that you may as well.
The reason to show you something this unpolished is that the sophistication in this method isn't in the prompt. It's in running the prompt a hundred times, from fresh sessions, and being disciplined about what you do with the answers.
We ran this multiple times a day for three weeks straight — writing prompts, shipping fixes, re-auditing. Not weekly. Multiple times a day. Audits are cheap and the compounding is real. You can watch it in the timestamps on our own audit captures: 20 July the client trailed Sadati in most categories. 27 July it was at 9.3 overall. 1 August it was at 9.8 and ranked #1.

Nobody tells you how much of this is waiting
Here's the part the demo videos skip: this loop is slow.
A single agent task can sit and think for five minutes before it writes a line. Publishing can take another five. So the actual rhythm of a "fast AI workflow" is: write a careful prompt, wait, review a diff, merge, wait again to deploy, then ask a second AI to go look at it — which often takes a couple of minutes itself. An instant response is a warning sign, though freshness is only really settled by a canary change, as covered later.
The speed is inconsistent, which is its own problem. Sometimes a task lands in under a minute. Sometimes the identical kind of task takes ten. It depends on load, on model, on how much the agent decides to reason, and on nothing you can see or control. You cannot plan a day around an average when the variance is that wide — so you plan around the slow case and treat the fast ones as a bonus.
Call it fifteen-ish minutes per round trip for one fix, on a normal day.
Which is the actual reason this took three weeks rather than three days. Not because the work was hard — most individual fixes were small — but because the cycle time is long and there were a lot of cycles. The parallel-branch setup exists precisely to amortize that: if you're going to wait five minutes anyway, wait five minutes for six things at once.
Plan your day around it. The mistake is sitting there watching the spinner. The move is to keep four or five tasks in flight and treat each one's completion as an interrupt. AI made the work cheaper and better here; it did not make any individual step quick.
AI is fast the way a bus is fast: excellent top speed, and it stops constantly.
Keep a scorecard. Every run, every dimension, dated. That table makes the trend visible, makes an outlier run obvious, and makes a regression show up as a drop in one column instead of a vague sense that something got worse. It's also the single best client deliverable this process produces.
The traps
Everything above is the clean version. Here's what actually goes wrong.
1. Caching will silently invalidate your audits
The big one. Your audit session must use no-cache and live browsing. Otherwise it reads a cached copy and delivers a confident, detailed, obsolete audit — and you spend the day fixing things you fixed last week.
There are only two hard problems in computer science: cache invalidation, naming things, and persuading a language model that it is looking at a cached copy of your website.
Timing is a hint, not proof. A real fetch takes a noticeable amount of time, and an audit
that comes back instantly probably didn't retrieve anything. That heuristic served us well and
we still use it — but be clear about what it is. Saying no cache in a prompt is a request,
not a guarantee: it doesn't prove a Cache-Control header went anywhere, and it certainly
doesn't prove every intermediary cache and the model's own recent context were bypassed. Fast
responses are suggestive. They aren't evidence.
So verify with a canary instead. Publish a change only you know about — a new heading, a changed price, a fresh testimonial — then ask the auditor to tell you what it says. If it can quote the thing you shipped twenty minutes ago, it genuinely fetched the current page. If it can't, nothing else it tells you is worth reading.
That's the reliable check, and it costs one question. Ask for live browsing, ask it not to rely on cached knowledge, watch the timing as a soft signal — and then confirm with something only a fresh retrieval could know.
Press it and it will often admit the lapse — "you're right, I didn't do a live check." Useful, though also worth noting that a model agreeing it made a mistake isn't independent confirmation either; it's an agreeable machine agreeing with you. The canary is the part that actually settles it.
If that feels like an excessive amount of suspicion to direct at your own tooling: this article spends ten thousand words arguing you shouldn't take an AI's scoring at face value. It would be strange to then take its browsing on faith.
2. Sessions started on mobile often can't browse
Some sessions don't have live browsing available, or can't honor no-cache. In our experience this hit sessions started on mobile especially. Start audits on desktop and verify browsing works before trusting the first score.
Auditing a website from a phone session with no browser is a bit like proofreading through a mail slot. Confident, thorough, and describing a different document entirely.
3. The auditor's memory works against you
ChatGPT carries context between sessions. It remembers what your site used to be — and it would score a dimension lower for an inconsistency already fixed, because the stale version was in memory. The audit was grading a ghost.
The fix is blunt: explicitly tell it to update its memory. Don't assume a live fetch overrides a stale memory. Say the words.
4. A warm 10 is not a cold 10
Two different instruments, and conflating them will convince you you're done when you aren't:
- Warm re-check — "you flagged X, we fixed it, look again with no cache." Finds the fix, often hands out a 10. This verifies.
- Cold deep-dive — fresh session, no history, full audit from scratch. This discovers, and will surface a new batch the warm re-check never looked for.
Warm says 10, cold says "here are six more things." Both correct. Use warm re-checks to confirm deploys, cold audits to decide whether you're finished. You're only done when a cold audit comes up empty.
5. Audits vary — until you've run enough of them
This one we can show you. Here's a run from 20 July:

Note the categories: Owner biography, Case-study information, Video testimonials, Services and treatment pages, SEO and search depth, Authority and credentials. Now compare to the 1 August scorecard at the top of this article: Homepage overall, Hero / first screen, Brand distinction, Treatment clarity, Before-and-after proof, Reviews and authority, Conversion / booking, Local SEO relevance, Technical structure, Luxury presentation.
Different category names. Different counts. Different column order — Sadati is scored first in July, the client first in August. Same site, same auditor, same instruction.
Side by side, the drift is stark:
| 20 July rubric | 1 August rubric |
|---|---|
| Homepage | Homepage overall |
| Hero section | Hero / first screen |
| Owner biography | (folded into Reviews and authority) |
| Before-and-after gallery | Before-and-after proof |
| Case-study information | (gone) |
| Video testimonials | (gone) |
| Services and treatment pages | Treatment clarity |
| SEO and search depth | Local SEO relevance |
| Authority and credentials | Reviews and authority |
| (not scored) | Brand distinction |
| (not scored) | Conversion / booking |
| (not scored) | Technical structure |
| (not scored) | Luxury presentation |
This is the single most important caveat in the article, so we'll say it plainly: a fresh cold audit does not necessarily grade the same things as the last one.
Four categories that existed in July were gone by August. Four categories that mattered in August didn't exist in July. Two were renamed into near-synonyms. Nobody asked for any of this — we gave the same instruction both times.
The consequences are practical and they change how you're allowed to use the numbers:
- You cannot diff two cold audits category by category. "Video testimonials went from 9.5 to — " has no ending, because the second audit never scored video testimonials.
- A category appearing for the first time is not a regression. The first time "Technical structure" showed up we briefly thought something had broken. Nothing had broken; the rubric had simply grown a new opinion.
- Only two things survive across runs: the within-run ranking (all three sites judged by the same rubric in the same session) and the overall trend across many runs. Everything else is comparing apples to a slightly different fruit each week.
- If you want a stable rubric, you have to pin it yourself — paste the twelve dimensions into every fresh session rather than asking for "a deep dive audit" and accepting whatever categories come back. We eventually did this. It's the difference between a measurement and a conversation.
The upside: because the auditor keeps reinventing its own criteria, it keeps noticing things a fixed checklist would have stopped looking for. The drift is annoying and it is also where several of the best findings came from.
The drift settles down on its own
Here's the part that makes this liveable, and we only noticed it after weeks of runs: eventually the auditor starts reusing its own rubric.
Once enough context has accumulated — enough previous audits of the same three sites, in the same account, with the same recurring instruction — it stops inventing a fresh set of categories each time and starts scoring against roughly the same list it used last time. The rubric converges without being asked.
And you don't have to wait for it. You can simply tell it to use the same rubric as last time, and it will. Naming the previous categories, or just saying score it on the same dimensions as the last audit, is enough to lock the shape while leaving the judgment fresh.
Notice what that means: this is the same memory that was working against us three traps ago. When it remembered a stale version of the site, memory was the enemy and we had to explicitly tell it what had changed. When it remembers the rubric, memory is the whole point — it's what turns a series of unrelated opinions into something you can actually chart.
That gives you three usable positions, and they're genuinely different tools:
- Fully cold, no pinning. Maximum discovery, minimum comparability. Best early, when you still want it finding categories you hadn't thought of.
- Let context accumulate. The rubric stabilises by itself over repeated runs. This is where we ended up, and it's the least effort.
- Explicitly pin the rubric. Paste the dimensions, or tell it to reuse the last set. Most comparable, and the right call once you're tracking a number rather than hunting for problems.
Early on you want drift. Later you want consistency. The mistake is wanting both at once and then being confused by your own scorecard.
Two responses:
- Volume standardizes it. After enough runs the structure and scoring converge. Early scores are noisy; the trend is the signal.
- Don't argue with a bad run. If a score contradicted an established trend, we opened a new session and ran it again. Cheaper than negotiating with a model about its own output, and it doesn't pollute the session with your objection.
6. Non-visual work needs explicit verification
Technical SEO — JSON-LD structured data especially — has no visual signal. A browsing model
looking at a rendered page will not notice you added LocalBusiness schema, and will score
"technical structure" off whatever it can see.
Instruct it explicitly: go find the structured data, tell me what types are present, confirm the fields. Otherwise your best technical work is invisible to the grader.
7. Content changes cause regressions
We had it near-perfect. Then the client changed the pricing packages and added a new one. It didn't propagate everywhere. Pricing said different things on different pages, and the score dropped.
One price achieved independence somewhere between the brief and production and started quoting its own numbers. We found it eventually. It had been living on the pricing page under an assumed value.
This is why pricing consistency is a standing dimension, not a one-time task. Any client edit — pricing, hours, services, staff — can regress a finished site. The loop isn't a project you complete. It's a monitor you keep running.

8. The auditor is biased toward finding something
The flip side of trap 4: a model asked to audit will produce findings whether or not real problems remain. That's the job you gave it. As you approach 10, a growing share of findings are manufactured nitpicks rather than genuine defects.
So "it ran out of things to fix" is a judgment call you make, not a state the model announces. The signal to watch: findings stop being specific and start being generic advice.
9. You need a human veto
The auditor will sometimes recommend things that are right for the rubric and wrong for the business — publish full pricing when the business deliberately gates it, soften the exact differentiator that's the wedge, add a section that dilutes a focused page. Implement findings selectively. The branch-per-task workflow exists partly to make rejecting a change free.
10. Push back — it is not always right
The auditor states everything with equal confidence, including the parts it's wrong about. If something doesn't make sense, challenge it.
Interrogate every score movement. When a number dropped or jumped, we pressed on why. Sometimes there was a real cause. Sometimes there wasn't, and it folded.
Be most suspicious of good news. There was a run where it handed out 10/10 immediately after we told it a fix had been made. That's the model being agreeable, not the site being perfect — a warm session over-crediting a change you just announced to it.
It's the same reflex as the caching check: if the answer arrives too easily, verify it. A warm session that congratulates you right after you announce a fix is not evidence. A cold session that can't find anything is.
The single most useful prompt
When a category scored low, we stopped reading the finding and started asking:
What would make this a 10?
This converts a score into a work order. Instead of "before-and-after proof: 8.5," you get a specific list of what's missing and what would close the gap. The number stops being a grade and becomes a spec.
Do it for every dimension below target. It's the highest-leverage habit in the whole method, and it's the reason the loop terminates at all — you're not guessing what the auditor wants, you're asking it to write the ticket.
What the auditor caught that we didn't expect
It worked at three quite different altitudes, which is unusual for a single tool.
Hard technical facts. It pulled real Google Search Console data and found broken and unindexed URLs:

- 95 pages crawled but not indexed - 8 pages discovered but not indexed - 32 not-found URLs - 7 URLs blocked by robots.txt - 2 noindex URLs - 2 redirects
That's the "technical structure" dimension with receipts, not vibes.
Domain relevance. It identified local SEO gaps — what was thin or missing for the geographic target.
Subjective composition. It flagged homepage layout proportions — whether the page was too long or too big. This is the one worth pausing on, because it's a judgment about pacing and density, not a rule check. A homepage can pass every technical audit and still be exhausting to scroll. Most tools do exactly one of these three things. It did all three.
What actually moved the score
The AI proposed the sections, not just the copy
We didn't hand it a wireframe. The agent came up with the homepage section ideas itself: comparison reviews, a "who this treatment is for" qualification section, and a dedicated reviews section.

That's a different and more valuable mode than "make this div prettier." It was identifying missing page architecture and proposing the section to fill the gap. If you're only asking AI for styling and copy, you're using maybe a third of it.

Marketing verbiage → compliant medical verbiage
It also rewrote marketing language into compliant medical language. In aesthetics this isn't a tone preference — language that overpromises outcomes, or makes claims a non-physician practice isn't permitted to make, is genuine liability.
You can see the result in the disclaimers throughout the finished site. Under the candidacy criteria above: "Women outside these age ranges may also be appropriate candidates. Treatment suitability and results vary. A consultation is required." Under the press mentions:

General lesson: in regulated verticals, put compliance language review into the loop deliberately. Cheap to ask for, expensive to miss.
Authority: give the person a background
We rebuilt the About section around credentials, certificates, and a background timeline — not a paragraph of adjectives, a verifiable trajectory.

This is worth more than the score movement suggests. Medical aesthetics is squarely YMYL territory, where Google leans hardest on E-E-A-T signals. The auditor's rubric and Google's quality guidelines are pointed at the same thing, which is a large part of why the rubric works at all.

Proof: stack it and cross-link it
We added Google reviews, Yelp reviews, before-and-after galleries, and video testimonials — then cross-linked them to each other. This is the single best illustration of the idea:

The cross-linking mattered more than the individual additions. Proof sitting in isolated silos reads as decoration. Proof that references itself — a before-and-after linking to the video testimonial from the same patient, linking to their Google review, linking to the original Instagram post — reads as a verifiable trail. The audit noticed.

Authenticity beats polish
The homepage originally led with glamour shots: beautiful, professional, and interchangeable with every other site in the category. We replaced them with the client's real, non-AI before-and-after videos.

This is the counterintuitive lesson of the entire project. In a category where every competitor can now generate flawless imagery, flawless imagery is worthless and evidence is priceless. The AI-generated-everything era has quietly made "obviously real" a competitive advantage.
Page speed, and the fight it picked with everything else
We optimized performance on both mobile and desktop — not desktop only — and landed near 100 on both. It took several rounds: optimize, re-measure, feed the numbers back, go again.
Worth knowing: ChatGPT can run PageSpeed from inside the chat session. You don't have to leave the audit, run Lighthouse separately, and paste results back. That keeps the whole optimize → measure → optimize cycle in one place, which is most of why running it several times was practical rather than tedious.

There's a reason it wasn't a one-prompt fix, and it's the most interesting engineering problem in the project: we had just added a lot of video. Video testimonials, before-and-after videos, clips pulled from social. Video is the heaviest thing you can put on a page, and "near-100 on mobile" and "rich video proof everywhere" pull directly against each other.
This is the one place where two dimensions genuinely conflict, and resolving it is real work rather than prompting: lazy-load anything below the fold, poster images instead of autoplay, facade pattern for embeds, modern codecs, correct sizing per breakpoint, and never let a testimonial carousel block first paint.
The one place we can show you a real measurement
Everything above this point is a model's judgment. Page speed isn't. So here are all three sites run through Google PageSpeed Insights on mobile, the same day this article was written.
Mobile only, deliberately. Desktop isn't where the difficulty lives — everyone scores well on desktop, including both competitors, so a desktop table proves nothing about anyone. Mobile is the constrained test: slower simulated CPU, throttled network, smaller viewport, and it's where most of this client's traffic actually arrives. If you only publish one number, publish the mobile one. It's the one that was hard to earn.
| Mobile (PSI) | NB Skin Tightening | Dr. Kevin Sadati | Forever Ageless |
|---|---|---|---|
| Performance | 99 | 64–76 | 65–67 |
| Accessibility | 97 | 97–100 | 100 |
| Best Practices | 100 | 100 | 100 |
| SEO | 100 | 100 | 100 |



A 99 against a 64 and a 65 is a bigger gap than anything on the ChatGPT scorecard — and unlike those numbers, this one is reproducible by anyone reading this article. Go run it yourself.
Note the ranges in that table, though. We ran the competitors twice and got 76 then 64, and 67 then 65. Lighthouse lab scores drift between runs — network conditions, ad and tracker loading, server variance. So even the "objective" measurement has a confidence interval, and the honest read is "~99 versus mid-60s-to-70s," not "99 versus 64."
Which is a useful corrective to the whole article: the AI scores vary between runs, and so does the real instrument. The difference is that PageSpeed tells you its methodology and anyone can re-run it. That's the standard the model scores don't meet — and it's why the trend, not the reading, is what you should trust in both cases.
And now the part that complicates the win
Look closely at the Sadati screenshot above. It has a section the other two don't:
Core Web Vitals Assessment: Passed — LCP 1.6s · INP 80ms · CLS 0 · FCP 1.4s
That's Chrome User Experience Report field data — measurements from real people on real devices over a 28-day window. His site scores 64 in the lab and passes Core Web Vitals with actual users.
Our client's report, and Forever Ageless's, both say "No Data" in that section.
There are two things worth sitting with here.
First, lab score is not user experience. A 99 in a simulated test on a datacenter connection is a good sign, not a result. Sadati's mid-60s lab score with passing field data arguably describes a faster real-world site than a 99 with no field data at all. We can't claim otherwise, so we won't.
Second, "No Data" is a function of age, not quality. CrUX only reports once a site has accumulated roughly 28 days of real traffic at sufficient volume. Sadati's site has been around for years and clears that bar easily. Our client's site is substantially newer and rebuilt very recently, so there is simply no sample yet — and there couldn't be, no matter how good the page is.
This is worth being precise about, because it's easy to garble: the lab score is the thing you control and we have it (99). The field data is the thing that accrues, and it can't be built, prompted, or bought — only waited for. A brand-new site showing "No Data" isn't failing a test. It hasn't been given the test yet.
It is, though, the same moat surfacing in a third independent place. Reviews, before-and-afters, and now real-user performance history: all three are time-denominated. That's the consistent shape of what AI couldn't close.
So the honest version of this comparison: we built a technically faster page, and he has a more established website. Those are different achievements, and only one of them was ever available to us in three weeks. Check back in a year and the field data column will exist — that one just needs the clock to run.
The duplicate-build incident
There was a second cause, and it's the best single argument for this entire method.
To handle mobile and desktop, the builder had quietly produced two separate versions of the page — a mobile one and a desktop one, both shipped in the markup. It looked fine. It worked.
The auditor caught it and wanted it consolidated into one responsive version with the DOM element count reduced.
Duplicating markup is the path of least resistance for an agent: it's the fastest way to make both viewports look right. But it doubles the DOM (excessive DOM size is a direct Lighthouse flag), tanks mobile performance, and leaves you maintaining two copies of every future content change — a consistency regression waiting to happen, on a site whose rubric grades consistency.
The agent's answer to responsive design was to build the site twice. Somewhere out there, a perfectly good media query is still sitting by the phone waiting for a call that never came.
Sit with the shape of that. The builder created the problem and would never have flagged it — its own work looked correct. The auditor, which knew nothing about how the page was built and saw only the result, caught it immediately.
That's the separation-of-duties argument in one concrete incident. Not a philosophical preference — a real defect that existed because one model built it, and got found because a different model looked at it.
SEO and local pages
Alongside general SEO we added more local SEO pages, expanding location and service coverage rather than trying to rank one page for everything. That's what moved "local SEO relevance," and it fits the business: this isn't a national play, it's a few-mile-radius play against two competitors in the same city.
One caution, because AI makes the wrong version effortless. Mass-produced local pages that differ only by city name are doorway pages, and Google penalizes them. The version that works has genuinely distinct content per page — local landmarks, location-specific pricing, proof from that area. It's trivially easy to generate hundreds of the bad kind in an afternoon. Don't.

The media work nobody expects an agent to do
The agent could look at raw video and identify the shots and who was in each one. From there we directed edits by chat prompt:
- "Make the before and after eye-level and head position consistent."
- "Move the person's head higher in the frame."
It also pulled videos from social media and reformatted them for web — aspect ratio, encoding, sizing. No video editor in the loop.
Transcribe every video — it's free SEO
The one piece of this we'd tell everyone to do regardless of the rest of the method: run local Whisper over every video on the site. Replit's agent can do it, Claude Code can do it, and it costs nothing — the model runs on your own machine, no API, no per-minute billing.
Why it matters: a video is invisible to a search engine. You can have the most persuasive
patient testimonial in the county and, as far as Google is concerned, that page contains a
<video> tag and some whitespace. A transcript turns every spoken sentence into indexable text
— in the patient's own words, using the phrases real people actually say about the treatment,
which is exactly the long-tail vocabulary you'd otherwise pay a copywriter to guess at.
Look back at the testimonials figure above and you'll see the result: every card carries a "Read video transcript" link. Those are Whisper output, lightly cleaned up.
Three things you get from one cheap step:
- Indexable content where there was none, in natural patient language.
- Accessibility — captions and transcripts are the single most-cited WCAG gap on video-heavy sites, and remember there's no accessibility dimension in our rubric to catch it.
VideoObjectstructured data with a realtranscriptfield, which feeds the technical structure dimension.
It's the highest ratio of SEO benefit to effort anywhere in this project, and the reason it gets skipped is simply that transcribing video used to be tedious and expensive. It isn't either any more.
And then there's the benefit nobody expects. Once the video is text, the auditor can read it — which means video stops being a black box in your content and becomes something an AI can check like any other page element.
Specifically, it can tell you whether a video is actually aligned with the marketing message around it. That sounds abstract until you hit a real case, and they're common:
- A testimonial praises a treatment the page isn't selling, or names an older service you renamed six months ago.
- The page promises "no downtime" and the patient cheerfully mentions being red for two days.
- A client says something warm and enthusiastic that, written down in plain text, is an outcome claim a non-physician practice isn't allowed to make — the exact compliance problem covered earlier, hiding in a video where no one thought to look for it.
- The hero video's emotional register is completely different from the headline sitting next to it.
Every one of those is invisible while the content is locked inside an MP4. A human would have to sit and watch every video with the page copy open beside them. Nobody does that, which is why misaligned video sits on sites for years.
Transcribe first, and the auditor does that pass for free — on every video, every time you run it, forever.
Same story on stills. It cleaned up scanned certificates that came in skewed, yellowed and low-contrast — because having the credential is worthless if the scan looks like a fax — and removed an unwanted person from a photo that was otherwise usable — full Back to the Future treatment, except the erasure took about nine seconds and nobody's hand started fading out. One prompt, and a bystander who'd been standing in the shot since 2019 simply stopped having been there. The photograph now shows a past that never happened, which is either a triumph of accessible image editing or a small crime against the historical record, depending how strongly you feel about a stranger's elbow.
The broader point: the asset pipeline is part of the loop, not a prerequisite to it. The usual blocker on "add credentials and real proof" is that the raw material is a shoebox of bad scans and awkward photos, and cleaning it up traditionally means a designer and a budget. That blocker is mostly gone.
Between the video edits, the photo retouching, the copywriting and the front-end build, this project quietly did the work of about four specialists. Everyone asks whether AI is coming for our jobs. On this evidence: yes, obviously — except for the one job it still can't do, which is sitting there at 2am telling it that no, that's still wrong, do it again. Job security, of a sort. Not the sort anyone asked for.
One line worth holding, though: retouching presentation — deskew, contrast, remove a bystander — is fine. Retouching before-and-after results is not. It destroys the authenticity advantage the site was built on, and in medical aesthetics it's a regulatory problem. Fix the photo, never fix the result.
The part of this that isn't about AI at all
Here's the caveat that decides whether any of the above will work for you, and it has nothing to do with models, prompts or rubrics.
We got to the top of that scorecard because the client did the work.
He had real before-and-after photography. He had video of actual patients willing to appear on camera and say their names. He had certificates, a training history, a coherent professional timeline. When we asked for raw material, it arrived. When we asked him to film something, he filmed it. When we said the glamour shots had to go and be replaced with unretouched results, he agreed — which is not a small thing to agree to.
Most clients won't. In our experience the two most common walls are:
- "I don't think that content is good enough to put on the site." People are strange about their own proof. A practitioner will sit on a folder of genuinely persuasive before-and-afters because the lighting isn't perfect, then ask why the site doesn't convert. The best material is almost always already on someone's phone.
- Simply not doing it. Gathering assets is homework, and homework competes with running a business. A project can stall for a month waiting on twelve photos.
And then there's the review problem
Separately, and more often than you'd think: the client may not have review collection properly set up at all. No claimed Google Business Profile, no Yelp page under their control, no process for asking a happy client to leave a review, reviews scattered across a personal profile and a business one. Sometimes there's a Google listing nobody has ever logged into.
You cannot surface social proof that doesn't exist, and you cannot cross-link reviews that were never collected. Look back at the final scorecard: the client scored 9.7 on reviews and authority — a strong number, built on a real 4.9 from 65 Google reviews. With no review infrastructure that isn't a 9.7. It's a 5, and no amount of prompting moves it.
What this means practically
Before you quote this method to anyone, audit the client, not just the site:
- Do they have real before-and-afters, case studies, or results they're willing to publish?
- Will anyone go on camera?
- Is the Google Business Profile claimed, correct, and collecting reviews?
- Is Yelp claimed? Is anyone asking for reviews at all?
- Are there credentials, certificates, or a history worth documenting?
Those answers set the ceiling before you write a single prompt. AI can take a site with good raw material and make it excellent. It cannot conjure the raw material. A client with nothing to show gets a beautifully built, technically flawless, entirely unconvincing website — which will score well on mechanics and lose on everything that makes someone book.
The honest framing for the whole method: this is a multiplier on what a business already has. Multipliers are wonderful and they do not work on zero.
The playbook, condensed
- Separate the auditor from the builder. Non-negotiable.
- Pick 8–13 scoring dimensions for your vertical. Persuasion, proof, mechanics — and accessibility.
- Name your real competitors. Score against them, not against nothing.
- Split the score if a competitor has a structural advantage you can't replicate.
- Always request live browsing and a fresh retrieval. A fast response is a warning sign, not proof — verify freshness with a recent canary change.
- Start sessions on desktop. Verify browsing works before trusting a score.
- Tell the auditor to update its memory whenever something significant changes.
- Warm re-checks verify deploys. Cold audits decide if you're done.
- Explicitly ask it to verify non-visual work — schema, meta, structured data.
- One fix per branch. Review, merge, publish.
- Re-run a bad audit in a fresh session instead of arguing with it.
- Keep a dated scorecard of every run across every dimension.
- Apply a human veto. Rubric-right and business-right aren't the same thing.
- Keep running it after you hit the top. Client edits regress.
Expect a non-linear curve
8.4 → ~9.5 is the fun part: add the missing sections, fix the hero, ship the schema, stack the proof. The last half point is almost entirely consistency and polish — pricing matching everywhere, framing consistent across before-and-afters, compliant verbiage on every page. Slower, less satisfying, and the stretch where most people quit.
This is a retainer, not a project
The pricing regression proves it. Any client content edit can regress a finished site. That makes this a recurring service — scheduled cold audit, scorecard delivered, regressions fixed. For an agency, that's the actual business model hiding in this post.
What a "10" actually is
We flagged at the top that the title is hook bait. Here's the fuller version of why that matters, now that you've seen the method.
A 10/10 website is just what the AI thinks a 10 is. It's an opinion, formed from the model's training and knowledge — not an objective measurement, not a law of physics. No external authority certifies it. We invented the rubric; the model applied it. And as the two rubrics side by side in this article show, it doesn't even apply it identically twice.
But it does a genuinely good job rationalizing the scores. It explains why a dimension is a 9.7 and what specifically would take it to a 10. The reasoning is legible, specific, and actionable — which is exactly what makes the loop work at all.
The score is an opinion. The reasoning behind the score is the deliverable.
You're not buying a certified grade. You're buying a structured, consistent, competitor-aware critique that tells you what to do next — plus a number whose value is that it moves in the right direction when you do the work.
That's also why the earlier advice holds: volume standardizes it, because enough runs make the opinion converge. And cold-vs-warm matters, because an opinion is only ever as good as the information it was formed on.
How we knew we were done
We kept asking the same question after every audit: what would make this a 10?
For three weeks, the answer was a list of fixes.
Now the answer contains no fixes at all. Nothing about the site. The only things left are:
- More reviews
- More before-and-afters
- Standardized before-and-afters — consistent framing, angle, lighting and eye level across the whole library, so the set reads as a controlled series rather than assorted snapshots
That's the real ending, and it closes the loop on the caveat from the top of this article. AI closed every gap it could close. What's left is precisely the stuff it cannot manufacture: accumulated real reviews, and a deep library of real results.
It also matches the scorecard exactly. The only two categories the client didn't win were reviews and authority and luxury presentation — the two things you earn over years rather than build in three weeks. The auditor and the endpoint agree, which is the most convincing thing about either of them.
Note that one of those three items isn't an accumulation problem. Standardizing the framing on the before-and-afters you already have is craft work — exactly the tedious, high-impact, low-creativity task the agents turned out to be good at. So the remaining list is two things the business has to earn, and one thing we can still go do.

The website is finished. The business isn't, and never will be. The score isn't the product. The loop is. Ours is still running.
Reproduce it yourself
We'd rather you checked this than took our word for it, so here is the evidence, including the parts that are incomplete.
What we're publishing
| Artefact | |
|---|---|
| Final scorecard, 1 Aug 2026 | scorecard-2026-08-01.csv |
| Earlier rubric, 20 Jul 2026 | scorecard-2026-07-20.csv |
| PageSpeed runs, 1 Aug 2026 | pagespeed-2026-08-01.csv |
| Raw audit screenshots | the four unretouched captures shown throughout this article |
| Site screenshots | every figure here is an unedited capture from the live site, 1 Aug 2026 |
The audit prompt, verbatim, dictated by the client:
do a deep dive of a med spa newportbeachskintightening.com, dr sadati, and forever ageless, do a deep, deep dive, from 1-10 side by side
…with no cache and use browser skill appended when the session needed pushing.
The subjects
| Client | newportbeachskintightening.com |
| Benchmark 1 | drkevinsadati.com — physician-led |
| Benchmark 2 | foreveragelessinc.com — physician-led |
| Audit dates shown | 20 Jul, 24 Jul, 27 Jul, 1 Aug 2026 |
| PageSpeed date | 1 Aug 2026, mobile strategy, two runs per competitor |
| Auditor | ChatGPT, $20/mo plan — GPT-5.5 at the start, GPT-5.6 Sol by the end |
| Builder | Replit Agent, Power mode (Opus; Fable in rotation for part of it) |
What we can't give you, and why
Being straight about the gaps, because a reproducibility section that only lists strengths isn't one:
- No raw chat exports. The audit outputs exist as photographs of the screen, not text dumps. We've transcribed them faithfully and published the screenshots unretouched, but you're trusting our transcription. Set up your own audits with exports if you want cleaner evidence than that.
- No "before" screenshots. We didn't capture the pre-project site, which is the single most annoying omission here — the glamour-shots-to-real-video swap is described but not shown. If you're running this, screenshot everything on day one.
- No per-change revision log. Branch names and merge order weren't kept in a form worth publishing.
- The 8.4 starting score comes from an early audit we didn't screenshot. Treat it as reported, not evidenced. The scored trajectory we can show starts at 9.3 on 27 July.
- PageSpeed is lab data. Explained at length above: the client site has no Chrome UX Report field data because it's too new to have accumulated any, while Sadati's site does and passes.
The part that is fully reproducible right now: open PageSpeed Insights, run those three URLs on mobile, and compare. That's Google's instrument, not ours, and it will disagree with us by a few points because lab scores drift.
Go and try it on your own site
Genuinely — do. Open a fresh session, name two competitors you actually lose business to, ask for a side-by-side from 1 to 10, and read what comes back. It costs you twenty minutes and nothing else, and you'll learn more about your own site in that one audit than in a month of staring at analytics.
Fair warning: it's addictive. We didn't expect that part.
Something happens when a fuzzy worry — the homepage feels a bit flat — turns into "Hero: 9.2". Suddenly it's a number, and numbers want to go up. You fix the hero, re-run it, and it says 9.6. So you fix one more thing. Then you find yourself at eleven at night thinking the before-and-after section is probably a 9.4, and I know exactly what would make it a 9.7.
That's the loop's real trick. Not the intelligence — the scoreboard. Website work is usually a fog of opinions with no feedback until a quarter of traffic data eventually arrives. Give someone a number that moves the same day they do the work and the whole thing turns into a game they want to keep playing. Three weeks of daily audits didn't take discipline. It took the opposite: knowing when to stop.
Which is the one caution worth repeating. A scoreboard will happily keep you busy after the work stops paying. There's a difference between this fix wins bookings and this fix wins a tenth of a point, and around 9.7 they stop being the same thing. We kept going to the end because the remaining findings were still real. When they turned into polish for its own sake, we stopped — and the honest answer to "what would make this a 10" had become more reviews, which no amount of clicking refresh will produce.
Chase the score while it's pointing at real problems. Notice when it stops.
Want one of these?
If you've read this far, you already know the method isn't secret — it's written out above, and you're welcome to run it yourself. Most of it costs a $20 subscription and a lot of patience.
What it actually takes is three weeks of running the loop several times a day, the judgment to know which findings to ignore, and someone willing to keep pushing back on a confident machine at 2am. If you'd rather not do that part, that's what we're for.
If you want your own 10/10 website — get in touch. Tell us the two or three competitors you want to be measured against, and we'll run the first cold audit on your current site so you can see exactly where you stand before committing to anything.
You'll get the same thing this article is built on: a scorecard, a list of specific fixes, and an honest read on which of them are worth paying for.
Need Help With Your Website?
I fix these problems every day. Send me a message and I'll take a look.
Get Help Now