Give us the URL of a page you publish. We pull the checkable facts out of it, turn them into questions, and make an AI agent sit the same exam twice — closed book from memory alone, then open book with search and page fetching. Every answer is graded against your own words.
Same model, same questions, same phrasing, no system prompt in either. The only thing that changes is whether the tools are attached — which is the only way the gap between the papers means anything.
No search, no fetching. The model answers from whatever it absorbed about you during training — which may be a year stale, or drawn from a competitor's description of you.
This is what an agent falls back on when your page is slow, blocked, or simply not worth a tool call.
Search and page fetching available. The model may go and read your live page before it answers — if it thinks to look, if it finds you, and if it uses what it finds.
This is the realistic case, and the one most people assume is always happening.
Checked in order. A failure at any rung means the ones after it cannot be read, which is why the report tells you which one broke.
Does your page even come back when the model searches for the answer it contains?
If it never surfaces, nothing else on this page matters — the model is answering about you from somewhere else entirely.
Having found it, does the model actually read it — or answer from a stale memory of you?
Both look identical to the person asking. Only one survives you changing your pricing.
With your page in hand, does it get the answer right — or state something you never said?
Reading is not the same as understanding. A model can fetch the page and still misattribute a claim to you.
The same page, the same questions, a different model — and a different answer. One will fetch and get you right; another answers confidently from a year-old memory. You do not get to choose which one your customer asks.
So every run names its model, and runs are comparable across them: the question set is frozen per page, so the only thing that changed is the model.
Pick the model your customers actually use, or run several and compare.
Every checkable claim is extracted with the line it came from, then turned into a question with a known answer. The set is frozen, so two runs a month apart are comparable.
Each question goes to the model with search and fetch available, and again with nothing. Several samples per condition, because one answer is an anecdote.
A stronger model grades each answer against your page and writes down why. You get the outcome, the reasoning, the search results it saw, and the exact lines it should have used.
Pay what a run actually costs plus a flat margin — no seats, no subscription. Top up in $10 increments and draw it down; credit does not expire. A typical page is a couple of dollars.
Your first run is free, so you see a real report before deciding anything. Runs left public and searchable are 10% cheaper — the public index is what shows people this works.