Method · how the tests are run
How Makefield measures
This page sets out the rules a Makefield measurement follows. It writes the correct answer down before it asks an engine. It estimates how much answers move on their own before it calls any change a result. It scores each engine apart until the numbers mean the same thing on each. A field note on this site is one observation, not a measurement under these rules.
The loop these rules serve, step by step: ACE, explained.
The difference
An audit, or a measurement?
Anyone can ask ChatGPT about a company, take a screenshot, and call it an audit. The trouble is that an AI answer is a moving, noisy thing: it changes from one day to the next, it differs across engines, and two people reading the same answer will score it differently unless they agree in advance on what they're looking for. A screenshot doesn't survive any of that.
Makefield treats a reading as a small piece of research about a company — measured so the number means something, and written so someone else could check it. Here is what that involves, in plain terms.
Before we claim a win
We estimate the wobble first
AI answers drift on their own. Ask the same question twice and the wording, the names, even the recommendation can shift, with nothing changed on your side. So before we say a fix "moved the answer," we estimate how much the answer moves by itself — its noise floor. A change smaller than that wobble isn't a result; it's noise. Naming this floor is the quiet part most "audits" skip, and it's the thing that makes a later improvement believable.
Across engines
We don't blend engines until they mean the same thing
ChatGPT, Perplexity, Claude, Gemini, Copilot and Google AI Mode behave differently enough that a single blended score can hide more than it shows. Before we pool numbers across engines, we check that the measurement means the same thing on each one — and where it doesn't, we report them apart rather than averaging away the difference. A pooled figure you can't defend is worse than six honest ones.
So another person would agree
We score against a rubric, not a hunch
Every dimension is scored on a fixed, calibrated scale, applied to the actual answer with the company's name masked so brand familiarity can't tilt the rating. For a study, we measure how well independent scorers agree (a real agreement statistic, disclosed with its result) rather than asking you to trust one person's read. The point isn't to sound technical; it's that a reliable number is one that doesn't depend on who held the pen.
Proof, not assertion
We show the change moved, and how much
Measurement doesn't end when someone ships fixes. A later re-run of the same baseline shows whether anything moved. To connect a specific change to a specific gain (hard, because a round of work ships many changes at once), we lean on designs built for exactly this: tracking the metric over time and watching for a break at the moment of the change (an interrupted time-series), and single-case designs that change one thing at a time. And we report how big the effect was, not just that it went up. "Bigger than the noise floor, by this much" is a claim; "it improved" is a vibe.
Measuring the right thing
We check the score actually captures visibility
A score is only worth having if it measures what it claims to. Makefield's audit covers seven dimensions — recognition, accuracy, competitor framing and recommendation (the outcome), and documentation, structured data and citation authority (the machinery). As a dataset grows, we test whether those dimensions hold together the way the model says they should, and — once a real business outcome is in view — how much of that outcome the score can actually explain. Where the evidence isn't there yet, we say so plainly instead of dressing an assumption up as a finding.
The machinery · what a company can change
The outcome · what the engines say
In the open
Caveats first, and we publish what didn't work
Every study and audit leads with its sample and its limits, not its headline. Observations are dated, sourced, and backed by screenshots, so the work is reproducible. And when we expect an effect and find nothing, we publish that too — a negative result is honest, rare in this field, and worth more than a confident overclaim.
Independence
We measure. We never do the fix.
Makefield never does the fix, never takes a cut from whoever does, and no finding depends on who commissions it. The ranked list of what to change is part of the finding; the work belongs to someone else, and a re-run of the same measurement shows whether it moved the number, against the standard locked before the work. You can't be the referee and sell the thing you're grading.
Why it's here
Open, and free
This is published as educational material because the field is full of confident assertion and short on measurement, and that's a problem worth fixing in public. The method is on the open course, the findings in the field notes, the definitions in the reference, and the audit framework on GitHub. Published by Makefield, Berlin. Free to read, use, and cite with attribution — there's nothing to buy.