§ The three-act story the industry just lived through
Act 1 — the prosecution. Rand Fishkin (SparkToro) put AI answer consistency on trial: 2,961 prompts, hundreds of volunteers, three engines. The result was brutal for the dashboard industry, the same prompt, re-run, produced the same brand list less than 1% of the time, and the same order less than 0.1% of the time. His conclusion, verbatim: “Any tool that gives you a ‘ranking position in AI’ is full of baloney.” And it's getting worse for trackers, not better: ChatGPT's memory now personalizes answers per user, meaning the “clean” answer a tool's test account sees may not represent what any real buyer sees.
Act 2 — the twist. Here's what makes Fishkin worth trusting: his own data pushed back on his own prior. While rankings proved fiction, brand presence frequency, how often your brand appears at all, across many repeated runs, turned out to be statistically stable enough to track. The loudest skeptic in the industry updated his position in public. “It's ALL baloney” became “rank positions are baloney; frequency trends are real.”
Act 3 — the synthesis. Kevin Indig, working from a 3.7-million-citation dataset, arrived at the same place from the methodology side: prompt tracking is salvageable if you use repeated runs, fixed sampling and confidence intervals, and his data added a finding every business should know: 91% of cited URLs appear in only one engine. Winning ChatGPT says almost nothing about Gemini. Aleyda Solis completed the picture with her measurement framework, presence → readiness → business impact, and her own consistency numbers: only ~30% of brands stay visible run-to-run. Even Ahrefs, which sells a tracking product, published a piece titled “You Can't Track AI Like Traditional Search.”
When the skeptic, the methodologist, the framework-builder and a tool vendor all land in the same place, that's not an opinion anymore. That's the state of knowledge.
§ The 5 rules they all agree on
- Fix your prompt set. Choose the 10–25 questions your real buyers ask, and never change them mid-measurement. Rotating prompts makes every report incomparable with the last.
- Fix the model and settings. Same engines, same configuration, every time. A score built across shifting models measures the weather, not your visibility.
- Run repeatedly, on a schedule. One run is an anecdote. Frequency across 10+ runs is data. This is the single difference between measurement and theater.
- Read deltas against your own baseline, never absolute scores. “62/100” means nothing. “Named in 3 of 10 runs in April, 6 of 10 in July, on the same prompts” means everything. Your only valid comparison is you, earlier.
- Treat “AI rank position” as a red flag. Not a simplification, a fiction. Any vendor leading with it has told you how seriously to take the rest of their dashboard.
§ The 6 questions to ask any AEO tool before paying
- How many times do you run each prompt before reporting a number? (One = walk away.)
- Do you report frequency and trend, or a position? (Position = the baloney flag.)
- Which engines, which modes, and do you disclose when providers change models underneath you?
- How do you handle memory and personalization, do you acknowledge your test accounts aren't real users? (The honest answer includes the word “limitation.”)
- Can I export the raw runs, the actual answers, not just your score? (Receipts or it didn't happen.)
- What did my baseline look like before your work started? (No baseline = progress can never be proven. We've written about that grift before.)
A good vendor passes all six without flinching, several will. This isn't an argument against tools; it's an argument against single-number certainty. The tools aren't fake. The confidence is.
§ What this means if you're a business owner
You're going to be pitched an AI visibility score this year, by a tool, an agency, or both. The mature posture, in one paragraph: buy trendlines, never verdicts. Demand repeated-run methodology and raw receipts. Compare only against your own baseline. Cross-check the vendor's story against the first-party reports that now exist (Google's new gen-AI report and Bing's Citation Share) and against the crudest, honest metric, how many new customers say “an AI recommended you.” Where all of those point the same direction, believe the direction.
The measurement crisis is real. It's also survivable, the people who study this hardest just handed you the manual.
Measure the way the evidence says you must
FirePencil measures with fixed prompt sets, repeated scheduled runs and trendlines against your own baseline, receipts, not scores, and then does what no dashboard can: executes the fixes that move the trendline, weekly and owner-approved. Get your day-zero baseline with the free AEO Audit.
§ Frequently asked questions
Can you track your ranking position in AI answers?
What did Rand Fishkin's AI visibility study find?
What are the rules for reading an AI visibility score?
What questions should I ask an AEO tool before buying?
FirePencil.AI is an autonomous AEO agent. This article summarizes third-party research and public commentary (Rand Fishkin/SparkToro, Kevin Indig, Aleyda Solis and Ahrefs) for general information; figures are as reported by those sources. It is not a guarantee of any specific result. Third-party names (ChatGPT, Gemini, Perplexity, Claude, Google AI Overviews) are trademarks of their respective owners; use is descriptive.
