FirePencil.AIFirePencil.AI
HomeBlog › How to read an AI visibility score
AI Search / Measurement · Post 11

Fishkin Called AI Rankings “Baloney.” Then His Own Study Surprised Him.

How to actually read an AI visibility score, the five measurement rules Rand Fishkin, Kevin Indig and Aleyda Solis all converge on, and the six questions to ask any AEO tool before you pay.

Rand Fishkin called AI rankings baloney, then his own study on AI visibility scores surprised him
The most credible skeptic in search ran the biggest independent test of AI answer consistency, and the result cuts both ways.
Bottom line up front The most credible skeptic in search ran the biggest independent test of AI answer consistency, 2,961 prompts across ChatGPT, Claude and Gemini, and found less than a 1-in-100 chance that two runs return the same list of brands. His verdict: any tool selling you an “AI ranking position” is selling baloney. But the same data revealed what is measurable, and when you line his findings up against Kevin Indig's methodology work and Aleyda Solis's framework, the three most-followed voices in search converge on five rules. If you're being pitched an AEO dashboard this quarter, this is how to read the score, and the six questions to ask before paying.

§ The three-act story the industry just lived through

Act 1 — the prosecution. Rand Fishkin (SparkToro) put AI answer consistency on trial: 2,961 prompts, hundreds of volunteers, three engines. The result was brutal for the dashboard industry, the same prompt, re-run, produced the same brand list less than 1% of the time, and the same order less than 0.1% of the time. His conclusion, verbatim: “Any tool that gives you a ‘ranking position in AI’ is full of baloney.” And it's getting worse for trackers, not better: ChatGPT's memory now personalizes answers per user, meaning the “clean” answer a tool's test account sees may not represent what any real buyer sees.

Act 2 — the twist. Here's what makes Fishkin worth trusting: his own data pushed back on his own prior. While rankings proved fiction, brand presence frequency, how often your brand appears at all, across many repeated runs, turned out to be statistically stable enough to track. The loudest skeptic in the industry updated his position in public. “It's ALL baloney” became “rank positions are baloney; frequency trends are real.”

Act 3 — the synthesis. Kevin Indig, working from a 3.7-million-citation dataset, arrived at the same place from the methodology side: prompt tracking is salvageable if you use repeated runs, fixed sampling and confidence intervals, and his data added a finding every business should know: 91% of cited URLs appear in only one engine. Winning ChatGPT says almost nothing about Gemini. Aleyda Solis completed the picture with her measurement framework, presence → readiness → business impact, and her own consistency numbers: only ~30% of brands stay visible run-to-run. Even Ahrefs, which sells a tracking product, published a piece titled “You Can't Track AI Like Traditional Search.”

📊 The numbers everyone should carry: 2,961 prompts tested · same brand list <1% of re-runs · same order <0.1% · 91% of cited URLs appear in only one engine · only ~30% of brands stay visible run-to-run.

When the skeptic, the methodologist, the framework-builder and a tool vendor all land in the same place, that's not an opinion anymore. That's the state of knowledge.

§ The 5 rules they all agree on

  1. Fix your prompt set. Choose the 10–25 questions your real buyers ask, and never change them mid-measurement. Rotating prompts makes every report incomparable with the last.
  2. Fix the model and settings. Same engines, same configuration, every time. A score built across shifting models measures the weather, not your visibility.
  3. Run repeatedly, on a schedule. One run is an anecdote. Frequency across 10+ runs is data. This is the single difference between measurement and theater.
  4. Read deltas against your own baseline, never absolute scores. “62/100” means nothing. “Named in 3 of 10 runs in April, 6 of 10 in July, on the same prompts” means everything. Your only valid comparison is you, earlier.
  5. Treat “AI rank position” as a red flag. Not a simplification, a fiction. Any vendor leading with it has told you how seriously to take the rest of their dashboard.

§ The 6 questions to ask any AEO tool before paying

  1. How many times do you run each prompt before reporting a number? (One = walk away.)
  2. Do you report frequency and trend, or a position? (Position = the baloney flag.)
  3. Which engines, which modes, and do you disclose when providers change models underneath you?
  4. How do you handle memory and personalization, do you acknowledge your test accounts aren't real users? (The honest answer includes the word “limitation.”)
  5. Can I export the raw runs, the actual answers, not just your score? (Receipts or it didn't happen.)
  6. What did my baseline look like before your work started? (No baseline = progress can never be proven. We've written about that grift before.)

A good vendor passes all six without flinching, several will. This isn't an argument against tools; it's an argument against single-number certainty. The tools aren't fake. The confidence is.

§ What this means if you're a business owner

You're going to be pitched an AI visibility score this year, by a tool, an agency, or both. The mature posture, in one paragraph: buy trendlines, never verdicts. Demand repeated-run methodology and raw receipts. Compare only against your own baseline. Cross-check the vendor's story against the first-party reports that now exist (Google's new gen-AI report and Bing's Citation Share) and against the crudest, honest metric, how many new customers say “an AI recommended you.” Where all of those point the same direction, believe the direction.

The measurement crisis is real. It's also survivable, the people who study this hardest just handed you the manual.

Measure the way the evidence says you must

FirePencil measures with fixed prompt sets, repeated scheduled runs and trendlines against your own baseline, receipts, not scores, and then does what no dashboard can: executes the fixes that move the trendline, weekly and owner-approved. Get your day-zero baseline with the free AEO Audit.

§ Frequently asked questions

Can you track your ranking position in AI answers?
No. Fishkin's 2,961-prompt study across ChatGPT, Claude and Gemini found less than a 1-in-100 chance that two runs return the same brand list, and less than 1-in-1,000 for the same order. Any tool selling you an "AI ranking position" is selling a fiction. What is trackable is presence frequency, how often your brand is named at all across many repeated runs against your own baseline.
What did Rand Fishkin's AI visibility study find?
SparkToro's Rand Fishkin ran 2,961 prompts across ChatGPT, Claude and Gemini. The same prompt re-run produced the same brand list less than 1% of the time and the same order less than 0.1% of the time. But his own data showed brand presence frequency across repeated runs is stable enough to track, a finding echoed by Kevin Indig's 3.7-million-citation dataset and Aleyda Solis's framework.
What are the rules for reading an AI visibility score?
Five rules the top researchers agree on: fix your prompt set (10 to 25 real buyer questions, never changed mid-measurement); fix the model and settings every run; run repeatedly on a schedule (10+ runs, not one); read deltas against your own baseline rather than absolute scores; and treat any "AI rank position" number as a red flag, because it is a fiction.
What questions should I ask an AEO tool before buying?
Ask: how many times do you run each prompt before reporting a number (one means walk away); do you report frequency and trend or a position (position is the baloney flag); which engines and modes, and do you disclose provider model changes; how do you handle memory and personalization; can I export the raw runs, not just the score; and what did my baseline look like before your work started.

FirePencil.AI is an autonomous AEO agent. This article summarizes third-party research and public commentary (Rand Fishkin/SparkToro, Kevin Indig, Aleyda Solis and Ahrefs) for general information; figures are as reported by those sources. It is not a guarantee of any specific result. Third-party names (ChatGPT, Gemini, Perplexity, Claude, Google AI Overviews) are trademarks of their respective owners; use is descriptive.