A New Paradigm for AI Search Metrics: Why Current Visibility Tools Miss the Mark and What Truly Drives Business Value

In the rapidly evolving landscape of artificial intelligence, a new "vanity metric" has emerged, captivating marketing teams and leading many astray: AI search visibility. The proliferation of AI visibility tools in the market, designed to count how often a large language model (LLM) mentions or cites a brand in response to a prompt, creates a deceptive sense of progress. This measurement approach, which superficially resembles the rank tracking prevalent in traditional search engine optimization (SEO) for two decades, is fundamentally different. Experts argue that the growing disparity between what these tools measure and what genuinely impacts business outcomes necessitates a re-evaluation of current strategies. This article, informed by extensive industry discourse and recent data, aims to delineate the metrics that truly matter in AI search, distinguishing them from those that merely offer an illusion of success.
The Flawed Foundation: Prompt Tracking and Its Pitfalls
The prevailing method for assessing AI search performance, prompt tracking, is proving to be an inadequate instrument for most organizations. The sales pitch is compelling: a tool simulates user queries across major AI platforms like ChatGPT, Perplexity, and Google’s AI Overviews, then reports brand mentions. Its resemblance to traditional rank tracking makes it an easy sell, leading to an oversaturation of such tools in the market. However, this focus on the most visible and obvious metric often overlooks its actual relevance to business objectives.
Jono Alderson, a respected technical SEO consultant, articulates this critical objection succinctly. He states, "We need to instead try and influence how the machine perceives us. And that’s not prompt tracking, which is what everyone is doing at the moment." While acknowledging a minor role for prompt tracking, Alderson incisively diagnoses its core flaw: "It’s copy-paste the current modality of rank tracking into a new thing. It doesn’t really fit, but it’s better than nothing." This sentiment underscores a broader industry struggle to adapt established methodologies to a radically new technological environment.
Beyond its conceptual misalignment, prompt tracking often operates on an ungrounded assumption about user behavior. Teams typically construct a list of prompts they hope customers will use, then measure their brand’s appearance against these hypothetical queries. For many brands, this invented list bears little resemblance to actual user questions, primarily because an AI prompt is not equivalent to a traditional keyword. The complexity deepens when attempts are made to ground these prompts in real search data, as AI itself is rapidly corrupting this data, rendering it unreliable almost as quickly as it can be analyzed.
A firsthand account illustrates this distortion. Last year, an investigation into a peculiar data leak revealed real users’ ChatGPT prompts appearing within Google Search Console, the primary tool for website owners to monitor search traffic. Analytics consultant Jason Packer published these findings, which were subsequently covered by numerous outlets including Ars Technica. The leak was traced to a buggy prompt box that inadvertently triggered a Google search almost every time, with ChatGPT URLs leading the queries. Google then tokenized these, leading to website owners seeing strangers’ private prompts in their dashboards. This phenomenon contributed to a pattern dubbed "crocodile mouth" in Search Console, where impressions spike dramatically while corresponding clicks plummet.
This leak was a visible manifestation of a widespread, yet often invisible, problem. AI systems constantly query Google to ground their answers, expanding a single user prompt into multiple parallel searches. These machine-generated searches register as impressions on ranking pages, yet no human ever sees the results or clicks through. Consequently, a surge in impressions unaccompanied by a rise in clicks does not signify increased human demand; it increasingly represents machines consuming information without direct user interaction. This distortion is also evident in search-trend and keyword-volume data, where climbing curves no longer reliably indicate human interest. While Google has integrated AI visibility reporting directly into Search Console, this report primarily displays impressions—the very metric AI systems inflate—while withholding AI click data, preventing a crucial check on true engagement.
Citation vs. Recommendation: A Critical Distinction
Perhaps the single most vital distinction in AI search measurement is that a citation does not equate to a recommendation. A citation occurs when an LLM names a page as a source for its answer. A recommendation, conversely, is when the model explicitly advises the user to choose a particular brand or product. Most current tools count the former, allowing marketers to mistakenly infer the latter, a perilous assumption.
Recent studies provide compelling evidence for this disconnect. Lily Ray conducted an analysis of Google’s AI Overview answers for 100 "best of" business software queries across three checkpoints in 2026. Her findings were striking: when a brand’s own self-promotional listicle was cited as a source, that brand was omitted from the actual recommendation in a staggering 69% of cases (224 out of 323 cited self-promotional listicles). This suggests that Google’s AI was intelligently parsing the content, extracting valuable information, and then recommending competitors mentioned within the cited page rather than the citing brand itself.
Further reinforcing this, Jeff Oxford’s team at Visibility Labs tested 20,000 ChatGPT responses and discovered that product recommendations changed in 80.2% of cases once the search functionality was activated. Crucially, there was only a weak 0.4 correlation between being cited and being recommended. Separately, BrightEdge, examining five different search engines, observed a similar pattern: while source overlap between engine pairs ranged from 16% to 59%, the set of recommended brands remained within a much tighter 36% to 55% band. This indicates a more stable, curated set of recommendations independent of source citations. Kevin Indig’s extensive analysis of 3.7 million citations further revealed that 91% of cited URLs appear in only one engine, highlighting the lack of portability for citation footprints across platforms.
Alisa Scharf, Chief AI Officer at Seer Interactive, has long championed this perspective. She asserts that "citations are an even worse metric than page one visibility, because they don’t necessarily indicate that your brand is mentioned in that response." Scharf likens citations to being on page two or three of Google—a leading indicator, but far from a direct measure of success. She outlines a clear hierarchy: "There’s the citation where your webpage is mentioned. There’s the mention where you’ve got your brand in the response. But rarely is ChatGPT or Claude specifically saying, you should go with X." It is this final step, the explicit recommendation, that translates into business value, yet prompt-tracking tools often conflate it with a mere footnote.
Malte Landwehr, who oversees product and marketing at Peec AI, offers a vivid illustration of this divergence. He recounted a scenario where a now-defunct tool became one of the most frequently cited sources in its category by ChatGPT. Despite this high citation volume, "They didn’t gain visibility as a brand," Landwehr noted. Instead, they inadvertently gained "power over what brands are recommended by LLMs," demonstrating that being the source and being the chosen recommendation are entirely distinct measurements.
The Volatility of AI Answers: Beyond Single-Shot Measurement
A single measurement of an AI answer holds minimal value due to the inherent variability of LLM responses. This crucial aspect is often obscured by prompt-tracking dashboards, which present data as if it were stable and definitive.
Rand Fishkin, who founded the audience-research firm SparkToro, conducted a study to quantify this variability. As he explained, "You are not getting an answer when you ask. You are getting one of thousands or potentially millions of answers, and every time you ask, it’s gonna be different. Every different person who asks is gonna get a different list, a different number of items, a different order, and a different set of recommendations." To illustrate the extreme nature of this variance, Fishkin revealed that, on average, one would need to query Claude or ChatGPT 1,500 times to obtain two answers with the exact same list of brands in the same order.
This astonishing figure fundamentally undermines the validity of single-shot measurements. It does not imply that AI visibility is immeasurable; rather, it dictates that it must be approached with the statistical rigor applied to polling, not the deterministic checking of a search rank. Fishkin explicitly states that a reliable signal is achievable, but it requires diligent work: "If you ask the right number of prompts, the right number of times, with some variability, you can get a statistical number that’s basically plus or minus 5%, or plus or minus 1% if you go really hard." The methodology itself is sound; the problem lies in tools that perform a single query and present the result as a definitive ranking.
Towards Meaningful Metrics: Presence and Recommendation Share
The replacement for conventional prompt tracking is "presence": a measure of how often a brand is named across the entire answer space, critically assessed against whether that presence translates into a genuine recommendation and, ultimately, a user action.
Rand Fishkin champions this as the only honest metric. "Percent of visibility is the number that’s real, that an AI tracking tool should be giving you," he argues, drawing a parallel to 20th-century consumer surveys asking, "Have you heard of Nike shoes, have you heard of Adidas shoes?" Wil Reynolds, founder of Seer Interactive, further refines this by highlighting the importance of tracking not just appearance, but also the composition of the AI answer over time. He points out, "If you’re tracking visibility and you don’t also track things like the number of words or brands mentioned per model per prompt over time, you would not know that back in November ChatGPT doubled the length of the answer." When an answer’s length doubles, a brand’s raw visibility might increase without any actual enhancement in its perceived value or prominence; the user is simply presented with more words.
Underlying all these considerations is a crucial caveat, one Reynolds states bluntly: visibility is only meaningful if it is demonstrably linked to a tangible outcome. "You can be visible. That’s great," he concedes. "But somebody’s gotta actually take an action for you to make any money from that visibility. If you don’t track those two metrics against each other, you’re the sucker." This underscores the imperative for marketers to connect AI visibility to quantifiable business results, such as conversions, leads, or direct sales.
The author’s own experience with the "No Hacks" brand serves as a powerful testament to this approach. Through focused effort on influencing how AI systems perceive the brand entity, "No Hacks" successfully achieved recommendation as the best podcast for AI web strategy in Google’s AI Overviews—a status it did not hold just a month prior. This was not achieved by chasing prompt-tracking dashboards, but by fundamentally altering the machine’s understanding of the brand. Moreover, the quality of any recommendation share measurement hinges entirely on the authenticity of the prompts driving it. Unlike Search Console, which provides grounded query data, prompt tracking often begins with fabricated lists, disconnected from actual user behavior. Therefore, while measuring recommendation share is valuable, its accuracy is directly tied to the realism of the underlying prompts.
Lessons from the Past: The Recurring Vanity Metric Trap
The current fascination with superficial AI visibility mirrors a historical lesson the search industry took nearly two decades to internalize: impressions and clicks, in isolation, are often vanity metrics. They do not inherently guarantee revenue or business success. AI visibility, when pursued for its own sake, represents the same trap, merely cloaked in new technological garb. It might boost brand awareness, but without a clear link to action and revenue, it remains an easily inflated, yet ultimately hollow, metric.
Wil Reynolds, whose podcast episode title revolved around this very point, directly traces this cyclical pattern: "The vanity metric early was rankings, and then people went, wait, I gotta get traffic from those rankings, and then I need that traffic to turn into a business. So to me it’s just a regurgitation of what we did years ago." Jono Alderson extends this critique, suggesting that the precise attribution models that provided comfort for decades were never truly accurate: "the crutch and the lies that we’ve told ourselves for the last decade, that we can neatly attribute impression share through to clicks, through to actions, through to revenue. It’s never been true, and it’s getting less true." The fundamental task, he argues, predates modern tooling: to influence how people—and now machines—perceive a brand.
The Foundational Metric: Brand Accuracy
Before pursuing recommendation share, the most critical metric to establish and maintain is brand accuracy: whether the AI correctly describes a brand or entity at all. If an LLM harbors factual inaccuracies about a brand, any subsequent measurement, whether a recommendation or a refusal to recommend, is built on a flawed understanding of a non-existent version of that brand.
Achieving brand accuracy begins with clarity. This involves ensuring absolute consistency in brand representation across all platforms, fostering consistent external descriptions, and providing unambiguous answers to fundamental questions about the brand’s identity and offerings. Duane Forrester, a key figure in the launch of Schema.org and builder of Bing Webmaster Tools, frames the ultimate goal not as high ranking, but as becoming the "trusted source." He states, "Your goal should be to be seen as the canonical for whatever your question is. Not rankings, but that you are the source of knowledge." Forrester posits that machines exhibit a "useful laziness": "It costs money and cycles and tokens to go build trust. So if I’ve done all that work and I trust you, and you’re a good answer, and my consumer is happy with that answer, why would I change?" A brand established as a canonical source gains a significant advantage in sustained AI visibility.
Alisa Scharf has translated this imperative into a practical measurement strategy: the brand accuracy audit. "You come up with a list of objective criteria," she explains. "It can’t be, we want to rank for best X for Y. It’s got to be: when were you founded, where are you based, what do you sell, who do you compete against." This list of non-negotiable facts is then systematically queried across various AI engines on a regular schedule. The model is scored on its consistent accuracy, not on whether it flattering a brand with a mention. This audit provides a robust foundation for all subsequent AI optimization efforts.
Navigating the Blind Spots: Training Cutoff and Platform Data
Honest measurement requires acknowledging inherent blind spots. In AI search, two significant challenges persist. The first is the training-data cutoff. A substantial portion of AI answers derives from pre-trained knowledge, frozen at an uncontrollable date, before any live grounding occurs. Currently, there is no clear method to ascertain whether ongoing optimization efforts are impacting these baked-in, potentially stale, answers. Marketers could be diligently optimizing against a version of the model’s knowledge that is months out of date.
The second blind spot concerns platform data. The market structure of frontier AI models presents a dilemma: pure-play model companies like OpenAI or Anthropic have little inherent incentive to share granular usage data with external entities. There is no clear business model that compels them to expose how their algorithms arrive at recommendations. In contrast, companies with broader ecosystems to protect, such as Google and Microsoft, offer some data. Google, for instance, adds AI impressions to Search Console, and Microsoft provides similar insights through Bing Webmaster Tools. They do so because these measurement surfaces benefit from user attention, which aligns with their larger platform strategies. However, this data remains relatively weak and incomplete. Whether pure-play model companies will ever open up this crucial data remains a significant open question, and the future measurability of AI search hinges heavily on their eventual decisions.
The Future of AI Search: Certainty as the New Currency
The imperative for brands to clearly define themselves for AI systems is rapidly becoming a deterministic core of marketing work, moving beyond soft branding considerations. This shift is underscored by recent legal precedents. A German court recently held Google liable for false statements generated by its AI Overview about a business, ruling that the AI answer constituted Google’s own speech.
This legal development suggests a profound implication: a platform now legally accountable for its AI’s pronouncements about businesses has a powerful incentive to only surface entities about which it is absolutely confident. One can envision an internal "confidence threshold" or certainty score. If an AI system is sufficiently certain about a brand’s identity and facts, it includes it; if not, it opts to omit the brand rather than risk generating inaccurate information and facing legal repercussions. While the precise mechanics are speculative, the direction seems clear. If this hypothesis holds true, the most critical metric for brands to measure will not be how often they appear, but rather the degree of certainty the machine holds about their identity, as this certainty will increasingly dictate their very presence in AI search results. Therefore, ensuring clear, consistent, and verifiable brand information across all digital touchpoints is paramount for future AI visibility and business success.






