Every CMO now has a slide that says "our brand must be visible in AI." Platforms like Profound, Otterly, Peec, and Goodie have grown quickly on the back of that urgency, each offering some version of a Share of Voice dashboard that answers the question: how often does ChatGPT, Perplexity, or Google AI Overviews mention us compared to competitors? The category is legitimate. The dashboards are not yet good enough.
This article breaks down what current market-leader platforms measure, where their dashboards leave decision-making value on the table, and then builds an alternative framework — demonstrated with an interactive dashboard using mock data — that treats LLM brand visibility as a multi-dimensional signal rather than a single citation-frequency number.
How the market leaders work today
All four platforms share the same core mechanic: a prompt library is re-run against multiple LLM engines on a scheduled cadence (daily or near-real-time), responses are parsed for brand mentions, and results are aggregated into a Share of Voice percentage. The differences are in coverage, granularity, and what they do with the data beyond the mention count.
| Platform | LLM platforms tracked | Citation position | Sentiment | Source attribution | Hallucination flag | Topic-cluster SOV |
|---|---|---|---|---|---|---|
| Profound | ChatGPT 5.5, Perplexity, Claude Opus 4.7, Gemini 3.1, Google AIO | Partial | ✓ | ✓ | ✓ | ✗ |
| Otterly | 6 platforms, 50+ countries | ✗ | ✓ | Partial | ✗ | ✗ |
| Peec AI | ChatGPT 5.5, Perplexity, Gemini 3.1, Copilot, Google AIO, 115+ languages | ✗ | Partial | ✓ | ✗ | Partial |
| Goodie | Major platforms, SOC 2 certified | ✗ | ✓ | ✓ | ✗ | ✗ |
The three gaps that matter
1. Citation frequency ≠ citation influence
Every current dashboard counts mentions. None weights them by position in the LLM's response. A brand mentioned first — in the opening sentence of a comparative answer — has radically different influence on the user than the same brand listed fourth in a closing "you might also consider" clause. On Google Search, position 1 vs position 4 represents a click-through rate drop of roughly 8× (Backlinko, 2020). The dynamic in LLM responses is at least as large: users reading a conversational answer remember what came first.
A Citation Prominence Score needs to account for: whether the brand leads the response, whether it appears in the first third, the middle, or only at the end, and whether the mention is the subject of a comparative statement or a parenthetical aside.
2. Platform-level SOV hides divergent risk and opportunity
Aggregated SOV conceals the fact that different platforms behave differently. Research comparing the major LLMs shows consistent patterns: ChatGPT 5.5 tends to favour established category leaders, Perplexity returns more citations per answer and is more source-diverse, Google AI Overviews shows the highest brand diversity (reflecting its underlying web index), and Copilot has the most concentrated citation inequality (Nightwatch, 2026). A brand that has 32% aggregate SOV might be at 38% on Gemini 3.1 and 28% on ChatGPT 5.5 — a structural difference with completely different remediation strategies. Aggregated dashboards bury this (Profound vs LLM Pulse, 2026).
3. What kind of query drives the mention is more important than the mention count
Not all prompts are equal. A brand cited in response to "what is the best enterprise CRM?" has different commercial value than the same brand cited in response to "what is a CRM?" — one is commercial intent, one is informational. Current tools let you define a prompt library but do not automatically segment SOV by query intent category. A brand might dominate informational queries while being absent from comparison and commercial queries — which is precisely the pattern that fails to convert AI visibility into pipeline.
The leading indicator for AI-driven revenue is not overall SOV. It is commercial-intent SOV — the share of responses to queries with explicit buying or comparison signals where your brand appears, in a positive or neutral context, in the first half of the response. Current tools measure overall SOV; this metric requires combining intent classification, position scoring, and sentiment — none of which any current platform does end-to-end.
A better measurement framework
The dashboard at the end of this article proposes five dimensions that together produce an actionable picture. They build on what current tools measure and add the layers they are missing:
- Overall SOV — the standard citation share, tracked over time against a fixed competitor set
- Platform-level SOV delta — per-platform SOV for your brand versus the category leader, showing where you are structurally over- or under-indexed
- Citation Prominence Score — weighted mention rate where first-half mentions count more than late mentions
- Topic-cluster SOV — SOV segmented by query intent category (informational, comparison, commercial, brand)
- Sentiment quality — positive/neutral/negative split tracked over time, not just as a snapshot
The interactive dashboard below applies this framework to a fictional B2B software brand, Zentrava, competing against three equally fictional rivals — Korvant, Auralyx, and Drovetta. All brand names and all data are invented. Beyond the five core dimensions, the demo adds the layers the article argues current tools are missing: an auto-generated signals feed computed live from the underlying data, a switchable competitor benchmark, an answer-level explorer with response excerpts, citation source attribution, and a hallucination log. The implementation uses no external libraries — pure HTML, CSS, and vanilla JavaScript — so the pattern is directly replicable in any reporting environment.