Thunderdome B2B SaaS AI Visibility Index

← Back to the index

Full transparency

Methodology

The index answers one question: when a real buyer asks an AI assistant what software to buy, which brands come back in the answer? Here is exactly how we measure it.

1. Buyer-intent prompts

Each of the 103 categories has a fixed set of 8 prompts written the way real buyers ask: by company size, budget, use case, and "if you could only pick one." Prompts are brand-neutral. We never name a vendor in a question, because that would seed the answer. Every prompt is published verbatim on its category page.

2. Models tested

We query the production APIs of the three most-used AI assistants:

Each model answers every prompt 3 times per monthly snapshot, because LLMs are non-deterministic and a single run can mislead. The September 2026 snapshot covers 7,416 answers in total.

A note on surface. An assistant's API and its consumer app are not the same measurement surface: the app layers web search, memory, and personalization on top of the model, so app answers can differ from the API baseline measured here. We measure the API deliberately, because it is stable, reproducible, and isolates what the model itself believes about a market. App-surface and AI-search-surface tracking are on the roadmap as their own labeled dimensions, never a silent substitute, and where the two disagree, that divergence will be reported rather than averaged away.

3. Extraction

From every answer we extract the ordered list of products the model presents as recommendations, in order of first appearance. Products mentioned only as an integration, a comparison point, or a warning do not count. Brand names are normalized so "HubSpot CRM" and "HubSpot" are the same brand.

4. Scoring

Our metric definitions align with how the AI-visibility industry measures brand presence in LLM answers. Mention rate, share of voice, and average position are the standard vocabulary used by dedicated tracking platforms like Profound, Peec AI, and Otterly and by benchmark research such as Conductor's AEO/GEO benchmarks report and Discovered Labs' AEO benchmark guide. Multi-sampling the same prompt is standard practice across these tools because single runs of a non-deterministic model mislead.

We add our own flavor on top of the standard metrics: a position-weighted Visibility Score that rewards being the first name out of the model's mouth, strictly brand-neutral prompts, and full publication of every prompt we ask. Most commercial trackers keep their query sets private; we think an index you can audit is an index you can cite.

The index publishes a fresh snapshot monthly, and every point on a trend line is a live measurement from one of those published snapshots. The index launched with the September 2026 snapshot, so trend lines and month-over-month deltas appear once a brand has two or more real months of data. We do not model, estimate, or backfill history: months we did not measure simply do not exist in the charts.

6. Sentiment & stance

Being named in an answer is not the same as being recommended. So for every mention, we also grade how the answer portrays the brand, judged only from that answer's own words, on a three-point scale:

A binary positive/negative split would be uninformative here: in "what's the best X" answers almost everything named is at least somewhat positive. The three-point scale instead separates the brand that is genuinely recommended from the one that is merely listed, and surfaces the "mentioned-but-with-a-catch" pattern that a raw mention count hides.

Each brand's endorsement rate is the share of its mentions graded endorsed, across every answer, model, and sample in the category. Because that is a proportion over a finite number of mentions, we publish it with a 95% Wilson confidence interval, not a bare figure, and we suppress any brand with fewer than three graded mentions in a category. As with every other number here, wider bands mean fewer mentions and less certainty, and we say so rather than round the doubt away.

7. Market segmentation (prompt trees)

Models rarely answer a buying question with one flat ranking; in over 95% of collected answers they carve the market — "best for speed: X", "if you want open-source: Y" — and assign brands per segment. The prompt trees are built from exactly those assignments. Extraction is two-stage: a model pass pulls every explicit segment-to-brand assignment from each answer (a brand merely listed is not an assignment), then a second pass consolidates each category's raw segment labels into at most eight canonical segments, each tagged with one of nine standardized axes (use case, team size, budget, ease, power, deployment, speed, stack, industry) so trees compare across categories while the segment wording stays the market's own.

Honesty rules: a brand is only shown in a segment when the models place it there in at least two answers; segment evidence quotes are displayed only when the quoted span is mechanically verified to appear verbatim in the collected answer (formatting normalized, words untouched) — a candidate quote that fails that check is discarded, and a segment with no provable quote ships without one. A category where the models never segment simply has no segmentation branches.

8. Brand knowledge (reputation)

Visibility asks whether the models recommend a brand. A separate question is what they actually know about it. So we also probe each model directly — "what is this product, what is it known for, its strengths, weaknesses, and who should avoid it?" — and grade the answer for knowledge depth: specificity, concreteness, and confidence, explicitly not length. A model that names real features, use cases, and genuine trade-offs scores high; one that offers generic filler or hedges scores low.

The trap here is that a model can describe an obscure brand confidently and wrongly — fluent confabulation reads like knowledge. So the headline is anchored by cross-model agreement: we ask whether all three models describe the same product consistently. High depth with high agreement is real reputation; high depth with low agreement is a warning we flag, not a score we trust. We publish the depth, the range across models, the agreement, and any factual divergence — never a lone confident number.

9. Data sourcing

Every number on this site traces back to a real AI answer we collected ourselves. Answers come directly from the models' production APIs at measurement time; nothing is scraped from consumer apps, bought from panels, estimated from third-party data, or generated synthetically. The raw answers behind each monthly snapshot are archived verbatim, because an AI answer is a measurement of a moment: once the models move on, it can never be collected again.

Limitations

Independence

No vendor pays to be included, excluded, or ranked. The index is free, carries no sponsorships, and no commercial relationship changes what the models said or how it is scored.

Citing the index

Everything here is free to use: quote it, chart it, republish it. All we ask is attribution with a link to thunderdome.io, and a note of the snapshot month (September 2026) since the rankings change monthly.