Thunderdome B2B SaaS AI Perception Index

← LLM & Agent Evals ranking

Brand report · September 2026

Promptfoo Promptfoo: AI Visibility Report

promptfoo.dev ↗

⚔ Promptfoo appears in the Perception Lab: forced-choice runs vs Braintrust →

When buyers ask ChatGPT, Claude, and Gemini about llm & agent evals, Promptfoo ranks #7 of 15, a Visibility Score of 28 in LLM & Agent Evals .

These rankings are measured from what the three models know from training, not a live web search. The live-web grounded surface is rolling out across the index. Methodology.

How AI sees Promptfoo

Mid-pack in one race, with no other races entered.

Where it wins. Promptfoo earns a 36% mention rate in LLM & Agent Evals, and startup teams specifically pull it into consideration more than any other buyer segment.

Where it loses. At rank 7 of 15 with only an 11% first-pick rate, it loses the final decision to stronger signals from LangSmith and DeepEval, and Feature-led buying queries skip it entirely.

AI-generated analysis of the September 2026 measurements.

LLM & Agent Evals · #7 of 15

28 / 100 first snapshot (September 2026)
36%
mention rate (72 answers)
3
avg. position
11%
first pick
6%
share of voice
First snapshot: September 2026. The trend line appears with the second monthly measurement; no modeled history.
ChatGPT
5
Claude
35
Gemini
43

One of the index's bigger model splits: see where the models disagree →

How the models portray Promptfoo

Being named is not the same as being recommended. Each mention is graded endorsed (a strong pick), listed (a neutral option), or caveated (named with a reservation).

69% 19% 12%
endorsed · 95% band 50–83% listed caveated graded across 26 mentions in LLM & Agent Evals answers
Prompt themes Promptfoo would want to own

Share of each theme's answers that mention Promptfoo. Hover a row for the exact prompt. Compare shapes on the head-to-head page.

Startup & small team
7/9 · #1×3
If you could pick one
5/9 · #1×1
Best overall
3/9 · #1×2
Top tools in 2026
3/9 · #1×1
Budget & alternatives
2/9 · #1×1

Gaps Feature-led ask

Most co-mentioned competitors

Share of Promptfoo's mentions where the AI models name this brand in the same answer. That is the real competitive set in the AI channel.

LangSmith
96%
Ragas
73%
DeepEval
50%
Braintrust
50%
Weights & Biases
46%
Arize Phoenix
46%

Scores are measured from real AI answers, refreshed monthly. Methodology.

Get alerted on big movements

We rerun the index every month. Drop your email and pick what to watch in LLM & Agent Evals: the whole category or specific brands. Free, unsubscribe anytime.

Or watch specific brands:

Work at Promptfoo? Put this on your site: a live "AI-recommended" badge, grounded in this data, that updates monthly and links back here. Free. Get the badge →