Thunderdome B2B SaaS AI Perception Index

← All categories

September 2026 snapshot · 72 AI answers Locked up

LLM & Agent Evals: AI Visibility Ranking

Braintrust is in the Perception Lab → aided, forced-choice, and grounded measurement of what Braintrust is to the models, beyond this ranking.

Platforms for evaluating and observing LLM apps and AI agents: eval suites, LLM-as-judge scoring, tracing, and production monitoring. Right now, LangSmith is the brand AI assistants recommend most, with a Visibility Score of 77 and the #1 spot in 46% of answers. Head-to-head: LangSmith vs Weights & Biases →

These rankings are measured from what the three models know from training, not a live web search. When the models DO search the web, see which sources they cite. Methodology.

LangSmith dominates; no one else is close.

LangSmith's 76.9 score sits 26 points above second-place Weights & Biases, and it shows up in 85% of answers while being picked first nearly half the time. Gemini is its loudest advocate at 93.8, while ChatGPT and Claude score it more modestly, yet even those modest scores beat most competitors. Behind LangSmith, Weights & Biases holds a stable second on the back of ChatGPT and Claude, but Braintrust, Langfuse, and Arize are essentially splitting thin air with scores bunched between 30 and 36.

AI-generated analysis of the September 2026 measurements.

# Brand Visibility Score Mention rate Avg. position #1 pick rate Models
1 LangSmith
77
±5.8
85% 1.9 46% ChatGPT #2 Claude #1 Gemini #1
2 Weights & Biases
51
±6.3
64% 3.1 3% ChatGPT #1 Claude #3 Gemini #5
3 Braintrust
36
±6.1
44% 2.9 15% ChatGPT #5 Claude #2 Gemini #11
4 Langfuse
35
±6.2
43% 2.9 10% ChatGPT #3 Claude #5 Gemini #8
5 Arize AI
30
±6.1
36% 2.7 13% ChatGPT #4 Claude #7 Gemini #7
6 Arize Phoenix
29
±7.4
40% 3.7 1% ChatGPT #7 Claude #8 Gemini #2
7 Promptfoo
28
±6.2
36% 3.3 11% ChatGPT #11 Claude #4 Gemini #3
8 Ragas
21
±5.5
33% 4.8 0% ChatGPT #12 Claude #6 Gemini #6
9 DeepEval
19
±5.3
32% 5.2 0% ChatGPT #8 Claude #12 Gemini #4
10 Helicone
13
±4.6
24% 5.5 0% ChatGPT #9 Claude #9 Gemini #9
11 Humanloop
9
±4.3
14% 4.4 0% ChatGPT #6 Claude Gemini
12 Datadog
7
±3.4
10% 3.6 0% ChatGPT #14 Claude #11 Gemini #10
13 OpenAI
6
±4.6
10% 5.0 1% ChatGPT #10 Claude Gemini #13
14 Honeyhive
5
±3.3
10% 6.1 0% ChatGPT Claude #13 Gemini #12
15 Confident AI
4
±2.6
7% 4.6 0% ChatGPT #13 Claude #10 Gemini

The ± under each score is the reliability band: how much the number would move if we re-ran the identical measurement. It is tight because each prompt is sampled multiple times. It is not the same as how much the ranking depends on which prompts we ask — a separate, wider figure shown on hover.

Just below the cutoff: AgentOps (4), TruLens (3.9), Patronus AI (3.8), Portkey (3.5), Evidently AI (3.2). 42 brands were scored in this category; the leaderboard shows the top 15.

Visibility Score: position-weighted presence across all measured answers, 0 to 100. A score of 100 means the brand was the first recommendation in every answer. Every point in a trend line is a live monthly measurement; there is no modeled or backfilled history. Full methodology.

Openness 23/100: the share of this category's recommendation weight not held by the leader.

Compare brands in LLM & Agent Evals head-to-head →  ·  Explore the prompt tree →

Get alerted on big movements

We rerun the index every month. Drop your email and pick what to watch in LLM & Agent Evals: the whole category or specific brands. Free, unsubscribe anytime.

Or watch specific brands:

Do the models agree?

Per-model Visibility Scores for the top 10. A long gray bar means the three assistants disagree about the brand. Hover a row for exact values.

ChatGPT Claude Gemini
0 25 50 75 100 Visibility Score · 0 = never mentioned, 100 = always the first recommendation LangSmith Weights & Biases Braintrust Langfuse Arize AI Arize Phoenix Promptfoo Ragas DeepEval Helicone

Where they land in the answer

When a brand is mentioned, where does it appear? Box shows the middle 50% of positions, the thick line is the median, whiskers are the extremes. Hover a row for the detail.

1 2 3 4 5 6 7 8 9 10 Position in the answer · 1 = the first product recommended LangSmith Weights & Biases Braintrust Langfuse Arize AI Arize Phoenix Promptfoo Ragas DeepEval Helicone

What each model recommends first

ChatGPT

gpt-5.4

1 Weights & Biases
64
2 LangSmith
63
3 Langfuse
52
4 Arize AI
38
5 Braintrust
32

Claude

claude-sonnet-4-6

1 LangSmith
74
2 Braintrust
66
3 Weights & Biases
56
4 Promptfoo
35
5 Langfuse
32

Gemini

gemini-3-flash-preview

1 LangSmith
94
2 Arize Phoenix
44
3 Promptfoo
43
4 DeepEval
34
5 Weights & Biases
32

The questions we asked

Each model answered every prompt 3 times in September 2026. Prompts are brand-neutral so no vendor gets seeded into the question.

  1. Best overall
    What's the best platform for evaluating LLM and AI agent outputs? See who wins it →
  2. Use-case fit
    We're shipping an LLM-powered feature and need eval suites for prompts and models, plus tracing to debug failures in production. What should we use? See who wins it →
  3. Top tools in 2026
    What are the top LLM evaluation and observability platforms in 2026? See who wins it →
  4. Startup & small team
    Which LLM evals tool should an early-stage AI startup use? See who wins it →
  5. Enterprise pick
    Which LLM evaluation and observability platform should a large enterprise standardize on, considering security, compliance, and scale? See who wins it →
  6. Feature-led ask
    I want offline evals, LLM-as-judge scoring, and production tracing for our AI agents in one platform. What do you recommend? See who wins it →
  7. If you could pick one
    If you could only pick one LLM evaluation platform, which one and why? See who wins it →
  8. Budget & alternatives
    What's a good open-source or budget-friendly alternative to the big LLM evaluation platforms? See who wins it →