⚔ Arize Phoenix appears in the Perception Lab: forced-choice runs vs Braintrust →
When buyers ask ChatGPT, Claude, and Gemini about llm & agent evals, Arize Phoenix ranks #6 of 15, a Visibility Score of 29 in LLM & Agent Evals .
These rankings are measured from what the three models know from training, not a live web search. The live-web grounded surface is rolling out across the index. Methodology.
Middle of the eval pack, invisible to enterprise buyers.
Where it wins. Arize Phoenix earns a 40% mention rate in LLM & Agent Evals, meaning models surface it reliably when buyers ask about evaluation tooling. That keeps it competitive in a 15-tool field.
Where it loses. A 1.4% first-pick rate means models almost never open with Arize Phoenix as the go-to eval solution, and the total absence from Enterprise pick themes suggests it reads as a practitioner tool rather than a procurement-ready platform.
AI-generated analysis of the September 2026 measurements.
One of the index's bigger model splits: see where the models disagree →
Being named is not the same as being recommended. Each mention is graded endorsed (a strong pick), listed (a neutral option), or caveated (named with a reservation).
Share of each theme's answers that mention Arize Phoenix. Hover a row for the exact prompt. Compare shapes on the head-to-head page.
Gaps Enterprise pick
Share of Arize Phoenix's mentions where the AI models name this brand in the same answer. That is the real competitive set in the AI channel.
Scores are measured from real AI answers, refreshed monthly. Methodology.
We rerun the index every month. Drop your email and pick what to watch in LLM & Agent Evals: the whole category or specific brands. Free, unsubscribe anytime.