Searches every page: governance library, books, services, glossary, tools, insights.
AI Tools
Every Chat & General LLMs tool in the catalog, with its vendor, home jurisdiction, and the governance question it raises.
42 tools · 42 vendors · 7 jurisdictions · 9 full profiles
In short
Chat & General LLMs covers 42 tools from 42 vendors across 7 jurisdictions. 9 carry a full SRJ profile covering pricing tiers, data terms, and the governance question the tool raises; the rest link to the vendor's own product page. Live model rankings from arena.ai (formerly LMSYS Chatbot Arena) sit below, refreshed from the source rather than typed by hand. This is a reference list to inventory against, not a recommendation to buy.
Rankings below come from arena.ai (formerly LMSYS Chatbot Arena), fetched automatically and stamped with the date they were true. Model leaderboards move weekly, so any figure typed into a page by hand is wrong within days. These are not.
Read this before reading the tables
Crowdsourced blind pairwise voting, scored with a Bradley-Terry model and reported as an Elo-style rating with a 95 percent confidence interval. Gaps under roughly 10 points sit inside the noise floor and should not be read as a ranking.
Head-to-head human preference on general conversation. The closest thing the field has to a general-purpose ranking.
| # | Model | Vendor | Rating | ± 95% CI |
|---|---|---|---|---|
| 1 | claude-fable-5 | Anthropic | 1,507 | ± 5 |
| 2 | claude-opus-4-6-high | Anthropic | 1,505 | ± 4 |
| 3 | claude-opus-4-7-high | Anthropic | 1,502 | ± 4 |
| 4 | muse-spark-1.2 (xHigh) | Meta | 1,499 | ± 10 |
| 5 | claude-opus-4-6 | Anthropic | 1,497 | ± 3 |
| 6 | claude-opus-4-7 | Anthropic | 1,494 | ± 4 |
| 7 | claude-opus-5-high | Anthropic | 1,494 | ± 5 |
| 8 | claude-opus-5-max | Anthropic | 1,491 | ± 7 |
| 9 | qwen3.8-max | Alibaba | 1,491 | ± 8 |
| 10 | muse-spark-1.1 | Meta | 1,489 | ± 6 |
| 11 | kimi-k3-max | Moonshot | 1,489 | ± 6 |
| 12 | muse-spark | Meta | 1,488 | ± 6 |
Top 12 of 30 ranked. Full board at arena.ai (formerly LMSYS Chatbot Arena) →
Multi-step tool-using tasks rather than single answers. The board that matters if the model will act rather than reply. Published as an order only, with no ratings.
| # | Model | Vendor |
|---|---|---|
| 1 | Claude Fable 5 (High) | Anthropic |
| 2 | Claude Opus 5 (High) | Anthropic |
| 3 | Claude Opus 5 (Max) | Anthropic |
| 4 | GPT 5.6 Sol (xHigh) | OpenAI |
| 5 | Kimi K3 (Max) | Moonshot |
| 6 | Claude Opus 4.8 (High) | Anthropic |
| 7 | GPT 5.5 (xHigh) | OpenAI |
| 8 | Claude Opus 4.7 (High) | Anthropic |
| 9 | GPT 5.5 (High) | OpenAI |
| 10 | Claude Opus 4.7 | Anthropic |
Top 10 of 10 ranked. This board is published as an order only, with no ratings, so no rating column is shown rather than one being invented. Full board at arena.ai (formerly LMSYS Chatbot Arena) →
Image understanding and mixed text-image prompts.
| # | Model | Vendor | Rating | ± 95% CI |
|---|---|---|---|---|
| 1 | claude-fable-5 | Anthropic | 1,315 | ± 9 |
| 2 | qwen3.8-max | Alibaba | 1,301 | ± 9 |
| 3 | claude-opus-4-7-high | Anthropic | 1,301 | ± 7 |
| 4 | claude-opus-4-6-high | Anthropic | 1,300 | ± 7 |
| 5 | claude-opus-4-7 | Anthropic | 1,299 | ± 7 |
| 6 | claude-opus-5-high | Anthropic | 1,297 | ± 11 |
| 7 | gemini-3.6-flash-high | 1,295 | ± 18 | |
| 8 | muse-spark | Meta | 1,294 | ± 9 |
| 9 | claude-opus-4-6 | Anthropic | 1,293 | ± 7 |
| 10 | muse-spark-1.2 (xHigh) | Meta | 1,290 | ± 18 |
| 11 | gemini-3-pro | 1,289 | ± 8 | |
| 12 | gpt-5.5 | OpenAI | 1,286 | ± 7 |
Top 12 of 20 ranked. Full board at arena.ai (formerly LMSYS Chatbot Arena) →
Models answering with live retrieval, where citation quality matters as much as fluency.
| # | Model | Vendor | Rating | ± 95% CI |
|---|---|---|---|---|
| 1 | claude-opus-4-6-search | Anthropic | 1,253 | ± 5 |
| 2 | gpt-5.5-search | OpenAI | 1,240 | ± 5 |
| 3 | claude-fable-5 | Anthropic | 1,237 | ± 8 |
| 4 | claude-opus-4-7 | Anthropic | 1,233 | ± 5 |
| 5 | ernie-5.1 | Baidu | 1,226 | ± 10 |
| 6 | claude-sonnet-4-6-search | Anthropic | 1,221 | ± 5 |
| 7 | gemini-3.1-pro-grounding | 1,212 | ± 5 | |
| 8 | gemini-3-pro-grounding | 1,207 | ± 5 | |
| 9 | gpt-5.2-search | OpenAI | 1,206 | ± 6 |
| 10 | grok-4.20-multi-agent-beta-0309 | SpaceXAI | 1,205 | ± 5 |
| 11 | claude-opus-4-8 | Anthropic | 1,205 | ± 6 |
| 12 | gpt-5.1-search | OpenAI | 1,199 | ± 5 |
Top 12 of 32 ranked. Full board at arena.ai (formerly LMSYS Chatbot Arena) →
A leaderboard measures average preference across a crowd of prompts that are not your prompts. It says nothing about your data terms, your jurisdiction, your latency budget, or your integration surface, and those usually decide procurement. Use it to shortlist, then evaluate on your own workload. Benchmarks beyond these boards, including MMLU-Pro, SWE-bench Verified, GPQA Diamond, and throughput figures, are deliberately not reproduced here: no verified machine-readable source for them is wired up yet, and an unsourced number on this site would be worse than a missing one.
This catalog lists 42 tools in the Chat & General LLMs category, from 42 vendors across 7 jurisdictions. 9 carry a full SRJ profile; the rest link to the vendor's own product page.
On the overall text and chat board at arena.ai (formerly LMSYS Chatbot Arena), as of 2026-08-13, the leader is claude-fable-5 from Anthropic on 1507 Elo. Rankings move weekly, so the table on this page is refreshed from the source automatically rather than typed by hand. Read the confidence intervals before treating a small gap as a difference.
No. A leaderboard measures average preference across a crowd of prompts that are not your prompts. It says nothing about your data terms, your jurisdiction, your latency budget, or your integration surface, and those usually decide procurement. Use the rankings to shortlist, then evaluate on your own workload against your own governance requirements.
No. This is a reference catalog, not an endorsement list. Presence here means only that a tool is common enough in the field that it belongs on an inventory checklist. It is not a security review, a due-diligence result, or a recommendation to buy.
Jurisdiction is a governance fact. Where a vendor is based determines which data-transfer, privacy, and cross-border rules apply before any business data reaches the tool. Two tools that look identical in features can carry very different obligations once the vendor's home country is taken into account. This category spans 7 jurisdictions.
The AI Business Enablement Audit™ builds the inventory, measures your organization against every framework in the AI Governance Reference Library, and delivers a defensible governance dossier.
Start or finish your AI Audit →