AI Tools

Chat & General LLMs

Every Chat & General LLMs tool in the catalog, with its vendor, home jurisdiction, and the governance question it raises.

42 tools · 42 vendors · 7 jurisdictions · 9 full profiles

In short

Chat & General LLMs covers 42 tools from 42 vendors across 7 jurisdictions. 9 carry a full SRJ profile covering pricing tiers, data terms, and the governance question the tool raises; the rest link to the vendor's own product page. Live model rankings from arena.ai (formerly LMSYS Chatbot Arena) sit below, refreshed from the source rather than typed by hand. This is a reference list to inventory against, not a recommendation to buy.

Current model rankings

Rankings below come from arena.ai (formerly LMSYS Chatbot Arena), fetched automatically and stamped with the date they were true. Model leaderboards move weekly, so any figure typed into a page by hand is wrong within days. These are not.

As of 2026-08-13 Source: arena.ai (formerly LMSYS Chatbot Arena)

Read this before reading the tables

Crowdsourced blind pairwise voting, scored with a Bradley-Terry model and reported as an Elo-style rating with a 95 percent confidence interval. Gaps under roughly 10 points sit inside the noise floor and should not be read as a ranking.

Overall text and chat

Head-to-head human preference on general conversation. The closest thing the field has to a general-purpose ranking.

# Model Vendor Rating ± 95% CI
1 claude-fable-5 Anthropic 1,507 ± 5
2 claude-opus-4-6-high Anthropic 1,505 ± 4
3 claude-opus-4-7-high Anthropic 1,502 ± 4
4 muse-spark-1.2 (xHigh) Meta 1,499 ± 10
5 claude-opus-4-6 Anthropic 1,497 ± 3
6 claude-opus-4-7 Anthropic 1,494 ± 4
7 claude-opus-5-high Anthropic 1,494 ± 5
8 claude-opus-5-max Anthropic 1,491 ± 7
9 qwen3.8-max Alibaba 1,491 ± 8
10 muse-spark-1.1 Meta 1,489 ± 6
11 kimi-k3-max Moonshot 1,489 ± 6
12 muse-spark Meta 1,488 ± 6

Top 12 of 30 ranked. Full board at arena.ai (formerly LMSYS Chatbot Arena) →

Agentic use

Multi-step tool-using tasks rather than single answers. The board that matters if the model will act rather than reply. Published as an order only, with no ratings.

# Model Vendor
1 Claude Fable 5 (High) Anthropic
2 Claude Opus 5 (High) Anthropic
3 Claude Opus 5 (Max) Anthropic
4 GPT 5.6 Sol (xHigh) OpenAI
5 Kimi K3 (Max) Moonshot
6 Claude Opus 4.8 (High) Anthropic
7 GPT 5.5 (xHigh) OpenAI
8 Claude Opus 4.7 (High) Anthropic
9 GPT 5.5 (High) OpenAI
10 Claude Opus 4.7 Anthropic

Top 10 of 10 ranked. This board is published as an order only, with no ratings, so no rating column is shown rather than one being invented. Full board at arena.ai (formerly LMSYS Chatbot Arena) →

Vision and multimodal

Image understanding and mixed text-image prompts.

# Model Vendor Rating ± 95% CI
1 claude-fable-5 Anthropic 1,315 ± 9
2 qwen3.8-max Alibaba 1,301 ± 9
3 claude-opus-4-7-high Anthropic 1,301 ± 7
4 claude-opus-4-6-high Anthropic 1,300 ± 7
5 claude-opus-4-7 Anthropic 1,299 ± 7
6 claude-opus-5-high Anthropic 1,297 ± 11
7 gemini-3.6-flash-high Google 1,295 ± 18
8 muse-spark Meta 1,294 ± 9
9 claude-opus-4-6 Anthropic 1,293 ± 7
10 muse-spark-1.2 (xHigh) Meta 1,290 ± 18
11 gemini-3-pro Google 1,289 ± 8
12 gpt-5.5 OpenAI 1,286 ± 7

Top 12 of 20 ranked. Full board at arena.ai (formerly LMSYS Chatbot Arena) →

A leaderboard measures average preference across a crowd of prompts that are not your prompts. It says nothing about your data terms, your jurisdiction, your latency budget, or your integration surface, and those usually decide procurement. Use it to shortlist, then evaluate on your own workload. Benchmarks beyond these boards, including MMLU-Pro, SWE-bench Verified, GPQA Diamond, and throughput figures, are deliberately not reproduced here: no verified machine-readable source for them is wired up yet, and an unsourced number on this site would be worse than a missing one.

All 42 Chat & General LLMs tools

  • 01.AI Yi 01.AI, China Chinese open-source LLM; Kai-Fu Lee founded; cross-border data review advised
  • AI2 OLMo Allen Institute for AI, US Open-source language model; fully open training data; research transparency; nonprofit
  • Aleph Alpha Aleph Alpha, Germany European sovereign LLM; EU-hosted; GDPR-first; multilingual; Luminous model family
  • Apple Intelligence Apple, US On-device LLM; Private Cloud Compute; Apple silicon; minimal data exposure; privacy-first
  • Baichuan Baichuan AI, China Chinese open-source LLM; medical and legal fine-tunes; cross-border data review advised
  • Cerebras Cerebras, US Wafer-scale AI compute; inference API; on-prem options; high-performance computing
  • Character.ai Character.AI, US Persona chat; HR/appropriate-use risk
  • ChatGPT Profile OpenAI, US General LLM assistant
  • Claude Profile Anthropic, US General LLM assistant
  • Cohere Coral Cohere, Canada Enterprise chat AI; RAG-native; data privacy focus; Canadian vendor
  • Databricks DBRX Databricks, US Open-source general LLM; Mixture-of-Experts; enterprise training; already noted in Data & Analytics
  • DeepSeek Profile DeepSeek AI, China Open-weight MoE; cross-border data review advised
  • Falcon (TII) Technology Innovation Institute, UAE UAE-hosted open-source LLM; sovereign AI; Apache 2.0; Middle East data residency
  • Fireworks AI Fireworks AI, US Fast inference API; open-source models; enterprise deployment; cost optimization
  • Gemini Profile Google, US General LLM assistant
  • Grok xAI, US General LLM assistant
  • Huawei PanGu Huawei, China Chinese foundation model; cross-border data review advised; Ascend chip optimization
  • HyperWrite HyperWrite, US Writing assistant + agent
  • iFlytek Spark iFlytek, China Chinese voice+language AI; speech recognition; cross-border data review advised
  • InternLM (Shanghai AI Lab) Shanghai AI Laboratory, China Chinese open-source LLM; academic research; cross-border data review advised
  • JAIS (G42 / Inception) G42 / Inception, UAE Arabic-centric LLM; UAE sovereign AI; cultural alignment; Middle East data governance
  • Janitor AI Janitor AI, US Persona chat; HR/appropriate-use risk
  • Kimi Moonshot AI, China Long-context reasoning; cross-border data review advised
  • Llama (Meta) Profile Meta, US Open-weight model family; often self-hosted
  • LMSYS Arena Models LMSYS, US Model evaluation platform; crowdsourced benchmarking; research transparency; no production use
  • Microsoft Copilot Profile Microsoft, US Embedded across M365
  • Mistral / Le Chat Profile Mistral AI, France EU-hosted option; GDPR relevant
  • Notion AI Profile Notion Labs, US Embedded in workspace/knowledge base
  • NVIDIA Nemotron NVIDIA, US Enterprise LLM family; synthetic data generation; GPU-optimized; self-hosted options
  • Perplexity AI Profile Perplexity, US Search-augmented LLM
  • Pi Inflection AI, US Conversational assistant
  • Poe Quora, US Multi-model aggregator
  • Replicate Replicate, US Model hosting and inference API; open-source models; community models; content policy
  • SambaNova SambaNova Systems, US Enterprise AI hardware+software; DataScale platform; on-prem deployment; sovereign AI
  • Samsung Gauss Samsung, South Korea On-device LLM; Galaxy AI; personal data processing; Korean vendor; local processing
  • SaulLM (Saul) Saul, France French legal-domain LLM; specialized training; EU data governance
  • SenseTime SenseChat SenseTime, China Chinese multimodal LLM; computer vision integration; cross-border data review advised
  • Slack AI Salesforce, US Embedded in messaging; message-corpus exposure
  • Snowflake Arctic Snowflake, US Enterprise-focused LLM; Apache 2.0 license; data cloud integration; commercial use
  • Together AI Together AI, US Decentralized cloud for open models; inference API; model hub; data residency options
  • You.com You.com, US Search-augmented LLM
  • Zoom AI Companion Zoom, US Meeting summarization; recording consent relevant

Frequently asked questions

How many Chat & General LLMs tools are there?

This catalog lists 42 tools in the Chat & General LLMs category, from 42 vendors across 7 jurisdictions. 9 carry a full SRJ profile; the rest link to the vendor's own product page.

Which model ranks highest right now?

On the overall text and chat board at arena.ai (formerly LMSYS Chatbot Arena), as of 2026-08-13, the leader is claude-fable-5 from Anthropic on 1507 Elo. Rankings move weekly, so the table on this page is refreshed from the source automatically rather than typed by hand. Read the confidence intervals before treating a small gap as a difference.

Do benchmark rankings tell you which model to buy?

No. A leaderboard measures average preference across a crowd of prompts that are not your prompts. It says nothing about your data terms, your jurisdiction, your latency budget, or your integration surface, and those usually decide procurement. Use the rankings to shortlist, then evaluate on your own workload against your own governance requirements.

Does inclusion in this category mean SRJ recommends the tool?

No. This is a reference catalog, not an endorsement list. Presence here means only that a tool is common enough in the field that it belongs on an inventory checklist. It is not a security review, a due-diligence result, or a recommendation to buy.

Why does the catalog list each vendor's home jurisdiction?

Jurisdiction is a governance fact. Where a vendor is based determines which data-transfer, privacy, and cross-border rules apply before any business data reaches the tool. Two tools that look identical in features can carry very different obligations once the vendor's home country is taken into account. This category spans 7 jurisdictions.

Other categories

← Back to the full catalog

You cannot govern what you have not listed

The AI Business Enablement Audit™ builds the inventory, measures your organization against every framework in the AI Governance Reference Library, and delivers a defensible governance dossier.

Start or finish your AI Audit →