AI Tools

Coding & Developer Tools

Every Coding & Developer Tools tool in the catalog, with its vendor, home jurisdiction, and the governance question it raises.

30 tools · 28 vendors · 9 jurisdictions · 6 full profiles

In short

Coding & Developer Tools covers 30 tools from 28 vendors across 9 jurisdictions. 6 carry a full SRJ profile covering pricing tiers, data terms, and the governance question the tool raises; the rest link to the vendor's own product page. Live model rankings from arena.ai (formerly LMSYS Chatbot Arena) sit below, refreshed from the source rather than typed by hand. This is a reference list to inventory against, not a recommendation to buy.

Current model rankings

Rankings below come from arena.ai (formerly LMSYS Chatbot Arena), fetched automatically and stamped with the date they were true. Model leaderboards move weekly, so any figure typed into a page by hand is wrong within days. These are not.

As of 2026-08-13 Source: arena.ai (formerly LMSYS Chatbot Arena)

Read this before reading the tables

Crowdsourced blind pairwise voting, scored with a Bradley-Terry model and reported as an Elo-style rating with a 95 percent confidence interval. Gaps under roughly 10 points sit inside the noise floor and should not be read as a ranking.

Code generation

The same vote mechanic restricted to coding prompts. Ranks differ sharply from the text board, which is the reason to read both.

# Model Vendor Rating ± 95% CI
1 claude-opus-5-max Anthropic 1,691 ± 10
2 kimi-k3-max Moonshot 1,674 ± 11
3 qwen3.8-max Alibaba 1,669 ± 14
4 claude-opus-5-high Anthropic 1,664 ± 9
5 claude-fable-5 Anthropic 1,627 ± 9
6 gpt-5.6-sol-xhigh (codex-harness) OpenAI 1,622 ± 8
7 grok-4.6-high SpaceXAI 1,618 ± 21
8 glm-5.2-max Z.ai 1,587 ± 8
9 deepseek-v4-flash-high DeepSeek 1,582 ± 12
10 claude-opus-4-8-high Anthropic 1,564 ± 7
11 claude-opus-4-7 Anthropic 1,558 ± 6
12 claude-opus-4-7-high Anthropic 1,557 ± 6

Top 12 of 50 ranked. Full board at arena.ai (formerly LMSYS Chatbot Arena) →

A leaderboard measures average preference across a crowd of prompts that are not your prompts. It says nothing about your data terms, your jurisdiction, your latency budget, or your integration surface, and those usually decide procurement. Use it to shortlist, then evaluate on your own workload. Benchmarks beyond these boards, including MMLU-Pro, SWE-bench Verified, GPQA Diamond, and throughput figures, are deliberately not reproduced here: no verified machine-readable source for them is wired up yet, and an unsourced number on this site would be worse than a missing one.

All 30 Coding & Developer Tools tools

  • Anthropic Console Anthropic, US API management console; prompt testing; evaluation tools; usage monitoring; cost controls
  • AutoGen (Microsoft) Microsoft, US Multi-agent conversation framework; agent orchestration; code generation; Microsoft Research
  • Bolt.new StackBlitz, US App generation
  • Chroma Chroma, US Open-source embedding database; local-first; lightweight RAG; self-hosted option
  • Claude Code Profile Anthropic, US Agentic coding; source-code exposure
  • Comet ML Comet, US ML experiment management; model monitoring; production tracking; team collaboration
  • CrewAI CrewAI, US Multi-agent framework; agent collaboration; task delegation; emerging governance patterns
  • Cursor Profile Anysphere, US AI IDE; source-code exposure
  • Dify Dify, China LLM app development platform; workflow orchestration; open-source; cross-border data review advised
  • DSPy Stanford NLP, US Prompt optimization framework; Stanford research; systematic prompt engineering; reproducibility
  • Flowise Flowise, Open Source Visual LLM workflow builder; low-code RAG; self-hosted; no-code agent building
  • GitHub Copilot Profile Microsoft/GitHub, US Code completion; source-code exposure
  • Google AI Studio Google, US Model prototyping
  • Hugging Face Profile Hugging Face, US Model hub; supply-chain/AIBOM relevant
  • LangChain Profile LangChain, US LLM application framework; agent orchestration; supply-chain risk; rapid breaking changes
  • LlamaIndex Profile LlamaIndex, US RAG framework; data connectors; indexing strategies; enterprise RAG governance
  • Lovable Lovable, Sweden App generation
  • Neptune.ai Neptune, Poland ML metadata store; experiment tracking; model registry; EU-hosted options
  • OpenAI API / GPT models OpenAI, US Direct API; usage-based cost escalation risk
  • OpenAI Playground / Dashboard OpenAI, US API testing environment; fine-tuning; usage tracking; team management; cost escalation risk
  • Pinecone Pinecone, US Vector database; RAG infrastructure; data residency options; enterprise security
  • Relevance AI Relevance AI, Australia No-code AI workforce; agent teams; task automation; APAC data residency
  • Replit Agent Replit, US Agentic coding; source-code exposure
  • Stack AI Stack AI, US No-code AI app builder; enterprise deployment; workflow automation; data connector governance
  • Tabnine Tabnine, Israel Code completion; self-host option
  • v0 (Vercel) Vercel, US UI generation
  • Voiceflow Voiceflow, Canada Conversational AI design; voice/chat agent builder; enterprise team features
  • Weaviate Weaviate, Netherlands Open-source vector DB; GraphQL interface; EU-hosted; GDPR; hybrid search
  • Weights & Biases Weights & Biases, US ML experiment tracking; model registry; artifact lineage; MLOps governance
  • Windsurf Cognition, US AI IDE; source-code exposure

Frequently asked questions

How many Coding & Developer Tools tools are there?

This catalog lists 30 tools in the Coding & Developer Tools category, from 28 vendors across 9 jurisdictions. 6 carry a full SRJ profile; the rest link to the vendor's own product page.

Which model ranks highest right now?

On the code generation board at arena.ai (formerly LMSYS Chatbot Arena), as of 2026-08-13, the leader is claude-opus-5-max from Anthropic on 1691 Elo. Rankings move weekly, so the table on this page is refreshed from the source automatically rather than typed by hand. Read the confidence intervals before treating a small gap as a difference.

Do benchmark rankings tell you which model to buy?

No. A leaderboard measures average preference across a crowd of prompts that are not your prompts. It says nothing about your data terms, your jurisdiction, your latency budget, or your integration surface, and those usually decide procurement. Use the rankings to shortlist, then evaluate on your own workload against your own governance requirements.

Does inclusion in this category mean SRJ recommends the tool?

No. This is a reference catalog, not an endorsement list. Presence here means only that a tool is common enough in the field that it belongs on an inventory checklist. It is not a security review, a due-diligence result, or a recommendation to buy.

Why does the catalog list each vendor's home jurisdiction?

Jurisdiction is a governance fact. Where a vendor is based determines which data-transfer, privacy, and cross-border rules apply before any business data reaches the tool. Two tools that look identical in features can carry very different obligations once the vendor's home country is taken into account. This category spans 9 jurisdictions.

Other categories

← Back to the full catalog

You cannot govern what you have not listed

The AI Business Enablement Audit™ builds the inventory, measures your organization against every framework in the AI Governance Reference Library, and delivers a defensible governance dossier.

Start or finish your AI Audit →