Best AI LLM APIs & Models Tools
Foundation models, inference APIs, and model hosting platforms for developers building on LLMs
ChatGPT already recommends Claude Opus 4.8. Does it recommend yours?
These are the llm apis & models tools ChatGPT names when someone asks for a recommendation. If you build one and you're not on this list, run a free AI-visibility scan on your own product — 5 prompt angles, ~30 seconds, no signup, no card.
All LLM APIs & Models Tools (24)
Claude Opus 4.8
Anthropic's flagship model — stronger coding, agents, and honesty
Mistral Small 4
Mistral's unified open-source model — reasoning + vision + coding, Apache 2.0
Mistral Small 3.1
Mistral's 24B multimodal open-source model — beats GPT-4o Mini, Apache 2.0
Mistral Small 3
Mistral's 24B latency-optimized open model — faster than Llama 3.3 70B, Apache 2.0
Mistral Medium 3.5
Mistral's 128B merged flagship — open weights, coding+reasoning+instructions
Mistral 3
Mistral's MoE flagship + edge model family — Apache 2.0, multimodal, reasoning
North Mini Code
Cohere's open-source agentic coding model — 30B MoE, 3B active, Apache 2.0
Codestral 25.08
Mistral's low-latency code completion model — FIM, 80+ languages, 256k context
Codestral Embed
Mistral's code-specific embedding model — semantic code search and RAG for repos
Mistral OCR 3
Mistral's document OCR model — 74% win rate vs OCR 2, $2/1k pages, forms + handwriting
Devstral 2
Mistral's SOTA open-weight coding model — 72.2% SWE-bench, free API
Magistral
Mistral's first reasoning model — 73.6% AIME2024, native multilingual chain-of-thought, 10x faster via Le Chat
📬 Get the best new AI tools delivered weekly
One concise email with fresh launches, trending picks, and featured standouts.
Mistral Medium 3
Enterprise GPT-4o-class performance at $0.40/$2.00 per million tokens — Mistral's mid-tier powerhouse
Mistral Saba
Mistral's 24B regional LLM — Arabic & South Asian languages, 150+ tok/s, self-hostable
Mistral NeMo
Mistral × NVIDIA 12B open-weight model — 128k context, Tekken tokenizer, FP8 inference, Apache 2.0
Codestral Mamba
Mistral's 7B Mamba-architecture coding model — linear-time inference, 256k context, Apache 2.0
Mathstral 7B
Open-weight 7B math specialist from Mistral AI — STEM reasoning, MATH benchmark SOTA at release
Mixtral 8x22B
Mistral's largest open-weights MoE — 141B total / 39B active, Apache 2.0
Claude Fable 5
Anthropic's most capable model — state-of-the-art coding, vision, and long-horizon tasks
Understudy Labs
Captures traces from your production LLM work, then trains a cheaper open-weight model that beats the eval
Conifer
Local-first least-cost inference router that keeps routine requests on your own hardware
EvoLink
One OpenAI-compatible API key for 75+ LLM, image, video, and audio models with smart routing and failover
Yolo-Auto
Flat-rate, unmetered OpenAI-compatible API serving Qwen3.6-35B, priced by concurrency instead of tokens
Mood Metrics API
Real-time sentiment API with confidence scores, emotions and a financial mode
Want your tool featured here?
Featured tools appear first on this page and get surfaced to AI search engines like ChatGPT and Perplexity. Every plan includes a permanent dofollow backlink to your site.