Compare API providers and downloadable models in one searchable directory, with source-qualified token prices, licenses, context windows and self-hosting estimates.
API services and downloadable models
AI directory
28 of 28 source records · 18 API offers · 14 downloadable
Records retain separate source identities, not a deduplicated model count. API availability does not establish whether weights are open or closed.
Anthropic
Claude Fable 5
Input: $10USD per 1M tokens
Weights unknownAPI availableDownload not verifiedreasoningContext: Unknown tokens
Input: $10 · Cached input: $1 · Output: $50 — USD per 1M tokens. Unknown cache pricing is not zero.
The June 9, 2026 system card describes Claude Fable 5 as a general-access configuration with safeguards. General access and discussion of model-weight theft do not establish public weight-download availability.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
Anthropic
Claude Haiku 4.5
Input: $1USD per 1M tokens
Weights unknownAPI availableDownload not verifiedtextContext: Unknown tokens
Input: $1 · Cached input: $0.1 · Output: $5 — USD per 1M tokens. Unknown cache pricing is not zero.
The official Claude Haiku 4.5 overview identifies the model and lists hosted platforms; this review did not establish explicit weight-release or non-downloadability evidence.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
Anthropic
Claude Opus 5
Input: $5USD per 1M tokens
Weights unknownAPI availableDownload not verifiedreasoningContext: Unknown tokens
Input: $5 · Cached input: $0.5 · Output: $25 — USD per 1M tokens. Unknown cache pricing is not zero.
The official Claude Opus 5 overview identifies the model and lists hosted platforms; this review did not establish explicit weight-release or non-downloadability evidence.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
Anthropic
Claude Sonnet 5
Input: $2USD per 1M tokens
Weights unknownAPI availableDownload not verifiedreasoningContext: Unknown tokens
Input: $2 · Cached input: $0.2 · Output: $10 — USD per 1M tokens. Unknown cache pricing is not zero.
The official Claude Sonnet 5 overview identifies the model and lists hosted platforms; this review did not establish explicit weight-release or non-downloadability evidence.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
DeepSeek · DeepSeek R1
DeepSeek R1
No API price listedSelf-hosting costs not included
Open weightsDownloadabletextContext: 131,072 tokens
The official model card identifies Gemini 2.5 Flash-Lite and describes its capabilities and safety evaluations; this review did not establish explicit weight-release or non-downloadability evidence.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
This review applies to the named model, not a separate checkpoint inferred from the pricing record's prompt-length or input-modality tier. It does not refresh pricing or resolve unverified snapshot identity.
Google
Gemini 2.5 Pro (≤200K prompt)
Input: $1.25USD per 1M tokens
Weights unknownAPI availableDownload not verifiedreasoningContext: Unknown tokens
Input: $1.25 · Cached input: $0.125 · Output: $10 — USD per 1M tokens. Unknown cache pricing is not zero.
The official model card identifies Gemini 2.5 Pro and describes its capabilities and safety evaluations; this review did not establish explicit weight-release or non-downloadability evidence.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
This review applies to the named model, not a separate checkpoint inferred from the pricing record's prompt-length or input-modality tier. It does not refresh pricing or resolve unverified snapshot identity.
Google
Gemini 3.1 Pro Preview (≤200K prompt)
Input: $2USD per 1M tokens
Weights unknownAPI availableDownload not verifiedreasoningContext: Unknown tokens
Input: $2 · Cached input: $0.2 · Output: $12 — USD per 1M tokens. Unknown cache pricing is not zero.
The official model card for Gemini 3.1 Pro describes distribution through hosted applications and APIs; those channels alone do not establish whether weights are downloadable.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
This review applies to the named model, not a separate checkpoint inferred from the pricing record's prompt-length or input-modality tier. It does not refresh pricing or resolve unverified snapshot identity.
Google
Gemini 3.5 Flash-Lite
Input: $0.3USD per 1M tokens
Weights unknownAPI availableDownload not verifiedmultimodalContext: Unknown tokens
Input: $0.3 · Cached input: $0.03 · Output: $2.5 — USD per 1M tokens. Unknown cache pricing is not zero.
The official model card for Gemini 3.5 Flash-Lite describes distribution through hosted applications and APIs; those channels alone do not establish whether weights are downloadable.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
Google
Gemini 3.7 Flash
Input: $0.75USD per 1M tokens
Weights unknownAPI availableDownload not verifiedmultimodalContext: Unknown tokens
Input: $0.75 · Cached input: $0.075 · Output: $3.75 — USD per 1M tokens. Unknown cache pricing is not zero.
The official model card for Gemini 3.7 Flash describes distribution through hosted applications and APIs; those channels alone do not establish whether weights are downloadable.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
Google · Gemma 3
Gemma 3 1B IT
No API price listedSelf-hosting costs not included
Open weightsDownloadabletextContext: 32,768 tokens
The official GPT-5.4 mini model page describes API endpoints, features and model aliases; this review did not establish explicit weight-release or non-downloadability evidence.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
OpenAI
GPT-5.4 nano
Input: $0.2USD per 1M tokens
Weights unknownAPI availableDownload not verifiedtextContext: Unknown tokens
Input: $0.2 · Cached input: $0.02 · Output: $1.25 — USD per 1M tokens. Unknown cache pricing is not zero.
The official GPT-5.4 nano model page describes API endpoints, features and model aliases; this review did not establish explicit weight-release or non-downloadability evidence.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
OpenAI
GPT-5.6 Luna
Input: $0.2USD per 1M tokens
Weights unknownAPI availableDownload not verifiedreasoningContext: Unknown tokens
Input: $0.2 · Cached input: $0.02 · Output: $1.2 — USD per 1M tokens. Unknown cache pricing is not zero.
The official GPT-5.6 Luna model page describes API endpoints, features and model aliases; this review did not establish explicit weight-release or non-downloadability evidence.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
OpenAI
GPT-5.6 Sol
Input: $5USD per 1M tokens
Weights unknownAPI availableDownload not verifiedreasoningContext: Unknown tokens
Input: $5 · Cached input: $0.5 · Output: $30 — USD per 1M tokens. Unknown cache pricing is not zero.
The official GPT-5.6 Sol model page describes API endpoints, features and model aliases; this review did not establish explicit weight-release or non-downloadability evidence.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
OpenAI
GPT-5.6 Terra
Input: $2USD per 1M tokens
Weights unknownAPI availableDownload not verifiedreasoningContext: Unknown tokens
Input: $2 · Cached input: $0.2 · Output: $12 — USD per 1M tokens. Unknown cache pricing is not zero.
The official GPT-5.6 Terra model page describes API endpoints, features and model aliases; this review did not establish explicit weight-release or non-downloadability evidence.
Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
Moonshot AI
Kimi K2.6
Input: $0.95USD per 1M tokens
Open weightsAPI availableDownloadablereasoningContext: 262,144 tokens
Input: $0.95 · Cached input: $0.16 · Output: $4 — USD per 1M tokens. Unknown cache pricing is not zero.
The creator publishes the model weights under the Modified MIT license.
Open weights are downloadable under the model license; this is not a claim of unrestricted open-source licensing or that the model fits consumer hardware.
Moonshot AI
Kimi K2.7 Code
Input: $0.95USD per 1M tokens
Open weightsAPI availableDownloadablereasoningContext: 262,144 tokens
Input: $0.95 · Cached input: $0.19 · Output: $4 — USD per 1M tokens. Unknown cache pricing is not zero.
The creator publishes the model weights under the Modified MIT license.
Open weights are downloadable under the model license; this is not a claim of unrestricted open-source licensing or that the model fits consumer hardware.
Moonshot AI
Kimi K2.7 Code Highspeed
Input: $1.9USD per 1M tokens
Open weightsAPI availableDownloadablereasoningContext: 262,144 tokens
Input: $1.9 · Cached input: $0.38 · Output: $8 — USD per 1M tokens. Unknown cache pricing is not zero.
The provider explicitly describes HighSpeed as the same model as Kimi K2.7 Code with a faster serving tier.
The download link is the underlying Kimi K2.7 Code model, not a separate HighSpeed checkpoint; self-hosting does not reproduce the hosted speed guarantee.
Kimi K2.7 Code weights use the Modified MIT license, as documented in the linked creator model card.
Moonshot AI
Kimi K3
Input: $3USD per 1M tokens
Open weightsAPI availableDownloadablereasoningContext: 1,048,576 tokens
Input: $3 · Cached input: $0.3 · Output: $15 — USD per 1M tokens. Unknown cache pricing is not zero.
The creator publishes the model weights under the Kimi K3 License license.
Open weights are downloadable under the model license; this is not a claim of unrestricted open-source licensing or that the model fits consumer hardware.
Meta · Llama 3.2
Llama 3.2 11B Vision
No API price listedSelf-hosting costs not included
Open weightsDownloadabletext · imageContext: 131,072 tokens
Mixture-of-experts model with 3.3B parameters activated per token; weight memory reflects 30.5B total parameters. Validated to 131,072 tokens with YaRN.
Read before comparing
The same unit is not the same offer.
API prices are standard paid-tier USD per one million tokens. Batch, priority, negotiated and cloud-marketplace pricing are excluded. Prompt length, input modality, cache writes, geography and service tier can change a listed rate. Output charges can include reasoning or thinking tokens. Keep each record's source conditions in view; price sorting is not a quality ranking.
Unknown cached-input prices are not zero. Download-only records have no fabricated API price. Open weights do not automatically mean an OSI-approved open-source license; review the license and any gated download terms.
Q4 and Q8 estimates use 5 and 9 bits per weight respectively, plus 20% loading overhead in decimal GB. These are weight-capacity estimates, not performance benchmarks. KV cache, activations, runtime scratch space, drivers and backend compatibility require additional planning. See the methodology and source register.
Estimated local deployment
How model size maps to accelerator memory
Estimated weights plus 20% loading overhead. Runtime, KV cache, context length, and software can require substantially more memory.
Estimated 4-bit weight memoryGB, including 20% overhead · logarithmic bar scale
The estimate uses parameters × planning bits ÷ 8 × 1.20: 5 planning bits for Q4 and 9 for Q8 to allow for scales, metadata, and mixed-precision tensors. “Comfortable” reserves at least 20% of device memory. “Tight” means estimated weights fit but leave less headroom.
This is a capacity screen—not a speed benchmark or guarantee that a particular runtime, context length, multimodal projector, or GPU configuration will work.
Estimated 4-bit model weight fit by available accelerator memory