API services and downloadable models

AI directory

64 of 64 source records · 44 API offers · 14 downloadable

Records retain separate source identities, not a deduplicated model count. API availability does not establish whether weights are open or closed.

Name order groups only evidence-linked checkpoint records and seller offers. Price order lists individual records; matching names never establish a shared checkpoint.

Anthropic

Claude Fable 5

Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedreasoningContext: Unknown tokens

Input: USD 10 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache write: USD 12.50 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: 300 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache write: USD 20 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: 3600 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache read: USD 1 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Cache hits and refreshes: Same duration as the preceding write; no single read TTL is asserted.

Output: USD 50 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Output tokens as listed; separate reasoning/thinking billing treatment is unknown in this source.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Anthropic

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-29T04:37:41.581Z · Price source checked: 2026-09-29T04:37:41.581Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: identity · anthropic-claude-fable-5: pricing · anthropic-claude-fable-5: availability · anthropic-claude-fable-5: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Cached input is the cache-hit/refresh rate; cache writes cost more. US-only inference is 1.1×.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

  • The June 9, 2026 system card describes Claude Fable 5 as a general-access configuration with safeguards. General access and discussion of model-weight theft do not establish public weight-download availability.
  • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T05:34:42.597Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:23.592Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:19.422Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:41.581Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 245b5d0472659075015551126902de479990475fa61deb58c5fc20ef5e51a909; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Anthropic

Claude Fable 5.1

Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

Input: USD 10 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache write: USD 12.50 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: 300 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache write: USD 20 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: 3600 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache read: USD 0.25 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Cache hits and refreshes: Same duration as the preceding write; no single read TTL is asserted.
  • Cache hits and refreshes on Claude Fable 5.1 are priced at 0.025x the base input price (other models: 0.1x).

Output: USD 50 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Output tokens as listed; separate reasoning/thinking billing treatment is unknown in this source.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Anthropic

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-29T04:37:41.581Z · Price source checked: 2026-09-29T04:37:41.581Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T05:34:42.597Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:23.592Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:19.422Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:41.581Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 245b5d0472659075015551126902de479990475fa61deb58c5fc20ef5e51a909; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Anthropic

Claude Haiku 4.5

Input: USD 1 / 1000000 tokens · StandardOutput: USD 5 / 1000000 tokens · StandardHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedtextContext: Unknown tokens

Input: USD 1 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Haiku 4.5 context pricing threshold is not established by this source.
  • Haiku 4.5 does not support inference_geo; standard pricing applies.

Cache write: USD 1.25 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: 300 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Haiku 4.5 context pricing threshold is not established by this source.
  • Haiku 4.5 does not support inference_geo; standard pricing applies.

Cache write: USD 2 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: 3600 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Haiku 4.5 context pricing threshold is not established by this source.
  • Haiku 4.5 does not support inference_geo; standard pricing applies.

Cache read: USD 0.10 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Haiku 4.5 context pricing threshold is not established by this source.
  • Haiku 4.5 does not support inference_geo; standard pricing applies.
  • Cache hits and refreshes: Same duration as the preceding write; no single read TTL is asserted.

Output: USD 5 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Haiku 4.5 context pricing threshold is not established by this source.
  • Haiku 4.5 does not support inference_geo; standard pricing applies.
  • Output tokens as listed; separate reasoning/thinking billing treatment is unknown in this source.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Anthropic

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-29T04:37:41.581Z · Price source checked: 2026-09-29T04:37:41.581Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: anthropic-claude-haiku-4-5: identity · anthropic-claude-haiku-4-5: pricing · anthropic-claude-haiku-4-5: availability · anthropic-claude-haiku-4-5: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Cached input is the cache-hit/refresh rate; cache writes cost more.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-haiku-4-5: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

  • The official Claude Haiku 4.5 overview identifies the model and lists hosted platforms; this review did not establish explicit weight-release or non-downloadability evidence.
  • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T05:34:42.597Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:23.592Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:19.422Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:41.581Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 245b5d0472659075015551126902de479990475fa61deb58c5fc20ef5e51a909; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Anthropic

Claude Opus 5

Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedreasoningContext: Unknown tokens

Input: USD 5 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache write: USD 6.25 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: 300 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache write: USD 10 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: 3600 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache read: USD 0.50 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Cache hits and refreshes: Same duration as the preceding write; no single read TTL is asserted.

Output: USD 25 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Output tokens as listed; separate reasoning/thinking billing treatment is unknown in this source.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Anthropic

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-29T04:37:41.581Z · Price source checked: 2026-09-29T04:37:41.581Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: anthropic-claude-opus-5: identity · anthropic-claude-opus-5: pricing · anthropic-claude-opus-5: availability · anthropic-claude-opus-5: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Cached input is the cache-hit/refresh rate; Fast mode and cache writes have separate prices.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-opus-5: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

  • The official Claude Opus 5 overview identifies the model and lists hosted platforms; this review did not establish explicit weight-release or non-downloadability evidence.
  • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T05:34:42.597Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:23.592Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:19.422Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:41.581Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 245b5d0472659075015551126902de479990475fa61deb58c5fc20ef5e51a909; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Anthropic

Claude Opus 5.5

Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

Input: USD 4 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache write: USD 5 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: 300 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache write: USD 8 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: 3600 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.

Cache read: USD 0.20 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Cache hits and refreshes: Same duration as the preceding write; no single read TTL is asserted.
  • Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price (other models: 0.1x).

Output: USD 20 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Output tokens as listed; separate reasoning/thinking billing treatment is unknown in this source.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Anthropic

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-29T04:37:41.581Z · Price source checked: 2026-09-29T04:37:41.581Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T05:34:42.597Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:23.592Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:19.422Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:41.581Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 245b5d0472659075015551126902de479990475fa61deb58c5fc20ef5e51a909; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Anthropic

Claude Sonnet 5

Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedreasoningContext: Unknown tokens

Input: USD 2 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Sonnet 5 launch pricing is now standard, not an expiring promotion.

Cache write: USD 2.50 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: 300 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Sonnet 5 launch pricing is now standard, not an expiring promotion.

Cache write: USD 4 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: 3600 seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Sonnet 5 launch pricing is now standard, not an expiring promotion.

Cache read: USD 0.20 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Cache hits and refreshes: Same duration as the preceding write; no single read TTL is asserted.
  • Sonnet 5 launch pricing is now standard, not an expiring promotion.

Output: USD 10 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Claude API direct global Standard only; excludes Batch, Fast, partner billing, negotiated discounts and regional premiums.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.
  • US-only inference_geo has a 1.1x multiplier; excluded from this global rate.
  • Output tokens as listed; separate reasoning/thinking billing treatment is unknown in this source.
  • Sonnet 5 launch pricing is now standard, not an expiring promotion.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Anthropic

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-29T04:37:41.581Z · Price source checked: 2026-09-29T04:37:41.581Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: anthropic-claude-sonnet-5: identity · anthropic-claude-sonnet-5: pricing · anthropic-claude-sonnet-5: availability · anthropic-claude-sonnet-5: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Cached input is the cache-hit/refresh rate; US-only inference is 1.1×.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-sonnet-5: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

  • The official Claude Sonnet 5 overview identifies the model and lists hosted platforms; this review did not establish explicit weight-release or non-downloadability evidence.
  • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T05:34:42.597Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:23.592Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:19.422Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: cd36fcc6d6d9e17ac2a8643a3f769e42ec6cf33a8becdbfd98bc3ba8832f1aee; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Primary source ↗ · vendor-docs · Claims: anthropic-claude-fable-5: pricing · anthropic-claude-haiku-4-5: pricing · anthropic-claude-opus-5: pricing · anthropic-claude-sonnet-5: pricing · anthropic-claude-opus-5-5: pricing · anthropic-claude-fable-5-1: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:41.581Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 245b5d0472659075015551126902de479990475fa61deb58c5fc20ef5e51a909; collector anthropic-markdown-2.
  • Pricing observation only. Display-name mapping is not official API identity evidence. No new access, availability, weight, retirement or model metadata assertion.
  • This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. For the most current pricing information, visit claude.com/pricing.
  • The following table shows pricing for all Claude models: 1 Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. 2 Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price. All other models use the standard 0.1x multiplier. 3 The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur. MTok: Million tokens. $5 / MTok is $5 for every million tokens. 5m cache writes: Writing a prompt prefix to the 5-minute prompt cache. 1h cache writes: Writing a prompt prefix to the 1-hour prompt cache. Cache hits and refreshes: Reading a prompt prefix from the prompt cache, which also refreshes it. Limited access: Offered separately, by invitation only, as part of Project Glasswing. For access, contact your Anthropic, AWS, or Google Cloud account team. Retired: May still be available on other cloud platforms. See Model deprecations for more. Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
  • Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. Instead of reprocessing the same large system prompt, document, or conversation history on every request, the API reads from cache at a fraction of the standard input price. There are two ways to enable prompt caching: Automatic caching: Add a single cache_control field at the top level of your request. The system automatically manages cache breakpoints as conversations grow. This is the recommended starting point for most use cases. Explicit cache breakpoints: Place cache_control directly on individual content blocks for fine-grained control over exactly what gets cached. Prompt caching uses the following pricing multipliers relative to base input token rates: Cache operation Multiplier Duration 5-minute cache write 1.25x base input price Cache valid for 5 minutes 1-hour cache write 2x base input price Cache valid for 1 hour Cache read (hit) 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1; 0.05x on Claude Opus 5.5) Same duration as the preceding write Cache write tokens are charged when content is first stored. Cache read tokens are charged when a subsequent request retrieves the cached content. A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write). On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens). On Claude Opus 5.5, a cache hit costs 5% of the standard input price ($0.20 USD per million tokens). These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. For implementation details, supported models, and code examples, see Prompt caching.
  • For Claude 4.6 and later models, specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories, including input tokens, output tokens, cache writes, and cache reads. Global routing (the default) uses standard pricing. This applies to the Claude API (first-party) and Claude Platform on AWS. On Claude in Microsoft Foundry, the same 1.1x multiplier applies to deployments that use the US Data Zone Standard deployment type (see Inference geography). Partner-operated platforms (Bedrock and Google Cloud) have independent regional pricing. See Bedrock and Google Cloud for details. Earlier models do not support the inference_geo parameter and always use standard pricing; requests that include the parameter on these models return a 400 error. For more information, see Data residency.
  • Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window.

Checkpoint: DeepSeek R1

DeepSeek · DeepSeek R1

DeepSeek R1

No API price listedSelf-hosting costs not included
Open weightsDownloadabletextContext: Unknown tokens

Weight license: MIT · Hugging Face / download weights ↗

671B parameters · Estimated weight memory: Q4 503.3 GB · Q8 905.9 GB. Includes 20% loading overhead; excludes KV cache and runtime memory. Compare hardware.

Developer: DeepSeek

Checkpoint: DeepSeek R1 · Revision: 56d4cbbb4d29f4355bab4b9a39ccb717a14ad5ad · Lineage: Unknown / not verified

Creator model source ↗ Creator model source ↗ Weight-license evidence ↗ Weight-license evidence ↗

Model metadata observed: 2026-09-27T00:03:38.552Z / 2026-09-27T00:03:44.318Z · Metadata freshness: fresh

Weight-license reviewed: 2026-09-27T02:50:26.221Z. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: ungated · Code license: MIT · API availability: unknown (dated observation, not a live service check)

Active parameters: 37B; total weights: 671B.

Source, conditions and dates

Primary source ↗ · huggingface-model-card · Claims: deepseek-r1: identity · deepseek-r1: specifications · deepseek-r1: weightAccess · deepseek-r1: download · deepseek-r1: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

  • Full reasoning model, not a distilled checkpoint; 37B parameters are activated per token but all 671B weights must be stored.

Latest result: held. Retained evidence keeps its original observation and successful-check dates.

Primary source ↗ · huggingface-model-card · Claims: deepseek-v3: identity · deepseek-v3: specifications · deepseek-v3: weightAccess · deepseek-v3: download · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

  • Mixture-of-experts model with 37B parameters activated per token; local weight memory reflects 671B total parameters.

Latest result: held. Retained evidence keeps its original observation and successful-check dates.

Primary source ↗ · huggingface-model-card · Claims: deepseek-r1: identity · deepseek-r1: specifications · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:03:38.552Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Card specifies671B total and37B active model parameters; API tensor total684,489,845,504 has a different scope. Memory estimates retain total, not active, model parameters.
  • Card advertises128K context while config specifies163,840 positions. Exact supported token count remains unknown. Named distilled derivatives and their base-model/license exceptions are separate checkpoints; no lineage is inferred for this exact R1 from their table.
  • Metadata documents were read; weight binaries were not downloaded. Prior weight-access/download observations retain their original dates.
  • Revision 56d4cbbb4d29f4355bab4b9a39ccb717a14ad5ad; SHA256 d26d26ddb518fee60c6c6bf7a708bd751b1619d93a4944f188143693d956c77f; private manifest ai/v1/.model-metadata-development/36281114084/hf-deepseek-r1-readme-md/2026-09-27/complete.json.gz; manifest SHA256 9c64fafbfe74d6daed827efaa0f4cae55026bb5b8f3761fb4dd2866cf3482cc1.

Raw API evidence ↗ · vendor-api · Claims: deepseek-r1: gating · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:03:32.726Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Repository gating: false; listed weight files: true. An inventory and authenticated metadata read do not verify weight download availability.
  • Revision 56d4cbbb4d29f4355bab4b9a39ccb717a14ad5ad; SHA256 1bd51a07bca0d0706545c7bbaa21167e1cd5dc1713d84fb4793b00c8f8fcc101; private manifest ai/v1/.model-metadata-development/36281114084/hf-deepseek-r1-info/2026-09-27/complete.json.gz; manifest SHA256 db09a10a0492ea318429b6bcce310f151478d50fff55eb7456ec5c0a1006c9e6.

Primary source ↗ · vendor-docs · Claims: deepseek-r1: configuration · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:03:44.318Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Configuration is supporting evidence, not an automatic claim of advertised context or modality. Resolve differences using the explicit publication review.
  • Revision 56d4cbbb4d29f4355bab4b9a39ccb717a14ad5ad; SHA256 79ddea672a62e95d3f0be27be434375538e1988975971cffcf87c96c1de84a65; private manifest ai/v1/.model-metadata-development/36281114084/hf-deepseek-r1-config-json/2026-09-27/complete.json.gz; manifest SHA256 513d8f434a9a464410ff3796d8d3725bae4a7c79efac637bc789d61e55f3bd64.

Primary source ↗ · huggingface-model-card · Claims: deepseek-r1: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:03:38.552Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Exact R1 card explicitly licenses both code and weights under MIT; retain copyright/license notice. Distilled derivatives have separate exceptions and are not this checkpoint.
  • Revision 56d4cbbb4d29f4355bab4b9a39ccb717a14ad5ad; SHA256 d26d26ddb518fee60c6c6bf7a708bd751b1619d93a4944f188143693d956c77f; private manifest ai/v1/.model-metadata-development/36281114084/hf-deepseek-r1-readme-md/2026-09-27/complete.json.gz; manifest SHA256 9c64fafbfe74d6daed827efaa0f4cae55026bb5b8f3761fb4dd2866cf3482cc1.

Primary source ↗ · huggingface-model-card · Claims: deepseek-r1: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:03:50.109Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Exact R1 card explicitly licenses both code and weights under MIT; retain copyright/license notice. Distilled derivatives have separate exceptions and are not this checkpoint.
  • Revision 56d4cbbb4d29f4355bab4b9a39ccb717a14ad5ad; SHA256 1c8f573e830ca9b3ebfeb7ace1823146e22b66f99ee223840e7637c9e745e1c7; private manifest ai/v1/.model-metadata-development/36281114084/hf-deepseek-r1-license/2026-09-27/complete.json.gz; manifest SHA256 f9459f6fcb94c0ff78e42e7ef950ebf63df65e62fa523a9fd138ceca46c40e7b.

Primary source ↗ · huggingface-model-card · Claims: deepseek-r1: codeLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:03:38.552Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Exact R1 card explicitly applies MIT to code and weights; preserve copyright/license notice.
  • Revision 56d4cbbb4d29f4355bab4b9a39ccb717a14ad5ad; SHA256 d26d26ddb518fee60c6c6bf7a708bd751b1619d93a4944f188143693d956c77f; private manifest ai/v1/.model-metadata-development/36281114084/hf-deepseek-r1-readme-md/2026-09-27/complete.json.gz; manifest SHA256 9c64fafbfe74d6daed827efaa0f4cae55026bb5b8f3761fb4dd2866cf3482cc1.

Primary source ↗ · huggingface-model-card · Claims: deepseek-r1: codeLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:03:50.109Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Exact R1 card explicitly applies MIT to code and weights; preserve copyright/license notice.
  • Revision 56d4cbbb4d29f4355bab4b9a39ccb717a14ad5ad; SHA256 1c8f573e830ca9b3ebfeb7ace1823146e22b66f99ee223840e7637c9e745e1c7; private manifest ai/v1/.model-metadata-development/36281114084/hf-deepseek-r1-license/2026-09-27/complete.json.gz; manifest SHA256 f9459f6fcb94c0ff78e42e7ef950ebf63df65e62fa523a9fd138ceca46c40e7b.

Checkpoint: DeepSeek V3

DeepSeek · DeepSeek V3

DeepSeek V3

No API price listedSelf-hosting costs not included
Open weightsDownloadabletextContext: Unknown tokens

Weight license: DeepSeek License Agreement v1.0 (use restrictions) · Hugging Face / download weights ↗

671B parameters · Estimated weight memory: Q4 503.3 GB · Q8 905.9 GB. Includes 20% loading overhead; excludes KV cache and runtime memory. Compare hardware.

Developer: DeepSeek

Checkpoint: DeepSeek V3 · Revision: e815299b0bcbac849fa540c768ef21845365c9eb · Lineage: Unknown / not verified

Creator model source ↗ Creator model source ↗ Weight-license evidence ↗

Model metadata observed: 2026-09-27T00:03:08.723Z / 2026-09-27T00:03:14.406Z · Metadata freshness: fresh

Weight-license reviewed: 2026-09-26T04:22:10.283Z. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: ungated · Code license: MIT · API availability: unknown (dated observation, not a live service check)

Active parameters: 37B; total weights: 671B.

Source, conditions and dates

Primary source ↗ · huggingface-model-card · Claims: deepseek-r1: identity · deepseek-r1: specifications · deepseek-r1: weightAccess · deepseek-r1: download · deepseek-r1: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

  • Full reasoning model, not a distilled checkpoint; 37B parameters are activated per token but all 671B weights must be stored.

Latest result: held. Retained evidence keeps its original observation and successful-check dates.

Primary source ↗ · huggingface-model-card · Claims: deepseek-v3: identity · deepseek-v3: specifications · deepseek-v3: weightAccess · deepseek-v3: download · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

  • Mixture-of-experts model with 37B parameters activated per token; local weight memory reflects 671B total parameters.

Latest result: held. Retained evidence keeps its original observation and successful-check dates.

Primary source ↗ · vendor-docs · Claims: deepseek-v3: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-26T04:22:10.283Z · Reviewed: 2026-09-26T04:22:10.283Z · Review after: Unknown · Freshness: unknown · Confidence: A

  • The creator card expressly applies this agreement to DeepSeek V3 Base/Chat model weights. Commercial use is supported subject to its conditions; MIT is the separate code license.
  • Redistribution and hosted use must retain license/notices, mark modified files, and pass through enforceable use-based restrictions to downstream users.
  • Attachment A restricts unlawful or rights-infringing use, military use, harm to minors, harmful misinformation, regulated inappropriate content, unauthorized personal information, harassment, adverse automated legal decisions, discriminatory uses and exploitation of vulnerable groups. Consult the exact agreement for complete terms.
  • This document review does not verify actual weight downloads or establish OSI approval.
  • Exact document SHA256: ccfee4895df06bcab524151c278e8dde88bbe76165a24ecbcbcf9fafd71fd2b3. Reviewed by Codex against collection 36216997956 and zero-source replay 36217141490.

Primary source ↗ · vendor-docs · Claims: deepseek-v3: codeLicense · Retrieved · Effective: Unknown · Checked: 2026-09-26T04:22:10.283Z · Reviewed: 2026-09-26T04:22:10.283Z · Review after: Unknown · Freshness: unknown · Confidence: A

  • The creator card separately identifies MIT for code. Preserve copyright and permission notices when distributing copies or substantial portions; see the license for warranty/liability terms.
  • This code license does not replace the separate weight-license agreement.
  • Exact document SHA256: 1c8f573e830ca9b3ebfeb7ace1823146e22b66f99ee223840e7637c9e745e1c7. Reviewed by Codex against collection 36216997956 and zero-source replay 36217141490.

Primary source ↗ · huggingface-model-card · Claims: deepseek-v3: identity · deepseek-v3: specifications · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:03:08.723Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Card specifies671B core-model parameters and37B active. The repository includes additional multi-token-prediction material and API tensor total684,531,386,000; neither active parameters nor repository totals replace core-model size.
  • Card advertises128K context but config specifies163,840 positions. Exact supported token count remains unknown rather than silently substituting configured positions or an unproven K convention.
  • Card mentions distillation from one of the R1-series models, which does not establish an exact relationship to the curated R1 checkpoint. V3-Base is outside the registry.
  • Unchanged V3 weight/code license document hashes match the already-reviewed September26 evidence; retain that earlier legal review and observation date.
  • Metadata documents were read; weight binaries were not downloaded. Prior weight-access/download observations retain their original dates.
  • Revision e815299b0bcbac849fa540c768ef21845365c9eb; SHA256 3739763a851074c810f1565f9823956f3a095231f8c591dd3c254581bdd3cf08; private manifest ai/v1/.model-metadata-development/36281114084/hf-deepseek-v3-readme-md/2026-09-27/complete.json.gz; manifest SHA256 d474f1d7ffd3b501c05f910f1bdc40935a2fcb0346974e2c64a8e102165474d0.

Raw API evidence ↗ · vendor-api · Claims: deepseek-v3: gating · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:03:03.033Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Repository gating: false; listed weight files: true. An inventory and authenticated metadata read do not verify weight download availability.
  • Revision e815299b0bcbac849fa540c768ef21845365c9eb; SHA256 313b7961ccdc22be024b2272412d8b66a0d8c71f58ad7ee9a86b68c450522ae5; private manifest ai/v1/.model-metadata-development/36281114084/hf-deepseek-v3-info/2026-09-27/complete.json.gz; manifest SHA256 94fb540218d8d71ac8b75a271dbae49430b3a27db6f3505b0b19fbd448375dcb.

Primary source ↗ · vendor-docs · Claims: deepseek-v3: configuration · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:03:14.406Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Configuration is supporting evidence, not an automatic claim of advertised context or modality. Resolve differences using the explicit publication review.
  • Revision e815299b0bcbac849fa540c768ef21845365c9eb; SHA256 cbf0b95dc614de208a109bb5fd4e7eed11385e9c68411d2c17db5319443035d9; private manifest ai/v1/.model-metadata-development/36281114084/hf-deepseek-v3-config-json/2026-09-27/complete.json.gz; manifest SHA256 6bc6313ce85f839ddcdde5cd8ae0b422ebae1d49a501634082061423dce47016.

DeepSeek

DeepSeek V4 Pro 0813

Input: USD 1.32 / 1000000 tokens · PeakOutput: USD 3.96 / 1000000 tokens · PeakHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

Cache read: USD 0.044 / 1000000 tokens · Peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-v4-pro, served by model version DeepSeek-V4-Pro-0813.
  • Peak rate: 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays.
  • Input tokens that hit the context cache.

Cache read: USD 0.022 / 1000000 tokens · Off-peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: off-peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-v4-pro, served by model version DeepSeek-V4-Pro-0813.
  • Off-peak rate (half the peak rate): all other hours, including weekends and Chinese public holidays in full.
  • Input tokens that hit the context cache.

Input: USD 1.32 / 1000000 tokens · Peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-v4-pro, served by model version DeepSeek-V4-Pro-0813.
  • Peak rate: 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays.
  • Input tokens on cache miss.

Input: USD 0.66 / 1000000 tokens · Off-peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: off-peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-v4-pro, served by model version DeepSeek-V4-Pro-0813.
  • Off-peak rate (half the peak rate): all other hours, including weekends and Chinese public holidays in full.
  • Input tokens on cache miss.

Output: USD 3.96 / 1000000 tokens · Peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-v4-pro, served by model version DeepSeek-V4-Pro-0813.
  • Peak rate: 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays.
  • Output tokens.

Output: USD 1.98 / 1000000 tokens · Off-peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: off-peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-v4-pro, served by model version DeepSeek-V4-Pro-0813.
  • Off-peak rate (half the peak rate): all other hours, including weekends and Chinese public holidays in full.
  • Output tokens.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: DeepSeek

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-27T05:50:28.262Z · Price source checked: 2026-09-27T05:50:28.262Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: deepseek-api-flash: pricing · deepseek-api-flash: availability · deepseek-api-v4-pro: pricing · deepseek-api-v4-pro: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:50:28.262Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 210f102275ccf1a6542f08a3bc9e4b4c7c83278cb74b35217bffa112df6363b2; collector deepseek-usd-1.
  • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. The API model versions are not linked to downloadable checkpoints.

DeepSeek

DeepSeek V4.1 Flash

Input: USD 0.3 / 1000000 tokens · PeakOutput: USD 1.2 / 1000000 tokens · PeakHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

Cache read: USD 0.006 / 1000000 tokens · Peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-flash, served by model version DeepSeek-V4.1-Flash.
  • Peak rate: 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays.
  • Input tokens that hit the context cache.

Cache read: USD 0.003 / 1000000 tokens · Off-peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: off-peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-flash, served by model version DeepSeek-V4.1-Flash.
  • Off-peak rate (half the peak rate): all other hours, including weekends and Chinese public holidays in full.
  • Input tokens that hit the context cache.

Input: USD 0.3 / 1000000 tokens · Peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-flash, served by model version DeepSeek-V4.1-Flash.
  • Peak rate: 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays.
  • Input tokens on cache miss.

Input: USD 0.15 / 1000000 tokens · Off-peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: off-peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-flash, served by model version DeepSeek-V4.1-Flash.
  • Off-peak rate (half the peak rate): all other hours, including weekends and Chinese public holidays in full.
  • Input tokens on cache miss.

Output: USD 1.2 / 1000000 tokens · Peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-flash, served by model version DeepSeek-V4.1-Flash.
  • Peak rate: 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays.
  • Output tokens.

Output: USD 0.6 / 1000000 tokens · Off-peak · Modality: unspecified

Region: global · Endpoint scope: global · Tier: off-peak · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • DeepSeek direct API, USD list price per 1M tokens from the English pricing page; CNY prices on the Chinese-language page are out of scope.
  • Model ID deepseek-flash, served by model version DeepSeek-V4.1-Flash.
  • Off-peak rate (half the peak rate): all other hours, including weekends and Chinese public holidays in full.
  • Output tokens.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: DeepSeek

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-27T05:50:28.262Z · Price source checked: 2026-09-27T05:50:28.262Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: deepseek-api-flash: pricing · deepseek-api-flash: availability · deepseek-api-v4-pro: pricing · deepseek-api-v4-pro: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:50:28.262Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 210f102275ccf1a6542f08a3bc9e4b4c7c83278cb74b35217bffa112df6363b2; collector deepseek-usd-1.
  • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. The API model versions are not linked to downloadable checkpoints.

Checkpoint: DeepSeek-V3.1-Nex-N1

Nex-AGI

DeepSeek-V3.1-Nex-N1

No API price listedSelf-hosting costs not included
Weights unknownDownload not verifiedtextContext: Unknown tokens

Weight license: Unknown / not verified

Developer: Nex-AGI

Checkpoint: DeepSeek-V3.1-Nex-N1 · Revision: Unknown / not verified · Lineage: Unknown / not verified

Creator model source ↗

Model metadata observed: 2026-09-27T05:30:48.668Z · Metadata freshness: fresh

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: nex-n1: identity · nex-n1: specifications · nex-n1: download · nex-n1: weightAccess · nex-n1: gating · developer:nex-agi: identity · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:30:48.668Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

  • Creator explicitly lists this representative as a post-trained release. Exact upstream checkpoint lineage is not established beyond its name; no DeepSeek-V3.1 edge or inherited parameter count is invented. The series8B–671B range and benchmark128K are not checkpoint specifications.
  • The exact GitHub LICENSE endpoint returned404. Legal scope remains unknown. The creator links nex.sii.edu.cn; no independent-shop or verified legal-entity classification is asserted. Local inference examples are not commercial API evidence.
  • Official checkpoint/download reference: https://huggingface.co/nex-agi/DeepSeek-V3.1-Nex-N1. No model weights downloaded; access/download availability remains unverified.
  • Document revision kind git-blob; revision 6f716f2b47f36c7fc37b19fc85935cd8c4840cee; SHA256 0d75721a151569a85587ff0efc778739b51f71076db9b38b40e5cd18ceb77cd5; private manifest ai/v1/onboard-nex-n1-readme-md/2026-09-27/complete.json.gz; manifest SHA256 561baec2f1040131e7ba4788b512cdc7e7d164194b32c66405139ac5bb9aa663.

Checkpoint: dots.llm1.inst

rednote-hilab / Dots

dots.llm1.inst

No API price listedSelf-hosting costs not included
Weights unknownDownload not verifiedtextContext: 32,768 tokens

Weight license: Unknown / not verified

Developer: rednote-hilab / Dots

Checkpoint: dots.llm1.inst · Revision: Unknown / not verified · Lineage: Unknown / not verified

Creator model source ↗

Model metadata observed: 2026-09-27T05:30:38.167Z · Metadata freshness: fresh

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: MIT · API availability: unknown (dated observation, not a live service check)

Active parameters: 14B; total weights: 142B.

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: dots-llm1: identity · dots-llm1: specifications · dots-llm1: download · dots-llm1: weightAccess · dots-llm1: gating · developer:rednote-hilab: identity · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:30:38.167Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

  • Choose the exact instruction release from the creator download table; do not collapse it with dots.llm1.base. The explicit context value is32768.
  • The GitHub studio-dots-ai project links rednote-hilab model repositories and its LICENSE credits rednote-hilab. The model summary declares MIT, but separate weight scope is conservatively left unknown. No independent-company classification is made.
  • Official checkpoint/download reference: https://huggingface.co/rednote-hilab/dots.llm1.inst. No model weights downloaded; access/download availability remains unverified.
  • Document revision kind git-blob; revision c017d67bfe61cad3478bec9b662fbbc58c00d4e9; SHA256 642ecd79e806353f38a886ba2454e43d8a03fb1c32374bdc34a57b7641f6dbf8; private manifest ai/v1/onboard-dots-llm1-readme-md/2026-09-27/complete.json.gz; manifest SHA256 065be32027b29a78b28c35421d6d7bba1954a14b9aafbb5f508ce7bffc9c33ba.

Primary source ↗ · vendor-docs · Claims: dots-llm1: codeLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:30:43.339Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

  • Choose the exact instruction release from the creator download table; do not collapse it with dots.llm1.base. The explicit context value is32768.
  • The GitHub studio-dots-ai project links rednote-hilab model repositories and its LICENSE credits rednote-hilab. The model summary declares MIT, but separate weight scope is conservatively left unknown. No independent-company classification is made.
  • Official checkpoint/download reference: https://huggingface.co/rednote-hilab/dots.llm1.inst. No model weights downloaded; access/download availability remains unverified.
  • Document revision kind git-blob; revision 4cdd25386a33c12f2af65fb06c49c2018f01000b; SHA256 8060e25630dd2af626d870352630274fec7568f2b324260726460209c2e13afd; private manifest ai/v1/onboard-dots-llm1-license/2026-09-27/complete.json.gz; manifest SHA256 c8fa7a6cfff1e669e55e4b2c571c16fdfb286d0b36196893974ba715113a95ca.

Google

Gemini 2.5 Flash-Lite (text/image/video)

Input: USD 0.10 / 1000000 tokens · StandardOutput: USD 0.40 / 1000000 tokens · StandardHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedmultimodalContext: Unknown tokens

Input: USD 0.10 / 1000000 tokens · Standard · Modality: text, image, video

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-2.5-flash-lite. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.10 (text / image / video) $0.30 (audio)
  • Output (including thinking tokens) paid cell: $0.40
  • Context caching paid cell: $0.01 (text / image / video) $0.03 (audio) $1.00 / 1,000,000 tokens per hour (storage price)

Cache read: USD 0.01 / 1000000 tokens · Standard · Modality: text, image, video

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-2.5-flash-lite. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.10 (text / image / video) $0.30 (audio)
  • Output (including thinking tokens) paid cell: $0.40
  • Context caching paid cell: $0.01 (text / image / video) $0.03 (audio) $1.00 / 1,000,000 tokens per hour (storage price)

Storage: USD 1.00 / 1000000 token-hours · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-2.5-flash-lite. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.10 (text / image / video) $0.30 (audio)
  • Output (including thinking tokens) paid cell: $0.40
  • Context caching paid cell: $0.01 (text / image / video) $0.03 (audio) $1.00 / 1,000,000 tokens per hour (storage price)

Output: USD 0.40 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-2.5-flash-lite. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.10 (text / image / video) $0.30 (audio)
  • Output (including thinking tokens) paid cell: $0.40
  • Context caching paid cell: $0.01 (text / image / video) $0.03 (audio) $1.00 / 1,000,000 tokens per hour (storage price)

Cache write: Unknown / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-2.5-flash-lite. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.10 (text / image / video) $0.30 (audio)
  • Output (including thinking tokens) paid cell: $0.40
  • Context caching paid cell: $0.01 (text / image / video) $0.03 (audio) $1.00 / 1,000,000 tokens per hour (storage price)

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Google

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-26T00:43:16.524Z · Price source checked: 2026-09-26T00:43:16.524Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: google-gemini-2-5-flash-lite-text: identity · google-gemini-2-5-flash-lite-text: pricing · google-gemini-2-5-flash-lite-text: availability · google-gemini-2-5-flash-lite-text: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Text, image, and video input rate; audio input costs more.

Primary source ↗ · vendor-docs · Claims: google-gemini-2-5-flash-lite-text: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

  • The official model card identifies Gemini 2.5 Flash-Lite and describes its capabilities and safety evaluations; this review did not establish explicit weight-release or non-downloadability evidence.
  • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
  • This review applies to the named model, not a separate checkpoint inferred from the pricing record's prompt-length or input-modality tier. It does not refresh pricing or resolve unverified snapshot identity.

Primary source ↗ · vendor-docs · Claims: google-gemini-2-5-flash-lite-text: pricing · google-gemini-2-5-pro-short: pricing · google-gemini-3-1-pro-preview-short: pricing · google-gemini-3-5-flash-lite: pricing · google-gemini-3-7-flash: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T00:43:16.524Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 01b274b8e94989b587b8b084e57f12f8169976b46411714c9b62df6fb0a2ac76; collector google-paid-standard-2.
  • Pricing observation only. Prior independent identity/access/availability/context metadata unchanged. Curated five offers, not the entire provider catalog.

Google

Gemini 2.5 Pro (≤200K prompt)

Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedreasoningContext: Unknown tokens

Input: USD 1.25 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–200000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-2.5-pro. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $1.25, prompts <= 200k tokens $2.50, prompts > 200k tokens
  • Output (including thinking tokens) paid cell: $10.00, prompts <= 200k tokens $15.00, prompts > 200k
  • Context caching paid cell: $0.125, prompts <= 200k tokens $0.25, prompts > 200k $4.50 / 1,000,000 tokens per hour (storage price)

Output: USD 10.00 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–200000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-2.5-pro. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $1.25, prompts <= 200k tokens $2.50, prompts > 200k tokens
  • Output (including thinking tokens) paid cell: $10.00, prompts <= 200k tokens $15.00, prompts > 200k
  • Context caching paid cell: $0.125, prompts <= 200k tokens $0.25, prompts > 200k $4.50 / 1,000,000 tokens per hour (storage price)

Cache read: USD 0.125 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–200000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-2.5-pro. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $1.25, prompts <= 200k tokens $2.50, prompts > 200k tokens
  • Output (including thinking tokens) paid cell: $10.00, prompts <= 200k tokens $15.00, prompts > 200k
  • Context caching paid cell: $0.125, prompts <= 200k tokens $0.25, prompts > 200k $4.50 / 1,000,000 tokens per hour (storage price)

Storage: USD 4.50 / 1000000 token-hours · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-2.5-pro. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $1.25, prompts <= 200k tokens $2.50, prompts > 200k tokens
  • Output (including thinking tokens) paid cell: $10.00, prompts <= 200k tokens $15.00, prompts > 200k
  • Context caching paid cell: $0.125, prompts <= 200k tokens $0.25, prompts > 200k $4.50 / 1,000,000 tokens per hour (storage price)

Cache write: Unknown / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-2.5-pro. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $1.25, prompts <= 200k tokens $2.50, prompts > 200k tokens
  • Output (including thinking tokens) paid cell: $10.00, prompts <= 200k tokens $15.00, prompts > 200k
  • Context caching paid cell: $0.125, prompts <= 200k tokens $0.25, prompts > 200k $4.50 / 1,000,000 tokens per hour (storage price)

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Google

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-26T00:43:16.524Z · Price source checked: 2026-09-26T00:43:16.524Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: google-gemini-2-5-pro-short: identity · google-gemini-2-5-pro-short: pricing · google-gemini-2-5-pro-short: availability · google-gemini-2-5-pro-short: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Higher rates apply when the prompt exceeds 200K tokens; output includes thinking tokens.

Primary source ↗ · vendor-docs · Claims: google-gemini-2-5-pro-short: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

  • The official model card identifies Gemini 2.5 Pro and describes its capabilities and safety evaluations; this review did not establish explicit weight-release or non-downloadability evidence.
  • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
  • This review applies to the named model, not a separate checkpoint inferred from the pricing record's prompt-length or input-modality tier. It does not refresh pricing or resolve unverified snapshot identity.

Primary source ↗ · vendor-docs · Claims: google-gemini-2-5-flash-lite-text: pricing · google-gemini-2-5-pro-short: pricing · google-gemini-3-1-pro-preview-short: pricing · google-gemini-3-5-flash-lite: pricing · google-gemini-3-7-flash: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T00:43:16.524Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 01b274b8e94989b587b8b084e57f12f8169976b46411714c9b62df6fb0a2ac76; collector google-paid-standard-2.
  • Pricing observation only. Prior independent identity/access/availability/context metadata unchanged. Curated five offers, not the entire provider catalog.

Google

Gemini 3.1 Pro Preview (≤200K prompt)

Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedreasoningContext: Unknown tokens

Input: USD 2.00 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–200000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.1-pro-preview and gemini-3.1-pro-preview-customtools. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $2.00, prompts <= 200k tokens $4.00, prompts > 200k tokens
  • Output (including thinking tokens) paid cell: $12.00, prompts <= 200k tokens $18.00, prompts > 200k
  • Context caching paid cell: $0.20, prompts <= 200k tokens $0.40, prompts > 200k $4.50 / 1,000,000 tokens per hour (storage price)

Output: USD 12.00 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–200000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.1-pro-preview and gemini-3.1-pro-preview-customtools. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $2.00, prompts <= 200k tokens $4.00, prompts > 200k tokens
  • Output (including thinking tokens) paid cell: $12.00, prompts <= 200k tokens $18.00, prompts > 200k
  • Context caching paid cell: $0.20, prompts <= 200k tokens $0.40, prompts > 200k $4.50 / 1,000,000 tokens per hour (storage price)

Cache read: USD 0.20 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–200000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.1-pro-preview and gemini-3.1-pro-preview-customtools. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $2.00, prompts <= 200k tokens $4.00, prompts > 200k tokens
  • Output (including thinking tokens) paid cell: $12.00, prompts <= 200k tokens $18.00, prompts > 200k
  • Context caching paid cell: $0.20, prompts <= 200k tokens $0.40, prompts > 200k $4.50 / 1,000,000 tokens per hour (storage price)

Storage: USD 4.50 / 1000000 token-hours · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.1-pro-preview and gemini-3.1-pro-preview-customtools. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $2.00, prompts <= 200k tokens $4.00, prompts > 200k tokens
  • Output (including thinking tokens) paid cell: $12.00, prompts <= 200k tokens $18.00, prompts > 200k
  • Context caching paid cell: $0.20, prompts <= 200k tokens $0.40, prompts > 200k $4.50 / 1,000,000 tokens per hour (storage price)

Cache write: Unknown / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.1-pro-preview and gemini-3.1-pro-preview-customtools. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $2.00, prompts <= 200k tokens $4.00, prompts > 200k tokens
  • Output (including thinking tokens) paid cell: $12.00, prompts <= 200k tokens $18.00, prompts > 200k
  • Context caching paid cell: $0.20, prompts <= 200k tokens $0.40, prompts > 200k $4.50 / 1,000,000 tokens per hour (storage price)

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Google

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-26T00:43:16.524Z · Price source checked: 2026-09-26T00:43:16.524Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: google-gemini-3-1-pro-preview-short: identity · google-gemini-3-1-pro-preview-short: pricing · google-gemini-3-1-pro-preview-short: availability · google-gemini-3-1-pro-preview-short: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Preview model; higher rates apply when the prompt exceeds 200K tokens.

Primary source ↗ · vendor-docs · Claims: google-gemini-3-1-pro-preview-short: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

  • The official model card for Gemini 3.1 Pro describes distribution through hosted applications and APIs; those channels alone do not establish whether weights are downloadable.
  • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.
  • This review applies to the named model, not a separate checkpoint inferred from the pricing record's prompt-length or input-modality tier. It does not refresh pricing or resolve unverified snapshot identity.

Primary source ↗ · vendor-docs · Claims: google-gemini-2-5-flash-lite-text: pricing · google-gemini-2-5-pro-short: pricing · google-gemini-3-1-pro-preview-short: pricing · google-gemini-3-5-flash-lite: pricing · google-gemini-3-7-flash: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T00:43:16.524Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 01b274b8e94989b587b8b084e57f12f8169976b46411714c9b62df6fb0a2ac76; collector google-paid-standard-2.
  • Pricing observation only. Prior independent identity/access/availability/context metadata unchanged. Curated five offers, not the entire provider catalog.

Google

Gemini 3.5 Flash-Lite

Input: USD 0.30 / 1000000 tokens · StandardOutput: USD 2.50 / 1000000 tokens · StandardHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedmultimodalContext: Unknown tokens

Input: USD 0.30 / 1000000 tokens · Standard · Modality: text, image, video, audio

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.5-flash-lite. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.30 (text / image / video / audio)
  • Output (including thinking tokens) paid cell: $2.50
  • Context caching paid cell: $0.03 $1.00 / 1,000,000 tokens per hour (storage price)

Output: USD 2.50 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.5-flash-lite. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.30 (text / image / video / audio)
  • Output (including thinking tokens) paid cell: $2.50
  • Context caching paid cell: $0.03 $1.00 / 1,000,000 tokens per hour (storage price)

Cache read: USD 0.03 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.5-flash-lite. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.30 (text / image / video / audio)
  • Output (including thinking tokens) paid cell: $2.50
  • Context caching paid cell: $0.03 $1.00 / 1,000,000 tokens per hour (storage price)

Storage: USD 1.00 / 1000000 token-hours · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.5-flash-lite. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.30 (text / image / video / audio)
  • Output (including thinking tokens) paid cell: $2.50
  • Context caching paid cell: $0.03 $1.00 / 1,000,000 tokens per hour (storage price)

Cache write: Unknown / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.5-flash-lite. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.30 (text / image / video / audio)
  • Output (including thinking tokens) paid cell: $2.50
  • Context caching paid cell: $0.03 $1.00 / 1,000,000 tokens per hour (storage price)

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Google

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-26T00:43:16.524Z · Price source checked: 2026-09-26T00:43:16.524Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: google-gemini-3-5-flash-lite: identity · google-gemini-3-5-flash-lite: pricing · google-gemini-3-5-flash-lite: availability · google-gemini-3-5-flash-lite: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Standard paid tier; output includes thinking tokens.

Primary source ↗ · vendor-docs · Claims: google-gemini-3-5-flash-lite: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

  • The official model card for Gemini 3.5 Flash-Lite describes distribution through hosted applications and APIs; those channels alone do not establish whether weights are downloadable.
  • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.

Primary source ↗ · vendor-docs · Claims: google-gemini-2-5-flash-lite-text: pricing · google-gemini-2-5-pro-short: pricing · google-gemini-3-1-pro-preview-short: pricing · google-gemini-3-5-flash-lite: pricing · google-gemini-3-7-flash: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T00:43:16.524Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 01b274b8e94989b587b8b084e57f12f8169976b46411714c9b62df6fb0a2ac76; collector google-paid-standard-2.
  • Pricing observation only. Prior independent identity/access/availability/context metadata unchanged. Curated five offers, not the entire provider catalog.

Google

Gemini 3.7 Flash

Input: USD 0.75 / 1000000 tokens · StandardOutput: USD 3.75 / 1000000 tokens · StandardHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedmultimodalContext: Unknown tokens

Limited-time price: until 2026-12-31. Later prices, where stated, are in each rate's conditions.

Input: USD 0.75 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: 2026-12-31

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.7-flash. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.75 through December 31, 2026. $1.50 starting January 1, 2027.
  • Output (including thinking tokens) paid cell: $3.75 through December 31, 2026. $7.50 starting January 1, 2027.
  • Context caching paid cell: $0.075 through December 31, 2026. $0.15 starting January 1, 2027. $0.50 / 1,000,000 tokens per hour (storage price) through December 31, 2026. $1.00 / 1,000,000 tokens per hour (storage price) starting January 1, 2027.

Output: USD 3.75 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: 2026-12-31

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.7-flash. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.75 through December 31, 2026. $1.50 starting January 1, 2027.
  • Output (including thinking tokens) paid cell: $3.75 through December 31, 2026. $7.50 starting January 1, 2027.
  • Context caching paid cell: $0.075 through December 31, 2026. $0.15 starting January 1, 2027. $0.50 / 1,000,000 tokens per hour (storage price) through December 31, 2026. $1.00 / 1,000,000 tokens per hour (storage price) starting January 1, 2027.

Cache read: USD 0.075 / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: 2026-12-31

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.7-flash. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.75 through December 31, 2026. $1.50 starting January 1, 2027.
  • Output (including thinking tokens) paid cell: $3.75 through December 31, 2026. $7.50 starting January 1, 2027.
  • Context caching paid cell: $0.075 through December 31, 2026. $0.15 starting January 1, 2027. $0.50 / 1,000,000 tokens per hour (storage price) through December 31, 2026. $1.00 / 1,000,000 tokens per hour (storage price) starting January 1, 2027.

Storage: USD 0.50 / 1000000 token-hours · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: 2026-12-31

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.7-flash. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.75 through December 31, 2026. $1.50 starting January 1, 2027.
  • Output (including thinking tokens) paid cell: $3.75 through December 31, 2026. $7.50 starting January 1, 2027.
  • Context caching paid cell: $0.075 through December 31, 2026. $0.15 starting January 1, 2027. $0.50 / 1,000,000 tokens per hour (storage price) through December 31, 2026. $1.00 / 1,000,000 tokens per hour (storage price) starting January 1, 2027.

Cache write: Unknown / 1000000 tokens · Standard · Modality: unspecified

Region: unknown · Endpoint scope: unknown · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Gemini API Paid Standard only; excludes Free, Batch, Flex, Priority and Vertex/cloud. Account and regional eligibility unknown.
  • Official service code line: gemini-3.7-flash. No additional customtools offer inferred.
  • Prompt brackets are billing thresholds, not model context-window capacity. Output includes thinking tokens.
  • Cache storage is billed per token-hour, not cache-write. Cache-write price and TTL unknown.
  • Document token billing: Tokens for the DOCUMENT modality (for example, PDFs) are billed at the image token rate. In API responses, these tokens appear under the DOCUMENT modality within promptTokensDetails.
  • Grounding/search/maps tools and quotas are not included in these token prices; consult the official pricing page. No separate PDF rate inferred.
  • Advertised promotion dates have no stated timezone; future rates require renewed review, not automatic activation.
  • Input paid cell: $0.75 through December 31, 2026. $1.50 starting January 1, 2027.
  • Output (including thinking tokens) paid cell: $3.75 through December 31, 2026. $7.50 starting January 1, 2027.
  • Context caching paid cell: $0.075 through December 31, 2026. $0.15 starting January 1, 2027. $0.50 / 1,000,000 tokens per hour (storage price) through December 31, 2026. $1.00 / 1,000,000 tokens per hour (storage price) starting January 1, 2027.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Google

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-26T00:43:16.524Z · Price source checked: 2026-09-26T00:43:16.524Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: google-gemini-3-7-flash: identity · google-gemini-3-7-flash: pricing · google-gemini-3-7-flash: availability · google-gemini-3-7-flash: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Promotional through December 31, 2026; output includes thinking tokens.

Primary source ↗ · vendor-docs · Claims: google-gemini-3-7-flash: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

  • The official model card for Gemini 3.7 Flash describes distribution through hosted applications and APIs; those channels alone do not establish whether weights are downloadable.
  • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.

Primary source ↗ · vendor-docs · Claims: google-gemini-2-5-flash-lite-text: pricing · google-gemini-2-5-pro-short: pricing · google-gemini-3-1-pro-preview-short: pricing · google-gemini-3-5-flash-lite: pricing · google-gemini-3-7-flash: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T00:43:16.524Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 01b274b8e94989b587b8b084e57f12f8169976b46411714c9b62df6fb0a2ac76; collector google-paid-standard-2.
  • Pricing observation only. Prior independent identity/access/availability/context metadata unchanged. Curated five offers, not the entire provider catalog.

Checkpoint: Gemma 3 1B IT

Google · Gemma 3

Gemma 3 1B IT

No API price listedSelf-hosting costs not included
Open weightsDownloadabletextContext: 32,768 tokens

Weight license: Gemma license · Hugging Face / download weights ↗

1B parameters · Estimated weight memory: Q4 0.8 GB · Q8 1.3 GB. Includes 20% loading overhead; excludes KV cache and runtime memory. Compare hardware.

Developer: Google

Checkpoint: Gemma 3 1B IT · Revision: dcc83ea841ab6100d6b47a070329e1ba4cf78752 · Lineage: Unknown / not verified

Creator model source ↗ Creator model source ↗ Weight-license evidence ↗

Model metadata observed: 2026-09-27T00:01:00.458Z / 2026-09-27T00:01:06.226Z · Metadata freshness: fresh

Weight-license reviewed: 2026-09-27T02:50:26.221Z. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: gated · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · huggingface-model-card · Claims: google-gemma-3-1b-it: identity · google-gemma-3-1b-it: specifications · google-gemma-3-1b-it: weightAccess · google-gemma-3-1b-it: download · google-gemma-3-1b-it: weightLicense · google-gemma-3-1b-it: gating · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

  • Text-only model with gated license acknowledgment.

Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

Primary source ↗ · huggingface-model-card · Claims: google-gemma-3-4b-it: identity · google-gemma-3-4b-it: specifications · google-gemma-3-4b-it: weightAccess · google-gemma-3-4b-it: download · google-gemma-3-4b-it: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

  • Multimodal input; images are normalized to 896×896 and encoded as 256 tokens.

Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

Primary source ↗ · huggingface-model-card · Claims: google-gemma-3-1b-it: identity · google-gemma-3-1b-it: specifications · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:00.458Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Card identifies the 1B variant and 32K input context; exact configuration specifies 32,768 positions.
  • The family card discusses images, but this exact checkpoint uses text-generation and Gemma3ForCausalLM with no vision configuration. Keep text-only capability. Base google/gemma-3-1b-pt is outside the registry.
  • Metadata documents were read; weight binaries were not downloaded. Prior weight-access/download observations retain their original dates.
  • Revision dcc83ea841ab6100d6b47a070329e1ba4cf78752; SHA256 60be259533fe3acba21d109a51673815a4c29aefdb7769862695086fcedbeb7a; private manifest ai/v1/.model-metadata-development/36281114084/hf-google-gemma-3-1b-it-readme-md/2026-09-27/complete.json.gz; manifest SHA256 7446ea235db2d011fc6dabae8ef7f926e3fd2c26ea53ea31a7cb4b417105ddcb.

Raw API evidence ↗ · vendor-api · Claims: google-gemma-3-1b-it: gating · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:00:54.711Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Repository gating: manual; listed weight files: true. An inventory and authenticated metadata read do not verify weight download availability.
  • Revision dcc83ea841ab6100d6b47a070329e1ba4cf78752; SHA256 6e06a706c5c418a7991f38103003c450443dbc8d72abef33ebfdf864d40b9096; private manifest ai/v1/.model-metadata-development/36281114084/hf-google-gemma-3-1b-it-info/2026-09-27/complete.json.gz; manifest SHA256 c24a80bfad159081137147f12b34a4431313de394e70a3d58743e055330788d2.

Primary source ↗ · vendor-docs · Claims: google-gemma-3-1b-it: configuration · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:06.226Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Configuration is supporting evidence, not an automatic claim of advertised context or modality. Resolve differences using the explicit publication review.
  • Revision dcc83ea841ab6100d6b47a070329e1ba4cf78752; SHA256 19cb5d28c97778271ba2b3c3df47bf76bdd6706724777a2318b3522230afe91e; private manifest ai/v1/.model-metadata-development/36281114084/hf-google-gemma-3-1b-it-config-json/2026-09-27/complete.json.gz; manifest SHA256 18f27b109c0b28e29b354959feac096abedc465350ccad4730722abfa53475e3.

Primary source ↗ · huggingface-model-card · Claims: google-gemma-3-1b-it: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:00.458Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Exact creator card declares the Gemma license and gated license acknowledgement. No separate license document is in this captured inventory; detailed restrictions and independent code-license scope are not verified here. Do not classify this as permissive Apache/MIT or unconditional commercial permission.
  • Revision dcc83ea841ab6100d6b47a070329e1ba4cf78752; SHA256 60be259533fe3acba21d109a51673815a4c29aefdb7769862695086fcedbeb7a; private manifest ai/v1/.model-metadata-development/36281114084/hf-google-gemma-3-1b-it-readme-md/2026-09-27/complete.json.gz; manifest SHA256 7446ea235db2d011fc6dabae8ef7f926e3fd2c26ea53ea31a7cb4b417105ddcb.

Checkpoint: Gemma 3 4B IT

Google · Gemma 3

Gemma 3 4B IT

No API price listedSelf-hosting costs not included
Open weightsDownloadabletext · imageContext: Unknown tokens

Weight license: Gemma license · Hugging Face / download weights ↗

4B parameters · Estimated weight memory: Q4 3.0 GB · Q8 5.4 GB. Includes 20% loading overhead; excludes KV cache and runtime memory. Compare hardware.

Developer: Google

Checkpoint: Gemma 3 4B IT · Revision: 093f9f388b31de276ce2de164bdc2081324b9767 · Lineage: Unknown / not verified

Creator model source ↗ Creator model source ↗ Weight-license evidence ↗

Model metadata observed: 2026-09-27T00:01:19.238Z / 2026-09-27T00:01:25.290Z · Metadata freshness: fresh

Weight-license reviewed: 2026-09-27T02:50:26.221Z. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: gated · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · huggingface-model-card · Claims: google-gemma-3-1b-it: identity · google-gemma-3-1b-it: specifications · google-gemma-3-1b-it: weightAccess · google-gemma-3-1b-it: download · google-gemma-3-1b-it: weightLicense · google-gemma-3-1b-it: gating · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

  • Text-only model with gated license acknowledgment.

Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

Primary source ↗ · huggingface-model-card · Claims: google-gemma-3-4b-it: identity · google-gemma-3-4b-it: specifications · google-gemma-3-4b-it: weightAccess · google-gemma-3-4b-it: download · google-gemma-3-4b-it: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

  • Multimodal input; images are normalized to 896×896 and encoded as 256 tokens.

Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

Primary source ↗ · huggingface-model-card · Claims: google-gemma-3-4b-it: identity · google-gemma-3-4b-it: specifications · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:19.238Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Creator card identifies the 4B model; API tensor total 4,300,079,472 is recorded separately and does not redefine its advertised model size.
  • Card advertises 128K input context and 8,192 output tokens; archived config omits an explicit maximum-position value. Exact supported total token count remains unknown; do not add the limits or invent K rounding. Text/image input is established by card and vision configuration. Base google/gemma-3-4b-pt is outside the registry.
  • Metadata documents were read; weight binaries were not downloaded. Prior weight-access/download observations retain their original dates.
  • Revision 093f9f388b31de276ce2de164bdc2081324b9767; SHA256 eb6df732fd79f239d46a2bbc75302795cf293acd5b75fcc01995344da02aa0bf; private manifest ai/v1/.model-metadata-development/36281114084/hf-google-gemma-3-4b-it-readme-md/2026-09-27/complete.json.gz; manifest SHA256 d7dafe3b8102e6ea6549f693c47ae158e3be28c311fe414092e52a6d95983643.

Raw API evidence ↗ · vendor-api · Claims: google-gemma-3-4b-it: gating · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:13.195Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Repository gating: manual; listed weight files: true. An inventory and authenticated metadata read do not verify weight download availability.
  • Revision 093f9f388b31de276ce2de164bdc2081324b9767; SHA256 c984aefb655693f3cede71c39b418bf395fbcfcf69ebb4088365b00835b12ec6; private manifest ai/v1/.model-metadata-development/36281114084/hf-google-gemma-3-4b-it-info/2026-09-27/complete.json.gz; manifest SHA256 10b4c54462155ed273e3558c832291c8f9e9df36a0f9027e408e4390ca4ec435.

Primary source ↗ · vendor-docs · Claims: google-gemma-3-4b-it: configuration · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:25.290Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Configuration is supporting evidence, not an automatic claim of advertised context or modality. Resolve differences using the explicit publication review.
  • Revision 093f9f388b31de276ce2de164bdc2081324b9767; SHA256 9059f680f4dbd1957f35cb44b9fdd6948f4792db7a7a35ee353ea42e68adf7ff; private manifest ai/v1/.model-metadata-development/36281114084/hf-google-gemma-3-4b-it-config-json/2026-09-27/complete.json.gz; manifest SHA256 bc44d0098440f4d4c9968c98f25b6456ec09bef37e1a5f864c2d87f30095bf3a.

Primary source ↗ · huggingface-model-card · Claims: google-gemma-3-4b-it: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:19.238Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

  • Exact creator card declares the Gemma license and gated license acknowledgement. No separate license document is in this captured inventory; detailed restrictions and independent code-license scope are not verified here. Do not classify this as permissive Apache/MIT or unconditional commercial permission.
  • Revision 093f9f388b31de276ce2de164bdc2081324b9767; SHA256 eb6df732fd79f239d46a2bbc75302795cf293acd5b75fcc01995344da02aa0bf; private manifest ai/v1/.model-metadata-development/36281114084/hf-google-gemma-3-4b-it-readme-md/2026-09-27/complete.json.gz; manifest SHA256 d7dafe3b8102e6ea6549f693c47ae158e3be28c311fe414092e52a6d95983643.

Novita AI

GLM 5

Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

Input: USD 1 / 1000000 tokens · serverless · Modality: text

Region: Unknown · Endpoint scope: unknown · Tier: serverless · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Novita hosted serverless listing, USD per 1M tokens; standard listed prices only. Batch, priority, dedicated, promotional and free-limited offers are excluded.
  • Native decimal values agree with integer ten-thousandths and original/listed amounts. No tiered billing is reported for this listing.
  • Account and regional eligibility, status-code semantics and effective dates are unverified; listing is not an inference service test.
  • Listed prompt/input token rate; separately listed cache-read rates apply only to cache hits.

Output: USD 3.2 / 1000000 tokens · serverless · Modality: text

Region: Unknown · Endpoint scope: unknown · Tier: serverless · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Novita hosted serverless listing, USD per 1M tokens; standard listed prices only. Batch, priority, dedicated, promotional and free-limited offers are excluded.
  • Native decimal values agree with integer ten-thousandths and original/listed amounts. No tiered billing is reported for this listing.
  • Account and regional eligibility, status-code semantics and effective dates are unverified; listing is not an inference service test.
  • Listed completion/output token rate.

Cache read: USD 0.2 / 1000000 tokens · serverless · Modality: text

Region: Unknown · Endpoint scope: unknown · Tier: serverless · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Novita hosted serverless listing, USD per 1M tokens; standard listed prices only. Batch, priority, dedicated, promotional and free-limited offers are excluded.
  • Native decimal values agree with integer ten-thousandths and original/listed amounts. No tiered billing is reported for this listing.
  • Account and regional eligibility, status-code semantics and effective dates are unverified; listing is not an inference service test.
  • Cache-read input only; cache-write pricing and cache TTL are unknown.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Novita AI

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-27T06:32:08.821Z · Price source checked: 2026-09-27T06:32:08.821Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

Source, conditions and dates

Raw API evidence ↗ · vendor-api · Claims: novita-llama-3-3-70b-instruct: pricing · novita-minimax-m2: pricing · novita-glm-5: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T06:32:08.821Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: B

  • Raw SHA-256: 9722d1130261fe9667b226c4817976e42d0087d11cf1eb7b948cecdeb7988f25; original collector novita-model-catalog-1. Pinned parse-failure recovery preserves original dates; no new acquisition.
  • USD/scale interpretation reviewed against Novita official pricing and its MiniMax M3 pricing explanation; exact original/decimal/integer values retained privately.
  • Three curated hosted listings only. No verified checkpoint/developer/license/weight-access link; availability remains unknown because numeric status semantics are undocumented.

Checkpoint: GLM-5

Z.ai

GLM-5

No API price listedSelf-hosting costs not included
Weights unknownDownload not verifiedtextContext: Unknown tokens

Weight license: Unknown / not verified

Developer: Z.ai

Checkpoint: GLM-5 · Revision: Unknown / not verified · Lineage: Unknown / not verified

Creator model source ↗

Model metadata observed: 2026-09-27T05:30:18.102Z · Metadata freshness: fresh

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Apache-2.0 · API availability: unknown (dated observation, not a live service check)

Active parameters: 40B; total weights: 744B.

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: glm-5: identity · glm-5: specifications · glm-5: download · glm-5: weightAccess · glm-5: gating · developer:z-ai: identity · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:30:18.102Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

  • Use only the exact GLM-5 section and download row. The repository also describes newer5.1/5.2/5.3 releases whose context, modalities and licenses are not inherited.
  • The captured GitHub repository LICENSE is Apache-2.0 for repository code; the older HF card has an MIT metadata declaration without scoped weight terms. Weight scope remains unknown. Z.ai/zai-org branding and ZhipuAI ModelScope links are observed; a separate legal-entity equivalence is not established.
  • Official checkpoint/download reference: https://huggingface.co/zai-org/GLM-5. No model weights downloaded; access/download availability remains unverified.
  • Document revision kind git-blob; revision b50da8f1ed05d96f27184200496acd59f3b345fb; SHA256 78ebb1936e94a4c4412bca2abc1c9cbe599455df12fb1373e75c495cb043651e; private manifest ai/v1/onboard-glm-5-readme-md/2026-09-27/complete.json.gz; manifest SHA256 ee944393b2bd40bd37600fe2a7d41f47a4897c7c3bff19532bdf8966110830d1.

Primary source ↗ · vendor-docs · Claims: glm-5: codeLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:30:23.160Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

  • Use only the exact GLM-5 section and download row. The repository also describes newer5.1/5.2/5.3 releases whose context, modalities and licenses are not inherited.
  • The captured GitHub repository LICENSE is Apache-2.0 for repository code; the older HF card has an MIT metadata declaration without scoped weight terms. Weight scope remains unknown. Z.ai/zai-org branding and ZhipuAI ModelScope links are observed; a separate legal-entity equivalence is not established.
  • Official checkpoint/download reference: https://huggingface.co/zai-org/GLM-5. No model weights downloaded; access/download availability remains unverified.
  • Document revision kind git-blob; revision ca04a51e70398d8fe19b23a8c317374dfe25408b; SHA256 086c6e7f5082b4640dfb03a892f32504f24afd247c702c38a3aae8c3474c54f0; private manifest ai/v1/onboard-glm-5-license/2026-09-27/complete.json.gz; manifest SHA256 4da697f3c332f86f8f309afbb4197ab4059a108b8d2b50d37879f2fd86f3ad03.

Z.ai

GLM-5.2

Input: USD 1.4 / 1000000 tokens · Tier unspecifiedOutput: USD 4.4 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

Input: USD 1.4 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.2.
  • Input tokens.

Cache read: USD 0.26 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.2.
  • Cached input tokens. The source lists cached-input storage as "Limited-time Free" (no end date); storage is not priced here.

Output: USD 4.4 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.2.
  • Output tokens.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Z.ai

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-27T13:56:23.785Z · Price source checked: 2026-09-27T13:56:23.785Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: zai-api-glm-5-3-flash: pricing · zai-api-glm-5-3-flash: availability · zai-api-glm-5-3-flashx: pricing · zai-api-glm-5-3-flashx: availability · zai-api-glm-5-3: pricing · zai-api-glm-5-3: availability · zai-api-glm-5-2: pricing · zai-api-glm-5-2: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T13:56:23.785Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 3b5587f7826bf8573174e721613745262cf9a891f8a1408385251ccbba0b5be2; collector zai-international-1.
  • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. API models are not linked to downloadable GLM checkpoints (GLM licensing is version-specific).

Z.ai

GLM-5.3

Input: USD 1.4 / 1000000 tokens · Tier unspecifiedOutput: USD 4.4 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

Input: USD 1.4 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.3.
  • Input tokens.

Cache read: USD 0.26 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.3.
  • Cached input tokens. The source lists cached-input storage as "Limited-time Free" (no end date); storage is not priced here.

Output: USD 4.4 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.3.
  • Output tokens.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Z.ai

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-27T13:56:23.785Z · Price source checked: 2026-09-27T13:56:23.785Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: zai-api-glm-5-3-flash: pricing · zai-api-glm-5-3-flash: availability · zai-api-glm-5-3-flashx: pricing · zai-api-glm-5-3-flashx: availability · zai-api-glm-5-3: pricing · zai-api-glm-5-3: availability · zai-api-glm-5-2: pricing · zai-api-glm-5-2: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T13:56:23.785Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 3b5587f7826bf8573174e721613745262cf9a891f8a1408385251ccbba0b5be2; collector zai-international-1.
  • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. API models are not linked to downloadable GLM checkpoints (GLM licensing is version-specific).

Z.ai

GLM-5.3-Flash

Input: USD 0.15 / 1000000 tokens · Tier unspecifiedOutput: USD 0.50 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

Input: USD 0.15 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.3-Flash.
  • Input tokens.

Cache read: USD 0.03 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.3-Flash.
  • Cached input tokens. The source lists cached-input storage as "Limited-time Free" (no end date); storage is not priced here.

Output: USD 0.50 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.3-Flash.
  • Output tokens.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Z.ai

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-27T13:56:23.785Z · Price source checked: 2026-09-27T13:56:23.785Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: zai-api-glm-5-3-flash: pricing · zai-api-glm-5-3-flash: availability · zai-api-glm-5-3-flashx: pricing · zai-api-glm-5-3-flashx: availability · zai-api-glm-5-3: pricing · zai-api-glm-5-3: availability · zai-api-glm-5-2: pricing · zai-api-glm-5-2: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T13:56:23.785Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 3b5587f7826bf8573174e721613745262cf9a891f8a1408385251ccbba0b5be2; collector zai-international-1.
  • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. API models are not linked to downloadable GLM checkpoints (GLM licensing is version-specific).

Z.ai

GLM-5.3-FlashX

Input: USD 0.37 / 1000000 tokens · Tier unspecifiedOutput: USD 1.25 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

Input: USD 0.37 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.3-FlashX.
  • Input tokens.

Cache read: USD 0.075 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.3-FlashX.
  • Cached input tokens. The source lists cached-input storage as "Limited-time Free" (no end date); storage is not priced here.

Output: USD 1.25 / 1000000 tokens · Modality: unspecified

Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Z.ai international platform (docs.z.ai) direct API, USD per 1M tokens; China-platform (BigModel, CNY) pricing is out of scope.
  • Model ID GLM-5.3-FlashX.
  • Output tokens.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: Z.ai

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-27T13:56:23.785Z · Price source checked: 2026-09-27T13:56:23.785Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: zai-api-glm-5-3-flash: pricing · zai-api-glm-5-3-flash: availability · zai-api-glm-5-3-flashx: pricing · zai-api-glm-5-3-flashx: availability · zai-api-glm-5-3: pricing · zai-api-glm-5-3: availability · zai-api-glm-5-2: pricing · zai-api-glm-5-2: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T13:56:23.785Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

  • Raw SHA-256: 3b5587f7826bf8573174e721613745262cf9a891f8a1408385251ccbba0b5be2; collector zai-international-1.
  • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. API models are not linked to downloadable GLM checkpoints (GLM licensing is version-specific).

OpenAI

GPT-5.4 mini

Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
Weights unknownAPI listedDownload not verifiedtextContext: Unknown tokens

Input: USD 0.75 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Short context; numeric threshold is unknown in this source.
  • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
  • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

Cache read: USD 0.075 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Short context; numeric threshold is unknown in this source.
  • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
  • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

Cache write: Unknown / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Short context; numeric threshold is unknown in this source.
  • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
  • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.
  • Source cell is unavailable (-); not a zero price.

Output: USD 4.50 / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Short context; numeric threshold is unknown in this source.
  • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
  • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

Input: Unknown / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Long context; numeric threshold is unknown in this source.
  • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
  • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.
  • Source cell is unavailable (-); not a zero price.

Cache read: Unknown / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Long context; numeric threshold is unknown in this source.
  • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
  • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.
  • Source cell is unavailable (-); not a zero price.

Cache write: Unknown / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Long context; numeric threshold is unknown in this source.
  • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
  • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.
  • Source cell is unavailable (-); not a zero price.

Output: Unknown / 1000000 tokens · Standard · Modality: unspecified

Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

  • Long context; numeric threshold is unknown in this source.
  • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
  • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.
  • Source cell is unavailable (-); not a zero price.

Weight license: Unknown / not verified

Developer: Unknown / not verified · API seller: OpenAI

Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

Official pricing ↗

Price observed: 2026-09-29T04:37:31.867Z · Price source checked: 2026-09-29T04:37:31.867Z · Price freshness: unknown

Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

Source, conditions and dates

Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: identity · openai-gpt-5-4-mini: pricing · openai-gpt-5-4-mini: availability · openai-gpt-5-4-mini: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

    Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

    • The official GPT-5.4 mini model page describes API endpoints, features and model aliases; this review did not establish explicit weight-release or non-downloadability evidence.
    • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.

    Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:27.451Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

    • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
    • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

    Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:11.289Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

    • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
    • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

    Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:31.867Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

    • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
    • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

    OpenAI

    GPT-5.4 nano

    Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
    Weights unknownAPI listedDownload not verifiedtextContext: Unknown tokens

    Input: USD 0.20 / 1000000 tokens · Standard · Modality: unspecified

    Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

    • Short context; numeric threshold is unknown in this source.
    • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
    • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
    • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

    Cache read: USD 0.02 / 1000000 tokens · Standard · Modality: unspecified

    Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

    • Short context; numeric threshold is unknown in this source.
    • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
    • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
    • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

    Cache write: Unknown / 1000000 tokens · Standard · Modality: unspecified

    Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

    • Short context; numeric threshold is unknown in this source.
    • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
    • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
    • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.
    • Source cell is unavailable (-); not a zero price.

    Output: USD 1.25 / 1000000 tokens · Standard · Modality: unspecified

    Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

    • Short context; numeric threshold is unknown in this source.
    • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
    • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
    • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

    Input: Unknown / 1000000 tokens · Standard · Modality: unspecified

    Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

    • Long context; numeric threshold is unknown in this source.
    • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
    • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
    • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.
    • Source cell is unavailable (-); not a zero price.

    Cache read: Unknown / 1000000 tokens · Standard · Modality: unspecified

    Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

    • Long context; numeric threshold is unknown in this source.
    • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
    • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
    • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.
    • Source cell is unavailable (-); not a zero price.

    Cache write: Unknown / 1000000 tokens · Standard · Modality: unspecified

    Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

    • Long context; numeric threshold is unknown in this source.
    • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
    • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
    • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.
    • Source cell is unavailable (-); not a zero price.

    Output: Unknown / 1000000 tokens · Standard · Modality: unspecified

    Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

    • Long context; numeric threshold is unknown in this source.
    • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
    • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
    • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.
    • Source cell is unavailable (-); not a zero price.

    Weight license: Unknown / not verified

    Developer: Unknown / not verified · API seller: OpenAI

    Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

    Official pricing ↗

    Price observed: 2026-09-29T04:37:31.867Z · Price source checked: 2026-09-29T04:37:31.867Z · Price freshness: unknown

    Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

    Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

    Source, conditions and dates

    Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-nano: identity · openai-gpt-5-4-nano: pricing · openai-gpt-5-4-nano: availability · openai-gpt-5-4-nano: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-nano: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

      • The official GPT-5.4 nano model page describes API endpoints, features and model aliases; this review did not establish explicit weight-release or non-downloadability evidence.
      • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:27.451Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:11.289Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:31.867Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      OpenAI

      GPT-5.6 Luna

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedreasoningContext: Unknown tokens

      Input: USD 0.20 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 0.02 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 0.25 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 1.20 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Input: USD 0.40 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 0.04 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 0.50 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 1.80 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: OpenAI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-29T04:37:31.867Z · Price source checked: 2026-09-29T04:37:31.867Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-6-luna: identity · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-luna: availability · openai-gpt-5-6-luna: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Standard short-context rate; cache writes and long-context requests have separate prices.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-6-luna: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

      • The official GPT-5.6 Luna model page describes API endpoints, features and model aliases; this review did not establish explicit weight-release or non-downloadability evidence.
      • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:27.451Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:11.289Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:31.867Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      OpenAI

      GPT-5.6 Sol

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedreasoningContext: Unknown tokens

      Input: USD 4.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 0.40 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 5.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 20.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Input: USD 8.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 0.80 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 10.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 30.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: OpenAI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-29T04:37:31.867Z · Price source checked: 2026-09-29T04:37:31.867Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-6-sol: identity · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-sol: availability · openai-gpt-5-6-sol: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Standard short-context rate; cache writes and long-context requests have separate prices.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-6-sol: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

      • The official GPT-5.6 Sol model page describes API endpoints, features and model aliases; this review did not establish explicit weight-release or non-downloadability evidence.
      • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:27.451Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:11.289Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:31.867Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      OpenAI

      GPT-5.6 Terra

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedreasoningContext: Unknown tokens

      Input: USD 2.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 0.20 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 2.50 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 12.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Input: USD 4.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 0.40 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 5.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 18.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: OpenAI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-29T04:37:31.867Z · Price source checked: 2026-09-29T04:37:31.867Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-6-terra: identity · openai-gpt-5-6-terra: pricing · openai-gpt-5-6-terra: availability · openai-gpt-5-6-terra: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Standard short-context rate; cache writes and long-context requests have separate prices.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-6-terra: weightAccess · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

      • The official GPT-5.6 Terra model page describes API endpoints, features and model aliases; this review did not establish explicit weight-release or non-downloadability evidence.
      • Review outcome: unknown. No explicit model-specific weight-availability statement was established in the reviewed source; this is not evidence of non-downloadability. No download URL or weight license is asserted.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:27.451Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:11.289Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:31.867Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      OpenAI

      GPT-6 Astra

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 10.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 1.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 12.50 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 50.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Input: USD 20.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 2.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 25.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 75.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: OpenAI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-29T04:37:31.867Z · Price source checked: 2026-09-29T04:37:31.867Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:27.451Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:11.289Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:31.867Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      OpenAI

      GPT-6 Luna

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.10 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 0.01 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 0.125 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 0.50 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Input: USD 0.20 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 0.02 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 0.25 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 0.75 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: OpenAI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-29T04:37:31.867Z · Price source checked: 2026-09-29T04:37:31.867Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:27.451Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:11.289Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:31.867Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      OpenAI

      GPT-6 Sol

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 2.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 0.20 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 2.50 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 10.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Short context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Input: USD 4.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache read: USD 0.40 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Cache write: USD 5.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Output: USD 15.00 / 1000000 tokens · Standard · Modality: unspecified

      Region: global · Endpoint scope: global · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Long context; numeric threshold is unknown in this source.
      • OpenAI direct Standard only; excludes Batch, Flex, Fast and AWS billing. Regional eligibility is not established.
      • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. For GPT-6 Sol and Luna, EU data residency is available only with Standard processing. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
      • FedRAMP endpoints are charged a 10% uplift over the corresponding standard model rates.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: OpenAI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-29T04:37:31.867Z · Price source checked: 2026-09-29T04:37:31.867Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:40:27.451Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-28T04:40:11.289Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: openai-gpt-5-4-mini: pricing · openai-gpt-5-4-nano: pricing · openai-gpt-5-6-luna: pricing · openai-gpt-5-6-sol: pricing · openai-gpt-5-6-terra: pricing · openai-gpt-6-astra: pricing · openai-gpt-6-sol: pricing · openai-gpt-6-luna: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-29T04:37:31.867Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 0de899a93b1d9a8c7cf7b3a01ed3553337636c7ff374153821f218a203a203c8; collector openai-standard-2.
      • Pricing observation only; no new access, availability, context threshold, regional eligibility or release metadata assertions.

      Checkpoint: KAT-Coder-V2.5-Dev

      Kwaipilot / KAT

      KAT-Coder-V2.5-Dev

      No API price listedSelf-hosting costs not included
      Weights unknownDownload not verifiedtextContext: 262,144 tokens

      Weight license: Unknown / not verified

      Developer: Kwaipilot / KAT

      Checkpoint: KAT-Coder-V2.5-Dev · Revision: 7be56fe773e72b6f5ca93c1ae45d828ddb893922 · Lineage: post-trained-from: Qwen3.6-35B-A3B (lineage reference)

      Creator model source ↗

      Model metadata observed: 2026-09-27T04:40:13.000Z · Metadata freshness: fresh

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

      Active parameters: 3B; total weights: 35B.

      Source, conditions and dates

      Primary source ↗ · huggingface-model-card · Claims: kat-coder-v2-5-dev: identity · kat-coder-v2-5-dev: specifications · kat-coder-v2-5-dev: download · kat-coder-v2-5-dev: weightAccess · kat-coder-v2-5-dev: gating · developer:kwaipilot: identity · lineage:qwen3-6-35b-a3b: identity · kat-coder-v2-5-dev: lineage:post-trained-from:lineage:qwen3-6-35b-a3b · Retrieved · Effective: Unknown · Checked: 2026-09-27T04:40:13.000Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

      • Exact card explicitly ships language-only weights without vision components. Native context is262144; optional YaRN scaling is separate. No KAT-Pro hosted-offer join is inferred.
      • The card explicitly describes post-training from Qwen3.6-35B-A3B. That parent is represented only as an external lineage reference; its other specifications and access are not inferred. Apache-2.0 YAML does not establish separate legal scopes; preserve the #92 unknown decisions. Kwaipilot/KAT branding is observed; legal-entity aliases remain unverified.
      • Official checkpoint/download reference: https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev. No model weights downloaded; access/download availability remains unverified.
      • Document revision kind repository-commit; revision 7be56fe773e72b6f5ca93c1ae45d828ddb893922; SHA256 e9b49bdc7b5f00a5a1f985143eb43148304a0a20cf0f3e4e935bd8d369897561; private manifest ai/v1/.access-evidence-development/36294950609/access-kat-coder-v2-5-dev-readme-md/2026-09-27/complete.json.gz; manifest SHA256 be03c87e7e3e7e8810d87a57543a477919c2681a5ae4c0f4faec30b57bb79d69.

      Moonshot AI

      Kimi K2.6

      Input: USD 0.95 / 1000000 tokens · Tier unspecifiedOutput: USD 4.00 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
      Open weightsAPI listedDownloadablereasoningContext: 262,144 tokens

      Input: USD 0.95 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.
      • Input price on cache miss.

      Cache read: USD 0.16 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.
      • Input that hits the context cache is billed at this price.

      Output: USD 4.00 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.

      Weight license: Modified MIT · Hugging Face / download weights ↗

      Developer: Unknown / not verified · API seller: Moonshot AI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗ Weight-license evidence ↗

      Price observed: 2026-09-27T00:43:17.997Z · Price source checked: 2026-09-27T00:43:17.997Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k2-6-international: identity · moonshot-kimi-k2-6-international: pricing · moonshot-kimi-k2-6-international: availability · moonshot-kimi-k2-6-international: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • International direct-API rate; reasoning and preserved historical reasoning consume billed tokens.

      Primary source ↗ · huggingface-model-card · Claims: moonshot-kimi-k2-6-international: weightAccess · moonshot-kimi-k2-6-international: download · moonshot-kimi-k2-6-international: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

      • The creator publishes the model weights under the Modified MIT license.
      • Open weights are downloadable under the model license; this is not a claim of unrestricted open-source licensing or that the model fits consumer hardware.

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k3-international: pricing · moonshot-kimi-k2-7-code-international: pricing · moonshot-kimi-k2-7-code-highspeed-international: pricing · moonshot-kimi-k2-6-international: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T01:43:21.824Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: fc5aeeff9664371c5391ace28f8594a86250e0f363a74b250e783adb4e3c88bc; collector kimi-international-1.
      • Pricing observation only; no new access, availability, context, reasoning-billing or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k3-international: pricing · moonshot-kimi-k2-7-code-international: pricing · moonshot-kimi-k2-7-code-highspeed-international: pricing · moonshot-kimi-k2-6-international: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:43:17.997Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: fc5aeeff9664371c5391ace28f8594a86250e0f363a74b250e783adb4e3c88bc; collector kimi-international-2.
      • Pricing observation only; no new access, availability, context, reasoning-billing or release metadata assertions.

      Moonshot AI

      Kimi K2.7 Code

      Input: USD 0.95 / 1000000 tokens · Tier unspecifiedOutput: USD 4.00 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
      Open weightsAPI listedDownloadablereasoningContext: 262,144 tokens

      Input: USD 0.95 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.
      • Input price on cache miss.

      Cache read: USD 0.19 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.
      • Input that hits the context cache is billed at this price.

      Output: USD 4.00 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.

      Weight license: Modified MIT · Hugging Face / download weights ↗

      Developer: Unknown / not verified · API seller: Moonshot AI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗ Weight-license evidence ↗

      Price observed: 2026-09-27T00:43:17.997Z · Price source checked: 2026-09-27T00:43:17.997Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k2-7-code-international: identity · moonshot-kimi-k2-7-code-international: pricing · moonshot-kimi-k2-7-code-international: availability · moonshot-kimi-k2-7-code-international: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • International direct-API rate; thinking is always enabled and billed.

      Primary source ↗ · huggingface-model-card · Claims: moonshot-kimi-k2-7-code-international: weightAccess · moonshot-kimi-k2-7-code-international: download · moonshot-kimi-k2-7-code-international: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

      • The creator publishes the model weights under the Modified MIT license.
      • Open weights are downloadable under the model license; this is not a claim of unrestricted open-source licensing or that the model fits consumer hardware.

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k3-international: pricing · moonshot-kimi-k2-7-code-international: pricing · moonshot-kimi-k2-7-code-highspeed-international: pricing · moonshot-kimi-k2-6-international: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T01:43:21.824Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: fc5aeeff9664371c5391ace28f8594a86250e0f363a74b250e783adb4e3c88bc; collector kimi-international-1.
      • Pricing observation only; no new access, availability, context, reasoning-billing or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k3-international: pricing · moonshot-kimi-k2-7-code-international: pricing · moonshot-kimi-k2-7-code-highspeed-international: pricing · moonshot-kimi-k2-6-international: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:43:17.997Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: fc5aeeff9664371c5391ace28f8594a86250e0f363a74b250e783adb4e3c88bc; collector kimi-international-2.
      • Pricing observation only; no new access, availability, context, reasoning-billing or release metadata assertions.

      Moonshot AI

      Kimi K2.7 Code Highspeed

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Open weightsAPI listedDownloadablereasoningContext: 262,144 tokens

      Input: USD 1.90 / 1000000 tokens · highspeed · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: highspeed · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.
      • Input price on cache miss.

      Cache read: USD 0.38 / 1000000 tokens · highspeed · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: highspeed · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.
      • Input that hits the context cache is billed at this price.

      Output: USD 8.00 / 1000000 tokens · highspeed · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: highspeed · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.

      Weight license: Modified MIT · Hugging Face / download weights ↗

      Developer: Unknown / not verified · API seller: Moonshot AI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗ Weight-license evidence ↗

      Price observed: 2026-09-27T00:43:17.997Z · Price source checked: 2026-09-27T00:43:17.997Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k2-7-code-highspeed-international: identity · moonshot-kimi-k2-7-code-highspeed-international: pricing · moonshot-kimi-k2-7-code-highspeed-international: availability · moonshot-kimi-k2-7-code-highspeed-international: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Higher-throughput K2.7 Code variant; international direct-API rate.

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k2-7-code-highspeed-international: weightAccess · moonshot-kimi-k2-7-code-highspeed-international: download · moonshot-kimi-k2-7-code-highspeed-international: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

      • The provider explicitly describes HighSpeed as the same model as Kimi K2.7 Code with a faster serving tier.
      • The download link is the underlying Kimi K2.7 Code model, not a separate HighSpeed checkpoint; self-hosting does not reproduce the hosted speed guarantee.
      • Kimi K2.7 Code weights use the Modified MIT license, as documented in the linked creator model card.

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k3-international: pricing · moonshot-kimi-k2-7-code-international: pricing · moonshot-kimi-k2-7-code-highspeed-international: pricing · moonshot-kimi-k2-6-international: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T01:43:21.824Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: fc5aeeff9664371c5391ace28f8594a86250e0f363a74b250e783adb4e3c88bc; collector kimi-international-1.
      • Pricing observation only; no new access, availability, context, reasoning-billing or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k3-international: pricing · moonshot-kimi-k2-7-code-international: pricing · moonshot-kimi-k2-7-code-highspeed-international: pricing · moonshot-kimi-k2-6-international: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:43:17.997Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: fc5aeeff9664371c5391ace28f8594a86250e0f363a74b250e783adb4e3c88bc; collector kimi-international-2.
      • Pricing observation only; no new access, availability, context, reasoning-billing or release metadata assertions.

      Moonshot AI

      Kimi K3

      Input: USD 3.00 / 1000000 tokens · Tier unspecifiedOutput: USD 15.00 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
      Open weightsAPI listedDownloadablereasoningContext: 1,048,576 tokens

      Input: USD 3.00 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.
      • Input price; input that hits the context cache is billed at the cached input price instead.

      Cache read: USD 0.30 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.
      • Input that hits the context cache is billed at this price.

      Output: USD 15.00 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.

      Cache write: USD 3.00 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: 300 seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.
      • Cache write, 5min TTL tier (default when no TTL is specified).

      Cache write: USD 6.00 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: 3600 seconds · Valid through: Unknown

      • Kimi international platform direct API (USD per 1M tokens); China-mainland pricing is out of scope.
      • Prices exclude applicable taxes; tax is calculated at checkout by jurisdiction.
      • Cache write, 1h TTL tier.

      Weight license: Kimi K3 License · Hugging Face / download weights ↗

      Developer: Unknown / not verified · API seller: Moonshot AI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗ Weight-license evidence ↗

      Price observed: 2026-09-27T00:43:17.997Z · Price source checked: 2026-09-27T00:43:17.997Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k3-international: identity · moonshot-kimi-k3-international: pricing · moonshot-kimi-k3-international: availability · moonshot-kimi-k3-international: specifications · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • International direct-API rate; reasoning content is billed and applicable taxes are excluded.

      Primary source ↗ · huggingface-model-card · Claims: moonshot-kimi-k3-international: weightAccess · moonshot-kimi-k3-international: download · moonshot-kimi-k3-international: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: Unknown

      • The creator publishes the model weights under the Kimi K3 License license.
      • Open weights are downloadable under the model license; this is not a claim of unrestricted open-source licensing or that the model fits consumer hardware.

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k3-international: pricing · moonshot-kimi-k2-7-code-international: pricing · moonshot-kimi-k2-7-code-highspeed-international: pricing · moonshot-kimi-k2-6-international: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-26T01:43:21.824Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: fc5aeeff9664371c5391ace28f8594a86250e0f363a74b250e783adb4e3c88bc; collector kimi-international-1.
      • Pricing observation only; no new access, availability, context, reasoning-billing or release metadata assertions.

      Primary source ↗ · vendor-docs · Claims: moonshot-kimi-k3-international: pricing · moonshot-kimi-k2-7-code-international: pricing · moonshot-kimi-k2-7-code-highspeed-international: pricing · moonshot-kimi-k2-6-international: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:43:17.997Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: fc5aeeff9664371c5391ace28f8594a86250e0f363a74b250e783adb4e3c88bc; collector kimi-international-2.
      • Pricing observation only; no new access, availability, context, reasoning-billing or release metadata assertions.

      Checkpoint: Ling-lite-1.5

      InclusionAI

      Ling-lite-1.5

      No API price listedSelf-hosting costs not included
      Weights unknownDownload not verifiedtextContext: Unknown tokens

      Weight license: Unknown / not verified

      Developer: InclusionAI

      Checkpoint: Ling-lite-1.5 · Revision: ef1ac33ce4c3dca4bc01873b2030e6a264c5ef99 · Lineage: Unknown / not verified

      Creator model source ↗

      Model metadata observed: 2026-09-27T05:30:33.065Z · Metadata freshness: fresh

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: MIT (creator-declared code license) · API availability: unknown (dated observation, not a live service check)

      Active parameters: 2.75B; total weights: 16.8B.

      Source, conditions and dates

      Primary source ↗ · huggingface-model-card · Claims: ling-lite-1-5: identity · ling-lite-1-5: specifications · ling-lite-1-5: download · ling-lite-1-5: weightAccess · ling-lite-1-5: gating · ling-lite-1-5: codeLicense · developer:inclusionai: identity · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:30:33.065Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

      • Creator names InclusionAI and explicitly declares MIT for the code repository, linking a separate LICENCE document not acquired here. No weight-license scope is inferred from YAML. Advertised128K remains an imprecise context note.
      • Bailing appears only in an asset filename; no separate Bailing company or legal alias is created. Corporate affiliation remains unverified by these documents.
      • Official checkpoint/download reference: https://huggingface.co/inclusionAI/Ling-lite-1.5. No model weights downloaded; access/download availability remains unverified.
      • Document revision kind repository-commit; revision ef1ac33ce4c3dca4bc01873b2030e6a264c5ef99; SHA256 6ca74720f6f064d4946fd7682e473cdb12e24a9e224fd27e25cae975f45e2a8f; private manifest ai/v1/onboard-ling-lite-1-5-readme-md/2026-09-27/complete.json.gz; manifest SHA256 61dcbc8fac4380ed1d452f792e2ce2d99d7fc18bbff1cefdf3900cea9f9d9ca8.

      Checkpoint: Llama 3.2 11B Vision

      Meta · Llama 3.2

      Llama 3.2 11B Vision

      No API price listedSelf-hosting costs not included
      Open weightsDownloadabletext · imageContext: 131,072 tokens

      Weight license: Llama 3.2 Community License · Hugging Face / download weights ↗

      10.6B parameters · Estimated weight memory: Q4 7.9 GB · Q8 14.3 GB. Includes 20% loading overhead; excludes KV cache and runtime memory. Compare hardware.

      Developer: Meta

      Checkpoint: Llama 3.2 11B Vision · Revision: 9eb2daaa8597bf192a8b0e73f848f3a102794df5 · Lineage: Unknown / not verified

      Creator model source ↗ Creator model source ↗ Weight-license evidence ↗ Weight-license evidence ↗

      Model metadata observed: 2026-09-27T00:02:21.072Z / 2026-09-27T00:02:26.919Z · Metadata freshness: fresh

      Weight-license reviewed: 2026-09-27T02:50:26.221Z. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: gated · Code license: Llama 3.2 Community License · API availability: unknown (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-11b-vision-instruct: identity · meta-llama-3-2-11b-vision-instruct: specifications · meta-llama-3-2-11b-vision-instruct: weightAccess · meta-llama-3-2-11b-vision-instruct: download · meta-llama-3-2-11b-vision-instruct: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Image-plus-text is officially supported only in English; custom license restrictions apply.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-1b-instruct: identity · meta-llama-3-2-1b-instruct: specifications · meta-llama-3-2-1b-instruct: weightAccess · meta-llama-3-2-1b-instruct: download · meta-llama-3-2-1b-instruct: weightLicense · meta-llama-3-2-1b-instruct: gating · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Custom commercial license with gated access and acceptable-use policy; actual architecture count is approximately 1.23B.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-3-70b-instruct: identity · meta-llama-3-3-70b-instruct: specifications · meta-llama-3-3-70b-instruct: weightAccess · meta-llama-3-3-70b-instruct: download · meta-llama-3-3-70b-instruct: weightLicense · meta-llama-3-3-70b-instruct: gating · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Custom commercial license, gated access, and acceptable-use policy apply.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-11b-vision-instruct: identity · meta-llama-3-2-11b-vision-instruct: specifications · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:21.072Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Card explicitly lists11B (10.6) and text/image input; text config specifies131,072 positions matching advertised128K.
      • Card describes a Llama3.1 text backbone and vision adapter, but no exact parent exists in the curated registry. Do not infer a link to the curated Llama3.2 1B or Llama3.3 records.
      • Metadata documents were read; weight binaries were not downloaded. Prior weight-access/download observations retain their original dates.
      • Revision 9eb2daaa8597bf192a8b0e73f848f3a102794df5; SHA256 645883e9ef871845396d00d23480affb5612fa80c52d5a73f595580e36cb99c9; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-11b-vision-instruct-readme-md/2026-09-27/complete.json.gz; manifest SHA256 e85ab6068336b7f9b39af56642e3a4e17930a00346156f4cd14a9e9fb5c30af6.

      Raw API evidence ↗ · vendor-api · Claims: meta-llama-3-2-11b-vision-instruct: gating · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:15.334Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Repository gating: manual; listed weight files: true. An inventory and authenticated metadata read do not verify weight download availability.
      • Revision 9eb2daaa8597bf192a8b0e73f848f3a102794df5; SHA256 03cbbc8e38453bf689a186a045543ae12fda34cd7e1d3a403fc5bb1f79e6e0e4; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-11b-vision-instruct-info/2026-09-27/complete.json.gz; manifest SHA256 2ca4005d4e97d1aae99adcaf1e51830d7d8b4703d358572978a1904bfd96cf31.

      Primary source ↗ · vendor-docs · Claims: meta-llama-3-2-11b-vision-instruct: configuration · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:26.919Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Configuration is supporting evidence, not an automatic claim of advertised context or modality. Resolve differences using the explicit publication review.
      • Revision 9eb2daaa8597bf192a8b0e73f848f3a102794df5; SHA256 db708a72a246253fdf8aa3837a5f0b80ac1960692495ce05cb3101833128e845; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-11b-vision-instruct-config-json/2026-09-27/complete.json.gz; manifest SHA256 a0aafb863ce596fbd6651d402cf56f615fe1faf1d9d5434c26e741b1c73a8a29.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-11b-vision-instruct: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:32.768Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • The Llama 3.2 Community License expressly covers model code and trained weights. Preserve agreement, attribution and Built with Llama notices; model naming, acceptable-use and redistribution conditions apply. A separate grant is required above the license's release-date700-million-monthly-active-user threshold. This is a custom commercial license, not an OSI/permissive classification.
      • The model card's incorporated acceptable-use policy excludes license grants for multimodal models to EU-domiciled individuals/EU-principal-place-of-business companies, with an end-user-of-incorporating-product exception. This vision restriction is not transferred to text-only checkpoints.
      • Revision 9eb2daaa8597bf192a8b0e73f848f3a102794df5; SHA256 0b4284c1f87029e67654c7953afa16279961632cf73dcfe33374c4c2f298fa35; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-11b-vision-instruct-license-txt/2026-09-27/complete.json.gz; manifest SHA256 529717ee11e05b0223e2c45b175f4a338accd3b39cb80cd7799276f416e07257.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-11b-vision-instruct: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:21.072Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • The Llama 3.2 Community License expressly covers model code and trained weights. Preserve agreement, attribution and Built with Llama notices; model naming, acceptable-use and redistribution conditions apply. A separate grant is required above the license's release-date700-million-monthly-active-user threshold. This is a custom commercial license, not an OSI/permissive classification.
      • The model card's incorporated acceptable-use policy excludes license grants for multimodal models to EU-domiciled individuals/EU-principal-place-of-business companies, with an end-user-of-incorporating-product exception. This vision restriction is not transferred to text-only checkpoints.
      • Revision 9eb2daaa8597bf192a8b0e73f848f3a102794df5; SHA256 645883e9ef871845396d00d23480affb5612fa80c52d5a73f595580e36cb99c9; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-11b-vision-instruct-readme-md/2026-09-27/complete.json.gz; manifest SHA256 e85ab6068336b7f9b39af56642e3a4e17930a00346156f4cd14a9e9fb5c30af6.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-11b-vision-instruct: codeLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:32.768Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Agreement explicitly includes model code. License/notice, naming, acceptable-use and large-platform terms apply; this does not license unrelated third-party libraries.
      • Revision 9eb2daaa8597bf192a8b0e73f848f3a102794df5; SHA256 0b4284c1f87029e67654c7953afa16279961632cf73dcfe33374c4c2f298fa35; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-11b-vision-instruct-license-txt/2026-09-27/complete.json.gz; manifest SHA256 529717ee11e05b0223e2c45b175f4a338accd3b39cb80cd7799276f416e07257.

      Checkpoint: Llama 3.2 1B Instruct

      Meta · Llama 3.2

      Llama 3.2 1B Instruct

      No API price listedSelf-hosting costs not included
      Open weightsDownloadabletextContext: 131,072 tokens

      Weight license: Llama 3.2 Community License · Hugging Face / download weights ↗

      1.23B parameters · Estimated weight memory: Q4 0.9 GB · Q8 1.7 GB. Includes 20% loading overhead; excludes KV cache and runtime memory. Compare hardware.

      Developer: Meta

      Checkpoint: Llama 3.2 1B Instruct · Revision: 9213176726f574b556790deb65791e0c5aa438b6 · Lineage: Unknown / not verified

      Creator model source ↗ Creator model source ↗ Weight-license evidence ↗ Weight-license evidence ↗

      Model metadata observed: 2026-09-27T00:01:57.164Z / 2026-09-27T00:02:02.850Z · Metadata freshness: fresh

      Weight-license reviewed: 2026-09-27T02:50:26.221Z. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: gated · Code license: Llama 3.2 Community License · API availability: unknown (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-11b-vision-instruct: identity · meta-llama-3-2-11b-vision-instruct: specifications · meta-llama-3-2-11b-vision-instruct: weightAccess · meta-llama-3-2-11b-vision-instruct: download · meta-llama-3-2-11b-vision-instruct: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Image-plus-text is officially supported only in English; custom license restrictions apply.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-1b-instruct: identity · meta-llama-3-2-1b-instruct: specifications · meta-llama-3-2-1b-instruct: weightAccess · meta-llama-3-2-1b-instruct: download · meta-llama-3-2-1b-instruct: weightLicense · meta-llama-3-2-1b-instruct: gating · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Custom commercial license with gated access and acceptable-use policy; actual architecture count is approximately 1.23B.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-3-70b-instruct: identity · meta-llama-3-3-70b-instruct: specifications · meta-llama-3-3-70b-instruct: weightAccess · meta-llama-3-3-70b-instruct: download · meta-llama-3-3-70b-instruct: weightLicense · meta-llama-3-3-70b-instruct: gating · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Custom commercial license, gated access, and acceptable-use policy apply.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-1b-instruct: identity · meta-llama-3-2-1b-instruct: specifications · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:57.164Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Exact text-only card lists1B (1.23B) and128K context; config specifies131,072 positions. Quantized8K variants are different checkpoints and are not substituted.
      • No exact parent checkpoint in the curated registry is established; lineage remains unknown.
      • Metadata documents were read; weight binaries were not downloaded. Prior weight-access/download observations retain their original dates.
      • Revision 9213176726f574b556790deb65791e0c5aa438b6; SHA256 18564977261167ff9f76d8ee3a94c8d1cc59c0e143ba054f1187744086004a93; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-1b-instruct-readme-md/2026-09-27/complete.json.gz; manifest SHA256 96a5ef6d646fcadf7e09eaceeebbd1201abe62126e82d4f0b2871d0058f328da.

      Raw API evidence ↗ · vendor-api · Claims: meta-llama-3-2-1b-instruct: gating · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:51.533Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Repository gating: manual; listed weight files: true. An inventory and authenticated metadata read do not verify weight download availability.
      • Revision 9213176726f574b556790deb65791e0c5aa438b6; SHA256 2ac8c1b9256a31b0f06992a11d5906345878faea58385a6354128c27b5a918fa; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-1b-instruct-info/2026-09-27/complete.json.gz; manifest SHA256 326df90d97d2b2fecdff03d4bedd47efe976d9ce149f7e5a2bf2bd3eac1772f4.

      Primary source ↗ · vendor-docs · Claims: meta-llama-3-2-1b-instruct: configuration · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:02.850Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Configuration is supporting evidence, not an automatic claim of advertised context or modality. Resolve differences using the explicit publication review.
      • Revision 9213176726f574b556790deb65791e0c5aa438b6; SHA256 2febf68cea25bf4611be02b7536f2488a5ba523bb1134986e3610152abe74fdb; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-1b-instruct-config-json/2026-09-27/complete.json.gz; manifest SHA256 59153fc6f9ae97528a5cd1f7740eb01cc3ba1bce7106fbdf2cf833ac0ceb8774.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-1b-instruct: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:08.594Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • The Llama 3.2 Community License expressly covers model code and trained weights. Preserve agreement, attribution and Built with Llama notices; model naming, acceptable-use and redistribution conditions apply. A separate grant is required above the license's release-date700-million-monthly-active-user threshold. This is a custom commercial license, not an OSI/permissive classification.
      • Revision 9213176726f574b556790deb65791e0c5aa438b6; SHA256 0b4284c1f87029e67654c7953afa16279961632cf73dcfe33374c4c2f298fa35; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-1b-instruct-license-txt/2026-09-27/complete.json.gz; manifest SHA256 246157794effc3ec6c494130a7795118f0fecf087fc5ed647a6099147bf5d938.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-1b-instruct: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:57.164Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • The Llama 3.2 Community License expressly covers model code and trained weights. Preserve agreement, attribution and Built with Llama notices; model naming, acceptable-use and redistribution conditions apply. A separate grant is required above the license's release-date700-million-monthly-active-user threshold. This is a custom commercial license, not an OSI/permissive classification.
      • Revision 9213176726f574b556790deb65791e0c5aa438b6; SHA256 18564977261167ff9f76d8ee3a94c8d1cc59c0e143ba054f1187744086004a93; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-1b-instruct-readme-md/2026-09-27/complete.json.gz; manifest SHA256 96a5ef6d646fcadf7e09eaceeebbd1201abe62126e82d4f0b2871d0058f328da.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-1b-instruct: codeLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:08.594Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Agreement explicitly includes model code. License/notice, naming, acceptable-use and large-platform terms apply; this does not license unrelated third-party libraries.
      • Revision 9213176726f574b556790deb65791e0c5aa438b6; SHA256 0b4284c1f87029e67654c7953afa16279961632cf73dcfe33374c4c2f298fa35; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-2-1b-instruct-license-txt/2026-09-27/complete.json.gz; manifest SHA256 246157794effc3ec6c494130a7795118f0fecf087fc5ed647a6099147bf5d938.

      Novita AI

      Llama 3.3 70B Instruct

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.135 / 1000000 tokens · serverless · Modality: text

      Region: Unknown · Endpoint scope: unknown · Tier: serverless · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Novita hosted serverless listing, USD per 1M tokens; standard listed prices only. Batch, priority, dedicated, promotional and free-limited offers are excluded.
      • Native decimal values agree with integer ten-thousandths and original/listed amounts. No tiered billing is reported for this listing.
      • Account and regional eligibility, status-code semantics and effective dates are unverified; listing is not an inference service test.
      • Listed prompt/input token rate; separately listed cache-read rates apply only to cache hits.

      Output: USD 0.4 / 1000000 tokens · serverless · Modality: text

      Region: Unknown · Endpoint scope: unknown · Tier: serverless · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Novita hosted serverless listing, USD per 1M tokens; standard listed prices only. Batch, priority, dedicated, promotional and free-limited offers are excluded.
      • Native decimal values agree with integer ten-thousandths and original/listed amounts. No tiered billing is reported for this listing.
      • Account and regional eligibility, status-code semantics and effective dates are unverified; listing is not an inference service test.
      • Listed completion/output token rate.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: Novita AI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T06:32:08.821Z · Price source checked: 2026-09-27T06:32:08.821Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

      Source, conditions and dates

      Raw API evidence ↗ · vendor-api · Claims: novita-llama-3-3-70b-instruct: pricing · novita-minimax-m2: pricing · novita-glm-5: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T06:32:08.821Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: B

      • Raw SHA-256: 9722d1130261fe9667b226c4817976e42d0087d11cf1eb7b948cecdeb7988f25; original collector novita-model-catalog-1. Pinned parse-failure recovery preserves original dates; no new acquisition.
      • USD/scale interpretation reviewed against Novita official pricing and its MiniMax M3 pricing explanation; exact original/decimal/integer values retained privately.
      • Three curated hosted listings only. No verified checkpoint/developer/license/weight-access link; availability remains unknown because numeric status semantics are undocumented.

      Checkpoint: Llama 3.3 70B Instruct

      Meta · Llama 3.3

      Llama 3.3 70B Instruct

      No API price listedSelf-hosting costs not included
      Open weightsDownloadabletextContext: 131,072 tokens

      Weight license: Llama 3.3 Community License · Hugging Face / download weights ↗

      70B parameters · Estimated weight memory: Q4 52.5 GB · Q8 94.5 GB. Includes 20% loading overhead; excludes KV cache and runtime memory. Compare hardware.

      Developer: Meta

      Checkpoint: Llama 3.3 70B Instruct · Revision: 6f6073b423013f6a7d4d9f39144961bfbfbc386b · Lineage: Unknown / not verified

      Creator model source ↗ Creator model source ↗ Weight-license evidence ↗ Weight-license evidence ↗

      Model metadata observed: 2026-09-27T00:02:44.948Z / 2026-09-27T00:02:50.649Z · Metadata freshness: fresh

      Weight-license reviewed: 2026-09-27T02:50:26.221Z. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: gated · Code license: Llama 3.3 Community License · API availability: unknown (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-11b-vision-instruct: identity · meta-llama-3-2-11b-vision-instruct: specifications · meta-llama-3-2-11b-vision-instruct: weightAccess · meta-llama-3-2-11b-vision-instruct: download · meta-llama-3-2-11b-vision-instruct: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Image-plus-text is officially supported only in English; custom license restrictions apply.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-2-1b-instruct: identity · meta-llama-3-2-1b-instruct: specifications · meta-llama-3-2-1b-instruct: weightAccess · meta-llama-3-2-1b-instruct: download · meta-llama-3-2-1b-instruct: weightLicense · meta-llama-3-2-1b-instruct: gating · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Custom commercial license with gated access and acceptable-use policy; actual architecture count is approximately 1.23B.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-3-70b-instruct: identity · meta-llama-3-3-70b-instruct: specifications · meta-llama-3-3-70b-instruct: weightAccess · meta-llama-3-3-70b-instruct: download · meta-llama-3-3-70b-instruct: weightLicense · meta-llama-3-3-70b-instruct: gating · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Custom commercial license, gated access, and acceptable-use policy apply.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-3-70b-instruct: identity · meta-llama-3-3-70b-instruct: specifications · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:44.948Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Exact creator card advertises70B and text-only capabilities; config specifies131,072 positions matching advertised128K.
      • Instruction tuning is documented but an exact parent in the curated registry is not established.
      • Metadata documents were read; weight binaries were not downloaded. Prior weight-access/download observations retain their original dates.
      • Revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b; SHA256 83bfcf4ac40ab3f4d30577e669599e94955b26d6926861410e75e43a07772dad; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-3-70b-instruct-readme-md/2026-09-27/complete.json.gz; manifest SHA256 839e8f11cc929f6eebf990c07599ab526ef7669e774ab198af85f6e18031c3cf.

      Raw API evidence ↗ · vendor-api · Claims: meta-llama-3-3-70b-instruct: gating · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:39.300Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Repository gating: manual; listed weight files: true. An inventory and authenticated metadata read do not verify weight download availability.
      • Revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b; SHA256 e0c892a01a538bdfc9326308526b2b492676bd107e897b03c264098beb1baa31; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-3-70b-instruct-info/2026-09-27/complete.json.gz; manifest SHA256 6a3d88957a7b82026126fbdfd2fcc827dc7e6e6392250794180d9cb42d17063c.

      Primary source ↗ · vendor-docs · Claims: meta-llama-3-3-70b-instruct: configuration · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:50.649Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Configuration is supporting evidence, not an automatic claim of advertised context or modality. Resolve differences using the explicit publication review.
      • Revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b; SHA256 95ef9768e4741543dbfaf0c274f101855883ff338b235c99eca2b6a4f4abee12; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-3-70b-instruct-config-json/2026-09-27/complete.json.gz; manifest SHA256 75caae86015102e20719ca54cb4ac4701a70f260de636ba0eb2d2dccde882899.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-3-70b-instruct: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:56.290Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • The Llama 3.3 Community License expressly covers model code and trained weights. Preserve agreement, attribution and Built with Llama notices; model naming, acceptable-use and redistribution conditions apply. A separate grant is required above the license's release-date700-million-monthly-active-user threshold. This is a custom commercial license, not an OSI/permissive classification.
      • Revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b; SHA256 fb58d9a630ccc1dc0d08f8c00232de56bc73309020f59c8181d0c12ef28d9f8c; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-3-70b-instruct-license/2026-09-27/complete.json.gz; manifest SHA256 fc5aebf5745a177e7b6e6606c61f4cd0266bd6e6d2f4661e2ab7960e9579d960.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-3-70b-instruct: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:44.948Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • The Llama 3.3 Community License expressly covers model code and trained weights. Preserve agreement, attribution and Built with Llama notices; model naming, acceptable-use and redistribution conditions apply. A separate grant is required above the license's release-date700-million-monthly-active-user threshold. This is a custom commercial license, not an OSI/permissive classification.
      • Revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b; SHA256 83bfcf4ac40ab3f4d30577e669599e94955b26d6926861410e75e43a07772dad; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-3-70b-instruct-readme-md/2026-09-27/complete.json.gz; manifest SHA256 839e8f11cc929f6eebf990c07599ab526ef7669e774ab198af85f6e18031c3cf.

      Primary source ↗ · huggingface-model-card · Claims: meta-llama-3-3-70b-instruct: codeLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:02:56.290Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Agreement explicitly includes model code. License/notice, naming, acceptable-use and large-platform terms apply; this does not license unrelated third-party libraries.
      • Revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b; SHA256 fb58d9a630ccc1dc0d08f8c00232de56bc73309020f59c8181d0c12ef28d9f8c; private manifest ai/v1/.model-metadata-development/36281114084/hf-meta-llama-3-3-70b-instruct-license/2026-09-27/complete.json.gz; manifest SHA256 fc5aebf5745a177e7b6e6606c61f4cd0266bd6e6d2f4661e2ab7960e9579d960.

      Checkpoint: LongCat-2.0

      Meituan LongCat

      LongCat-2.0

      No API price listedSelf-hosting costs not included
      Weights unknownDownload not verifiedtextContext: Unknown tokens

      Weight license: MIT

      Developer: Meituan LongCat

      Checkpoint: LongCat-2.0 · Revision: Unknown / not verified · Lineage: Unknown / not verified

      Creator model source ↗ Weight-license evidence ↗ Weight-license evidence ↗

      Model metadata observed: 2026-09-27T05:31:06.669Z · Metadata freshness: fresh

      Weight-license reviewed: 2026-09-27T05:40:06.050Z. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: MIT · API availability: unknown (dated observation, not a live service check)

      Active parameters: 48B; total weights: 1600B.

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: longcat-2-0: identity · longcat-2-0: specifications · longcat-2-0: download · longcat-2-0: weightAccess · longcat-2-0: gating · developer:meituan-longcat: identity · longcat-2-0: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:31:06.669Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

      • Creator explicitly releases model weights under MIT and LICENSE credits Meituan. Preserve notices; the card states no grant of Meituan trademark/patent rights and does not alter MIT terms.
      • Advertised1.6T total/~48B active values are approximate;1M-context training does not establish an exact maximum. Architectural inheritance is not a checkpoint lineage claim.
      • Official checkpoint/download reference: https://huggingface.co/meituan-longcat/LongCat-2.0. No model weights downloaded; access/download availability remains unverified.
      • Document revision kind git-blob; revision bd890adec0d3287930d54f248e02dd1c568f2642; SHA256 5c3272245c9ed3281fddf8d2834d020aed782fdc57029307b6f85f1cea0e7b62; private manifest ai/v1/onboard-longcat-2-0-readme-md/2026-09-27/complete.json.gz; manifest SHA256 837b18f260c7ea5a52788c4d3c14626afe9714e1e8613780f379d67ec09bac5a.

      Primary source ↗ · vendor-docs · Claims: longcat-2-0: codeLicense · longcat-2-0: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:31:11.759Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

      • Creator explicitly releases model weights under MIT and LICENSE credits Meituan. Preserve notices; the card states no grant of Meituan trademark/patent rights and does not alter MIT terms.
      • Advertised1.6T total/~48B active values are approximate;1M-context training does not establish an exact maximum. Architectural inheritance is not a checkpoint lineage claim.
      • Official checkpoint/download reference: https://huggingface.co/meituan-longcat/LongCat-2.0. No model weights downloaded; access/download availability remains unverified.
      • Document revision kind git-blob; revision dd3cf1c68e542dacdf0dad791035dddaa52e5c2a; SHA256 8547d477557f947f322a8d2d479fb797f815800f2cd51b21d0b7a0b3b2cac707; private manifest ai/v1/onboard-longcat-2-0-license/2026-09-27/complete.json.gz; manifest SHA256 4f87262e9f4a7d5bdbb68e4fc69bac7acd5638e67a069ecf024dc24d558f202c.

      Checkpoint: MiMo-V2.5

      Xiaomi MiMo

      MiMo-V2.5

      No API price listedSelf-hosting costs not included
      Weights unknownDownload not verifiedtext · image · video · audioContext: Unknown tokens

      Weight license: Unknown / not verified

      Developer: Xiaomi MiMo

      Checkpoint: MiMo-V2.5 · Revision: 63651580ca774f8504f676040460aed3e1244ac1 · Lineage: Unknown / not verified

      Creator model source ↗

      Model metadata observed: 2026-09-27T05:31:01.491Z · Metadata freshness: fresh

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

      Active parameters: 15B; total weights: 310B.

      Source, conditions and dates

      Primary source ↗ · huggingface-model-card · Claims: mimo-v2-5: identity · mimo-v2-5: specifications · mimo-v2-5: download · mimo-v2-5: weightAccess · mimo-v2-5: gating · developer:xiaomi-mimo: identity · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:31:01.491Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

      • Creator explicitly supports text, image, video and audio understanding, not a claim of audio/video generation. Advertised1M remains a note rather than an invented exact token count.
      • MIT is a metadata declaration without an acquired LICENSE or separately established code/weight terms. Architectural inheritance from MiMo-V2-Flash is not asserted as a checkpoint lineage edge. Xiaomi branding, official platform links and xiaomi.com contact are observed; no price is inferred.
      • Official checkpoint/download reference: https://huggingface.co/XiaomiMiMo/MiMo-V2.5. No model weights downloaded; access/download availability remains unverified.
      • Document revision kind repository-commit; revision 63651580ca774f8504f676040460aed3e1244ac1; SHA256 8c54ab372168ccf760a0fec8b8756b1405d0e137e2de1b5799e1ec0a2dc8a13f; private manifest ai/v1/onboard-mimo-v2-5-readme-md/2026-09-27/complete.json.gz; manifest SHA256 087974cf0c7e3a3d5841f90afc0a9aa2d6f2731cb4a0a744f936216b140d2aa9.

      Xiaomi MiMo

      mimo-v2.6-flash

      Input: USD 0.14 / 1000000 tokens · Tier unspecifiedOutput: USD 0.28 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.14 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Xiaomi MiMo API Open Platform overseas pay-as-you-go price (Real-time API), USD per 1M tokens. Domestic (China, CNY) prices, the Batch API and Token Plan subscriptions are out of scope; web search is billed separately per call.
      • Model ID mimo-v2.6-flash.
      • Input tokens that miss the prompt cache ("Input (Cache Miss)").

      Cache read: USD 0.0028 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Xiaomi MiMo API Open Platform overseas pay-as-you-go price (Real-time API), USD per 1M tokens. Domestic (China, CNY) prices, the Batch API and Token Plan subscriptions are out of scope; web search is billed separately per call.
      • Model ID mimo-v2.6-flash.
      • Input tokens that hit the prompt cache ("Input (Cache Hit)"). The source lists cache writes as "Limited-time Free" (no end date); cache writes are not priced here.

      Output: USD 0.28 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Xiaomi MiMo API Open Platform overseas pay-as-you-go price (Real-time API), USD per 1M tokens. Domestic (China, CNY) prices, the Batch API and Token Plan subscriptions are out of scope; web search is billed separately per call.
      • Model ID mimo-v2.6-flash.
      • Output tokens.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: Xiaomi MiMo

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T18:03:16.076Z · Price source checked: 2026-09-27T18:03:16.076Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: xiaomi-mimo-api-v2-6-pro: pricing · xiaomi-mimo-api-v2-6-pro: availability · xiaomi-mimo-api-v2-6-flash: pricing · xiaomi-mimo-api-v2-6-flash: availability · xiaomi-mimo-api-v2-6-pro-ultraspeed: pricing · xiaomi-mimo-api-v2-6-pro-ultraspeed: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T18:03:16.076Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 64dbe58790cfa2331c3df7df8a6948f53ec9b55dc685d94bd055895a3447f483; collector xiaomi-mimo-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. API models are not linked to downloadable MiMo checkpoints.

      Xiaomi MiMo

      mimo-v2.6-pro

      Input: USD 0.435 / 1000000 tokens · Tier unspecifiedOutput: USD 0.87 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.435 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Xiaomi MiMo API Open Platform overseas pay-as-you-go price (Real-time API), USD per 1M tokens. Domestic (China, CNY) prices, the Batch API and Token Plan subscriptions are out of scope; web search is billed separately per call.
      • Model ID mimo-v2.6-pro.
      • Input tokens that miss the prompt cache ("Input (Cache Miss)").

      Cache read: USD 0.0036 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Xiaomi MiMo API Open Platform overseas pay-as-you-go price (Real-time API), USD per 1M tokens. Domestic (China, CNY) prices, the Batch API and Token Plan subscriptions are out of scope; web search is billed separately per call.
      • Model ID mimo-v2.6-pro.
      • Input tokens that hit the prompt cache ("Input (Cache Hit)"). The source lists cache writes as "Limited-time Free" (no end date); cache writes are not priced here.

      Output: USD 0.87 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Xiaomi MiMo API Open Platform overseas pay-as-you-go price (Real-time API), USD per 1M tokens. Domestic (China, CNY) prices, the Batch API and Token Plan subscriptions are out of scope; web search is billed separately per call.
      • Model ID mimo-v2.6-pro.
      • Output tokens.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: Xiaomi MiMo

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T18:03:16.076Z · Price source checked: 2026-09-27T18:03:16.076Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: xiaomi-mimo-api-v2-6-pro: pricing · xiaomi-mimo-api-v2-6-pro: availability · xiaomi-mimo-api-v2-6-flash: pricing · xiaomi-mimo-api-v2-6-flash: availability · xiaomi-mimo-api-v2-6-pro-ultraspeed: pricing · xiaomi-mimo-api-v2-6-pro-ultraspeed: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T18:03:16.076Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 64dbe58790cfa2331c3df7df8a6948f53ec9b55dc685d94bd055895a3447f483; collector xiaomi-mimo-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. API models are not linked to downloadable MiMo checkpoints.

      Xiaomi MiMo

      mimo-v2.6-pro-ultraspeed

      Input: USD 4.35 / 1000000 tokens · Tier unspecifiedOutput: USD 8.7 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 4.35 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Xiaomi MiMo API Open Platform overseas pay-as-you-go price (Real-time API), USD per 1M tokens. Domestic (China, CNY) prices, the Batch API and Token Plan subscriptions are out of scope; web search is billed separately per call.
      • Model ID mimo-v2.6-pro-ultraspeed.
      • Input tokens that miss the prompt cache ("Input (Cache Miss)").

      Cache read: USD 0.036 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Xiaomi MiMo API Open Platform overseas pay-as-you-go price (Real-time API), USD per 1M tokens. Domestic (China, CNY) prices, the Batch API and Token Plan subscriptions are out of scope; web search is billed separately per call.
      • Model ID mimo-v2.6-pro-ultraspeed.
      • Input tokens that hit the prompt cache ("Input (Cache Hit)"). The source lists cache writes as "Limited-time Free" (no end date); cache writes are not priced here.

      Output: USD 8.7 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Xiaomi MiMo API Open Platform overseas pay-as-you-go price (Real-time API), USD per 1M tokens. Domestic (China, CNY) prices, the Batch API and Token Plan subscriptions are out of scope; web search is billed separately per call.
      • Model ID mimo-v2.6-pro-ultraspeed.
      • Output tokens.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: Xiaomi MiMo

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T18:03:16.076Z · Price source checked: 2026-09-27T18:03:16.076Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: xiaomi-mimo-api-v2-6-pro: pricing · xiaomi-mimo-api-v2-6-pro: availability · xiaomi-mimo-api-v2-6-flash: pricing · xiaomi-mimo-api-v2-6-flash: availability · xiaomi-mimo-api-v2-6-pro-ultraspeed: pricing · xiaomi-mimo-api-v2-6-pro-ultraspeed: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T18:03:16.076Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 64dbe58790cfa2331c3df7df8a6948f53ec9b55dc685d94bd055895a3447f483; collector xiaomi-mimo-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. API models are not linked to downloadable MiMo checkpoints.

      Novita AI

      MiniMax M2

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.3 / 1000000 tokens · serverless · Modality: text

      Region: Unknown · Endpoint scope: unknown · Tier: serverless · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Novita hosted serverless listing, USD per 1M tokens; standard listed prices only. Batch, priority, dedicated, promotional and free-limited offers are excluded.
      • Native decimal values agree with integer ten-thousandths and original/listed amounts. No tiered billing is reported for this listing.
      • Account and regional eligibility, status-code semantics and effective dates are unverified; listing is not an inference service test.
      • Listed prompt/input token rate; separately listed cache-read rates apply only to cache hits.

      Output: USD 1.2 / 1000000 tokens · serverless · Modality: text

      Region: Unknown · Endpoint scope: unknown · Tier: serverless · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Novita hosted serverless listing, USD per 1M tokens; standard listed prices only. Batch, priority, dedicated, promotional and free-limited offers are excluded.
      • Native decimal values agree with integer ten-thousandths and original/listed amounts. No tiered billing is reported for this listing.
      • Account and regional eligibility, status-code semantics and effective dates are unverified; listing is not an inference service test.
      • Listed completion/output token rate.

      Cache read: USD 0.03 / 1000000 tokens · serverless · Modality: text

      Region: Unknown · Endpoint scope: unknown · Tier: serverless · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Novita hosted serverless listing, USD per 1M tokens; standard listed prices only. Batch, priority, dedicated, promotional and free-limited offers are excluded.
      • Native decimal values agree with integer ten-thousandths and original/listed amounts. No tiered billing is reported for this listing.
      • Account and regional eligibility, status-code semantics and effective dates are unverified; listing is not an inference service test.
      • Cache-read input only; cache-write pricing and cache TTL are unknown.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: Novita AI

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T06:32:08.821Z · Price source checked: 2026-09-27T06:32:08.821Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

      Source, conditions and dates

      Raw API evidence ↗ · vendor-api · Claims: novita-llama-3-3-70b-instruct: pricing · novita-minimax-m2: pricing · novita-glm-5: pricing · Retrieved · Effective: Unknown · Checked: 2026-09-27T06:32:08.821Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: B

      • Raw SHA-256: 9722d1130261fe9667b226c4817976e42d0087d11cf1eb7b948cecdeb7988f25; original collector novita-model-catalog-1. Pinned parse-failure recovery preserves original dates; no new acquisition.
      • USD/scale interpretation reviewed against Novita official pricing and its MiniMax M3 pricing explanation; exact original/decimal/integer values retained privately.
      • Three curated hosted listings only. No verified checkpoint/developer/license/weight-access link; availability remains unknown because numeric status semantics are undocumented.

      Checkpoint: MiniMax M2

      MiniMax

      MiniMax M2

      No API price listedSelf-hosting costs not included
      Weights unknownDownload not verifiedtextContext: Unknown tokens

      Weight license: Unknown / not verified

      Developer: MiniMax

      Checkpoint: MiniMax M2 · Revision: Unknown / not verified · Lineage: Unknown / not verified

      Creator model source ↗

      Model metadata observed: 2026-09-27T05:30:10.560Z · Metadata freshness: fresh

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Modified MIT (MiniMax M2) · API availability: unknown (dated observation, not a live service check)

      Active parameters: 10B; total weights: 230B.

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: minimax-m2: identity · minimax-m2: specifications · minimax-m2: download · minimax-m2: weightAccess · minimax-m2: gating · developer:minimax: identity · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:30:10.560Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

      • The repository links this exact MiniMaxAI checkpoint.128K appears in benchmark configuration, not as a verified maximum context.
      • Retain copyright/permission notices; the modified MIT terms require MiniMax M2 branding above100M monthly active users orUSD30M annual recurring revenue. Do not inherit M2.7 terms. Separate weight scope remains unknown.
      • Official checkpoint/download reference: https://huggingface.co/MiniMaxAI/MiniMax-M2. No model weights downloaded; access/download availability remains unverified.
      • Document revision kind git-blob; revision 18f4e0a6e67973b942f4c98965290f036691435a; SHA256 e464eb0a1b4addabf59cc0084b22fdf1ecf84ffb77a35effeec32a39ca14cff0; private manifest ai/v1/onboard-minimax-m2-readme-md/2026-09-27/complete.json.gz; manifest SHA256 e91583d80e45b9c331c790577b6385cd63772075afedb2961fd2d0fb69fd6f64.

      Primary source ↗ · vendor-docs · Claims: minimax-m2: codeLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T03:05:21.024Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

      • The repository links this exact MiniMaxAI checkpoint.128K appears in benchmark configuration, not as a verified maximum context.
      • Retain copyright/permission notices; the modified MIT terms require MiniMax M2 branding above100M monthly active users orUSD30M annual recurring revenue. Do not inherit M2.7 terms. Separate weight scope remains unknown.
      • Official checkpoint/download reference: https://huggingface.co/MiniMaxAI/MiniMax-M2. No model weights downloaded; access/download availability remains unverified.
      • Document revision kind git-blob; revision c0139070cdfe30607cbd5b506c6f294fa9ec8aca; SHA256 1a75fb68199afb1fcf3896b070b20ea05fe628945b946489d97d50b6a72eb10e; private manifest ai/v1/access-minimax-m2-license/2026-09-27/complete.json.gz; manifest SHA256 97ce69f4cf6323ba1682fa159aa06fc2c7aacc64e3da3243ce5e57368ee35a4f.

      MiniMax

      MiniMax-M2.7

      Input: USD 0.3 / 1000000 tokens · Tier unspecifiedOutput: USD 1.2 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.3 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Model ID MiniMax-M2.7.
      • Input tokens.

      Output: USD 1.2 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Model ID MiniMax-M2.7.
      • Output tokens.

      Cache read: USD 0.06 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Model ID MiniMax-M2.7.
      • Prompt caching read.

      Cache write: USD 0.375 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Model ID MiniMax-M2.7.
      • Prompt caching write (the pricing table states no cache TTL).

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: MiniMax

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T12:58:05.337Z · Price source checked: 2026-09-27T12:58:05.337Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: minimax-api-m3: pricing · minimax-api-m3: availability · minimax-api-m2-7: pricing · minimax-api-m2-7: availability · minimax-api-m2-7-highspeed: pricing · minimax-api-m2-7-highspeed: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T12:58:05.337Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 983e1a21a05880f1f92e2cfb258cdb486184bf77f574eb4769a628aa49be4835; collector minimax-international-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. API models are not linked to downloadable checkpoints (M2/M2.7 weight licensing is version-specific and separate from API billing).

      MiniMax

      MiniMax-M2.7-highspeed

      Input: USD 0.6 / 1000000 tokens · Tier unspecifiedOutput: USD 2.4 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.6 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Model ID MiniMax-M2.7-highspeed.
      • Input tokens.

      Output: USD 2.4 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Model ID MiniMax-M2.7-highspeed.
      • Output tokens.

      Cache read: USD 0.06 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Model ID MiniMax-M2.7-highspeed.
      • Prompt caching read.

      Cache write: USD 0.375 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Model ID MiniMax-M2.7-highspeed.
      • Prompt caching write (the pricing table states no cache TTL).

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: MiniMax

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T12:58:05.337Z · Price source checked: 2026-09-27T12:58:05.337Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: minimax-api-m3: pricing · minimax-api-m3: availability · minimax-api-m2-7: pricing · minimax-api-m2-7: availability · minimax-api-m2-7-highspeed: pricing · minimax-api-m2-7-highspeed: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T12:58:05.337Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 983e1a21a05880f1f92e2cfb258cdb486184bf77f574eb4769a628aa49be4835; collector minimax-international-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. API models are not linked to downloadable checkpoints (M2/M2.7 weight licensing is version-specific and separate from API billing).

      MiniMax

      MiniMax-M3

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.30 / 1000000 tokens · Standard · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Standard tab (the source lists Priority separately at 1.5x standard).
      • Model ID MiniMax-M3.
      • Context tier: ≤ 512k input tokens (source notation; the pricing table does not define k as an exact token count).
      • Input tokens.
      • Charged price shown with a "Permanent 50% off" label against a struck-through list price of $0.60.

      Output: USD 1.20 / 1000000 tokens · Standard · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Standard tab (the source lists Priority separately at 1.5x standard).
      • Model ID MiniMax-M3.
      • Context tier: ≤ 512k input tokens (source notation; the pricing table does not define k as an exact token count).
      • Output tokens.
      • Charged price shown with a "Permanent 50% off" label against a struck-through list price of $2.40.

      Cache read: USD 0.06 / 1000000 tokens · Standard · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Standard tab (the source lists Priority separately at 1.5x standard).
      • Model ID MiniMax-M3.
      • Context tier: ≤ 512k input tokens (source notation; the pricing table does not define k as an exact token count).
      • Prompt caching read.
      • Charged price shown with a "Permanent 50% off" label against a struck-through list price of $0.12.

      Input: USD 0.60 / 1000000 tokens · Standard · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Standard tab (the source lists Priority separately at 1.5x standard).
      • Model ID MiniMax-M3.
      • Context tier: > 512k input tokens (source notation; the pricing table does not define k as an exact token count).
      • The source marks this tier with an asterisk; the page's only footnote describes the Priority tier, so the asterisk's meaning for this tier is not stated.
      • Input tokens.
      • Charged price shown with a "Permanent 50% off" label against a struck-through list price of $1.20.

      Output: USD 2.40 / 1000000 tokens · Standard · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Standard tab (the source lists Priority separately at 1.5x standard).
      • Model ID MiniMax-M3.
      • Context tier: > 512k input tokens (source notation; the pricing table does not define k as an exact token count).
      • The source marks this tier with an asterisk; the page's only footnote describes the Priority tier, so the asterisk's meaning for this tier is not stated.
      • Output tokens.
      • Charged price shown with a "Permanent 50% off" label against a struck-through list price of $4.80.

      Cache read: USD 0.12 / 1000000 tokens · Standard · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • MiniMax international platform (platform.minimax.io) pay-as-you-go, USD per 1M tokens; the Priority tier, Credits/Token Plan and China-platform pricing are out of scope.
      • Standard tab (the source lists Priority separately at 1.5x standard).
      • Model ID MiniMax-M3.
      • Context tier: > 512k input tokens (source notation; the pricing table does not define k as an exact token count).
      • The source marks this tier with an asterisk; the page's only footnote describes the Priority tier, so the asterisk's meaning for this tier is not stated.
      • Prompt caching read.
      • Charged price shown with a "Permanent 50% off" label against a struck-through list price of $0.24.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: MiniMax

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T12:58:05.337Z · Price source checked: 2026-09-27T12:58:05.337Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: minimax-api-m3: pricing · minimax-api-m3: availability · minimax-api-m2-7: pricing · minimax-api-m2-7: availability · minimax-api-m2-7-highspeed: pricing · minimax-api-m2-7-highspeed: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T12:58:05.337Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 983e1a21a05880f1f92e2cfb258cdb486184bf77f574eb4769a628aa49be4835; collector minimax-international-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. API models are not linked to downloadable checkpoints (M2/M2.7 weight licensing is version-specific and separate from API billing).

      Checkpoint: Mistral Small 3.1 24B

      Mistral AI · Mistral Small

      Mistral Small 3.1 24B

      No API price listedSelf-hosting costs not included
      Open weightsDownloadabletext · imageContext: 131,072 tokens

      Weight license: Apache-2.0 · Hugging Face / download weights ↗

      24B parameters · Estimated weight memory: Q4 18.0 GB · Q8 32.4 GB. Includes 20% loading overhead; excludes KV cache and runtime memory. Compare hardware.

      Developer: Mistral AI

      Checkpoint: Mistral Small 3.1 24B · Revision: 68faf511d618ef198fef186659617cfd2eb8e33a · Lineage: Unknown / not verified

      Creator model source ↗ Creator model source ↗ Weight-license evidence ↗

      Model metadata observed: 2026-09-27T00:01:38.884Z / 2026-09-27T00:01:44.681Z · Metadata freshness: fresh

      Weight-license reviewed: 2026-09-27T02:50:26.221Z. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: ungated · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · huggingface-model-card · Claims: mistral-small-3-1-24b: identity · mistral-small-3-1-24b: specifications · mistral-small-3-1-24b: weightAccess · mistral-small-3-1-24b: download · mistral-small-3-1-24b: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Supports image understanding, not image generation; requires Mistral-specific serving and tokenizer settings.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: mistral-small-3-1-24b: identity · mistral-small-3-1-24b: specifications · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:38.884Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Card advertises 24B parameters, text/image capabilities and 128K context; exact max_position_embeddings is131,072, consistent with that advertised limit.
      • Base mistralai/Mistral-Small-3.1-24B-Base-2503 is outside the curated registry.
      • Metadata documents were read; weight binaries were not downloaded. Prior weight-access/download observations retain their original dates.
      • Revision 68faf511d618ef198fef186659617cfd2eb8e33a; SHA256 68effd5695d7053fa8153feedaaae529ce10bd6481f86dad76374a59de3592e6; private manifest ai/v1/.model-metadata-development/36281114084/hf-mistral-small-3-1-24b-readme-md/2026-09-27/complete.json.gz; manifest SHA256 40766d3d77b0f976572fe175194b19a96a780c4667fd1bd44c9896579e386f8e.

      Raw API evidence ↗ · vendor-api · Claims: mistral-small-3-1-24b: gating · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:32.656Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Repository gating: false; listed weight files: true. An inventory and authenticated metadata read do not verify weight download availability.
      • Revision 68faf511d618ef198fef186659617cfd2eb8e33a; SHA256 1bc16923fdccdeff3b0dee2d4ef009a51f5166df8f5ea2c85c1ad9f4b487f151; private manifest ai/v1/.model-metadata-development/36281114084/hf-mistral-small-3-1-24b-info/2026-09-27/complete.json.gz; manifest SHA256 c71f76c76660e5b2e643daa69b8223905d49b66695b14d7ee379f16e6b86b52c.

      Primary source ↗ · vendor-docs · Claims: mistral-small-3-1-24b: configuration · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:44.681Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Configuration is supporting evidence, not an automatic claim of advertised context or modality. Resolve differences using the explicit publication review.
      • Revision 68faf511d618ef198fef186659617cfd2eb8e33a; SHA256 ce3ec410cac74da358f786c574b73b6624c50c8bb876bcb628f06500fe07adcc; private manifest ai/v1/.model-metadata-development/36281114084/hf-mistral-small-3-1-24b-config-json/2026-09-27/complete.json.gz; manifest SHA256 d7d7dea0f654ce935423043b3d0f657a2cf9d8cf9fb8b6d6a8684081582c800a.

      Primary source ↗ · huggingface-model-card · Claims: mistral-small-3-1-24b: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:01:38.884Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Exact model card expressly declares Apache-2.0 for usage and modification, including commercial and non-commercial purposes. The captured inventory contains no separate LICENSE; this is a reviewed creator declaration, not a claim that a standalone license document or external library license was inspected.
      • Revision 68faf511d618ef198fef186659617cfd2eb8e33a; SHA256 68effd5695d7053fa8153feedaaae529ce10bd6481f86dad76374a59de3592e6; private manifest ai/v1/.model-metadata-development/36281114084/hf-mistral-small-3-1-24b-readme-md/2026-09-27/complete.json.gz; manifest SHA256 40766d3d77b0f976572fe175194b19a96a780c4667fd1bd44c9896579e386f8e.

      Checkpoint: Qwen3 0.6B

      Qwen · Qwen3

      Qwen3 0.6B

      No API price listedSelf-hosting costs not included
      Open weightsDownloadabletextContext: 32,768 tokens

      Weight license: Apache-2.0 · Hugging Face / download weights ↗

      0.6B parameters · Estimated weight memory: Q4 0.4 GB · Q8 0.8 GB. Includes 20% loading overhead; excludes KV cache and runtime memory. Compare hardware.

      Developer: Qwen

      Checkpoint: Qwen3 0.6B · Revision: c1899de289a04d12100db370d81485cdf75e47ca · Lineage: Unknown / not verified

      Creator model source ↗ Creator model source ↗ Weight-license evidence ↗ Weight-license evidence ↗

      Model metadata observed: 2026-09-27T00:00:12.090Z / 2026-09-27T00:00:17.971Z · Metadata freshness: fresh

      Weight-license reviewed: 2026-09-27T02:50:26.221Z. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: ungated · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-0-6b: identity · qwen-qwen3-0-6b: specifications · qwen-qwen3-0-6b: weightAccess · qwen-qwen3-0-6b: download · qwen-qwen3-0-6b: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Repository name uses 0.6B; Hugging Face safetensors metadata reports approximately 752M parameters.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-30b-a3b: identity · qwen-qwen3-30b-a3b: specifications · qwen-qwen3-30b-a3b: weightAccess · qwen-qwen3-30b-a3b: download · qwen-qwen3-30b-a3b: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Mixture-of-experts model with 3.3B parameters activated per token; weight memory reflects 30.5B total parameters. Validated to 131,072 tokens with YaRN.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-0-6b: identity · qwen-qwen3-0-6b: specifications · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:00:12.090Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Creator card advertises 0.6B model parameters; repository tensor metadata reports 751,632,384 elements. The discrepancy is unresolved; the published model-size field follows the creator card, not a measured checkpoint byte size.
      • Native context explicitly 32,768 tokens. Base checkpoint Qwen/Qwen3-0.6B-Base is outside the curated registry; no identity link is inferred.
      • Metadata documents were read; weight binaries were not downloaded. Prior weight-access/download observations retain their original dates.
      • Revision c1899de289a04d12100db370d81485cdf75e47ca; SHA256 1ab64a26fcb3b461423b89a433a8c858f1bf8d4086f979cbb3ff878d47cf20e9; private manifest ai/v1/.model-metadata-development/36281114084/hf-qwen-qwen3-0-6b-readme-md/2026-09-27/complete.json.gz; manifest SHA256 412197809e9b332de9a0b09729295eff5ed2bd93041e8759ba340cdedc006d1d.

      Raw API evidence ↗ · vendor-api · Claims: qwen-qwen3-0-6b: gating · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:00:06.238Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Repository gating: false; listed weight files: true. An inventory and authenticated metadata read do not verify weight download availability.
      • Revision c1899de289a04d12100db370d81485cdf75e47ca; SHA256 35689400e35733468a959e53e5572c56f25331c0b7008a0460ee03d457195673; private manifest ai/v1/.model-metadata-development/36281114084/hf-qwen-qwen3-0-6b-info/2026-09-27/complete.json.gz; manifest SHA256 74f2cc2a57eb992ff5ad20ecdd3a487ce3ee83b2a6ce039e51b7b5047b45c986.

      Primary source ↗ · vendor-docs · Claims: qwen-qwen3-0-6b: configuration · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:00:17.971Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Configuration is supporting evidence, not an automatic claim of advertised context or modality. Resolve differences using the explicit publication review.
      • Revision c1899de289a04d12100db370d81485cdf75e47ca; SHA256 660db3b73d788119c04535e48cf9be5f55bc3100841a718637ae695b442f27dd; private manifest ai/v1/.model-metadata-development/36281114084/hf-qwen-qwen3-0-6b-config-json/2026-09-27/complete.json.gz; manifest SHA256 c0fa1c44da1d7337731b39a4c5b676e11b5e6f650f0ee3999a8ecbc3c3ae19af.

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-0-6b: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:00:12.090Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Exact model card declares Apache-2.0 and the repository license supplies those terms. Redistribution requires license/notice preservation and marking modifications; patent and trademark terms apply. This review establishes the declared model license, not an independent license for external inference libraries.
      • Revision c1899de289a04d12100db370d81485cdf75e47ca; SHA256 1ab64a26fcb3b461423b89a433a8c858f1bf8d4086f979cbb3ff878d47cf20e9; private manifest ai/v1/.model-metadata-development/36281114084/hf-qwen-qwen3-0-6b-readme-md/2026-09-27/complete.json.gz; manifest SHA256 412197809e9b332de9a0b09729295eff5ed2bd93041e8759ba340cdedc006d1d.

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-0-6b: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:00:23.648Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Exact model card declares Apache-2.0 and the repository license supplies those terms. Redistribution requires license/notice preservation and marking modifications; patent and trademark terms apply. This review establishes the declared model license, not an independent license for external inference libraries.
      • Revision c1899de289a04d12100db370d81485cdf75e47ca; SHA256 832dd9e00a68dd83b3c3fb9f5588dad7dcf337a0db50f7d9483f310cd292e92e; private manifest ai/v1/.model-metadata-development/36281114084/hf-qwen-qwen3-0-6b-license/2026-09-27/complete.json.gz; manifest SHA256 5e4d98d2d08fe5d97a5839920545be8e2dfc56fca094c45753e3c7fdbf3d7ab2.

      Checkpoint: Qwen3 30B-A3B

      Qwen · Qwen3

      Qwen3 30B-A3B

      No API price listedSelf-hosting costs not included
      Open weightsDownloadabletextContext: 32,768 tokens

      Weight license: Apache-2.0 · Hugging Face / download weights ↗

      30.5B parameters · Estimated weight memory: Q4 22.9 GB · Q8 41.2 GB. Includes 20% loading overhead; excludes KV cache and runtime memory. Compare hardware.

      Developer: Qwen

      Checkpoint: Qwen3 30B-A3B · Revision: ad44e777bcd18fa416d9da3bd8f70d33ebb85d39 · Lineage: Unknown / not verified

      Creator model source ↗ Creator model source ↗ Weight-license evidence ↗ Weight-license evidence ↗

      Model metadata observed: 2026-09-27T00:00:36.268Z / 2026-09-27T00:00:42.205Z · Metadata freshness: fresh

      Weight-license reviewed: 2026-09-27T02:50:26.221Z. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: ungated · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

      Active parameters: 3.3B; total weights: 30.5B.

      Source, conditions and dates

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-0-6b: identity · qwen-qwen3-0-6b: specifications · qwen-qwen3-0-6b: weightAccess · qwen-qwen3-0-6b: download · qwen-qwen3-0-6b: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Repository name uses 0.6B; Hugging Face safetensors metadata reports approximately 752M parameters.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-30b-a3b: identity · qwen-qwen3-30b-a3b: specifications · qwen-qwen3-30b-a3b: weightAccess · qwen-qwen3-30b-a3b: download · qwen-qwen3-30b-a3b: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Mixture-of-experts model with 3.3B parameters activated per token; weight memory reflects 30.5B total parameters. Validated to 131,072 tokens with YaRN.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-30b-a3b: identity · qwen-qwen3-30b-a3b: specifications · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:00:36.268Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Creator card specifies 30.5B total and 3.3B active parameters. Memory estimates use total, never active, parameters.
      • Native context is 32,768; 131,072 requires the documented YaRN configuration. Configured positions 40,960 are not a replacement for native context. Base checkpoint Qwen/Qwen3-30B-A3B-Base is outside the registry.
      • Metadata documents were read; weight binaries were not downloaded. Prior weight-access/download observations retain their original dates.
      • Revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39; SHA256 1897b7cdf7b5f45ae847b87f9b416d4635fe73228e0daac669a4e1f3e5948858; private manifest ai/v1/.model-metadata-development/36281114084/hf-qwen-qwen3-30b-a3b-readme-md/2026-09-27/complete.json.gz; manifest SHA256 7a58e7afe6aa4cae98f81dbc5ed230b2f17167093ac7e1ecbd76e3f441c00d26.

      Raw API evidence ↗ · vendor-api · Claims: qwen-qwen3-30b-a3b: gating · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:00:30.582Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Repository gating: false; listed weight files: true. An inventory and authenticated metadata read do not verify weight download availability.
      • Revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39; SHA256 417eaddef81df7e652bf5037b74b8dd9ebdcac2a2bbc660a2078f826c7d02c9b; private manifest ai/v1/.model-metadata-development/36281114084/hf-qwen-qwen3-30b-a3b-info/2026-09-27/complete.json.gz; manifest SHA256 66a466201ee86891df7504209c5a680d70ff8d76b4fc4f3e2615337efcb832f6.

      Primary source ↗ · vendor-docs · Claims: qwen-qwen3-30b-a3b: configuration · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:00:42.205Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Configuration is supporting evidence, not an automatic claim of advertised context or modality. Resolve differences using the explicit publication review.
      • Revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39; SHA256 2850ddb3bf7aecad20b611e2d44f3077fc8193f4827c93beddd4c02ad63c2297; private manifest ai/v1/.model-metadata-development/36281114084/hf-qwen-qwen3-30b-a3b-config-json/2026-09-27/complete.json.gz; manifest SHA256 e6d3fc13038ccf80dac8fb9fb19eacfbd99831cb14bf68089fbbb3883279819c.

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-30b-a3b: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:00:36.268Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Exact model card declares Apache-2.0 and the repository license supplies those terms. Redistribution requires license/notice preservation and marking modifications; patent and trademark terms apply. This review establishes the declared model license, not an independent license for external inference libraries.
      • Revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39; SHA256 1897b7cdf7b5f45ae847b87f9b416d4635fe73228e0daac669a4e1f3e5948858; private manifest ai/v1/.model-metadata-development/36281114084/hf-qwen-qwen3-30b-a3b-readme-md/2026-09-27/complete.json.gz; manifest SHA256 7a58e7afe6aa4cae98f81dbc5ed230b2f17167093ac7e1ecbd76e3f441c00d26.

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-30b-a3b: weightLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T00:00:47.901Z · Reviewed: 2026-09-27T02:50:26.221Z · Review after: 2026-10-27T02:50:26.221Z · Freshness: fresh · Confidence: A

      • Exact model card declares Apache-2.0 and the repository license supplies those terms. Redistribution requires license/notice preservation and marking modifications; patent and trademark terms apply. This review establishes the declared model license, not an independent license for external inference libraries.
      • Revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39; SHA256 832dd9e00a68dd83b3c3fb9f5588dad7dcf337a0db50f7d9483f310cd292e92e; private manifest ai/v1/.model-metadata-development/36281114084/hf-qwen-qwen3-30b-a3b-license/2026-09-27/complete.json.gz; manifest SHA256 a1b595a7ed862edbf5a057fe9ea790019986dd38f5f9da28afd5304cac980b01.

      Checkpoint: Qwen3.6-35B-A3B (lineage reference)

      Qwen

      Qwen3.6-35B-A3B (lineage reference)

      No API price listedSelf-hosting costs not included
      Weights unknownDownload not verifiedContext: Unknown tokens

      Weight license: Unknown / not verified

      Developer: Qwen

      Checkpoint: Qwen3.6-35B-A3B (lineage reference) · Revision: Unknown / not verified · Lineage: Unknown / not verified

      Creator model source ↗

      Model metadata observed: 2026-09-27T04:40:13.000Z · Metadata freshness: fresh

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: unknown (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-0-6b: identity · qwen-qwen3-0-6b: specifications · qwen-qwen3-0-6b: weightAccess · qwen-qwen3-0-6b: download · qwen-qwen3-0-6b: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Repository name uses 0.6B; Hugging Face safetensors metadata reports approximately 752M parameters.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: qwen-qwen3-30b-a3b: identity · qwen-qwen3-30b-a3b: specifications · qwen-qwen3-30b-a3b: weightAccess · qwen-qwen3-30b-a3b: download · qwen-qwen3-30b-a3b: weightLicense · Retrieved · Effective: Unknown · Checked: Unknown · Reviewed: Unknown · Review after: Unknown · Freshness: stale · Confidence: Unknown

      • Mixture-of-experts model with 3.3B parameters activated per token; weight memory reflects 30.5B total parameters. Validated to 131,072 tokens with YaRN.

      Latest result: failed. Retained evidence keeps its original observation and successful-check dates.

      Primary source ↗ · huggingface-model-card · Claims: kat-coder-v2-5-dev: identity · kat-coder-v2-5-dev: specifications · kat-coder-v2-5-dev: download · kat-coder-v2-5-dev: weightAccess · kat-coder-v2-5-dev: gating · developer:kwaipilot: identity · lineage:qwen3-6-35b-a3b: identity · kat-coder-v2-5-dev: lineage:post-trained-from:lineage:qwen3-6-35b-a3b · Retrieved · Effective: Unknown · Checked: 2026-09-27T04:40:13.000Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

      • Exact card explicitly ships language-only weights without vision components. Native context is262144; optional YaRN scaling is separate. No KAT-Pro hosted-offer join is inferred.
      • The card explicitly describes post-training from Qwen3.6-35B-A3B. That parent is represented only as an external lineage reference; its other specifications and access are not inferred. Apache-2.0 YAML does not establish separate legal scopes; preserve the #92 unknown decisions. Kwaipilot/KAT branding is observed; legal-entity aliases remain unverified.
      • Official checkpoint/download reference: https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev. No model weights downloaded; access/download availability remains unverified.
      • Document revision kind repository-commit; revision 7be56fe773e72b6f5ca93c1ae45d828ddb893922; SHA256 e9b49bdc7b5f00a5a1f985143eb43148304a0a20cf0f3e4e935bd8d369897561; private manifest ai/v1/.access-evidence-development/36294950609/access-kat-coder-v2-5-dev-readme-md/2026-09-27/complete.json.gz; manifest SHA256 be03c87e7e3e7e8810d87a57543a477919c2681a5ae4c0f4faec30b57bb79d69.

      Alibaba Cloud Model Studio

      Qwen3.7-Plus

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.4 / 1000000 tokens · Standard · Modality: unspecified

      Region: singapore · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–256000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Alibaba Cloud Model Studio International deployment (Singapore), USD standard real-time price per 1M tokens; China (Beijing) and other regions, batch, context-cache and free-quota terms are out of scope.
      • Model ID qwen3.7-plus.
      • Input tokens per request: 0 < tokens ≤ 256K (K = 1,000; M = 1,000,000). All tokens in a request are billed at the unit price of its tier.
      • List price; the source states a limited-time 20% discount without an end date, so the discounted amount is not published.

      Output: USD 1.6 / 1000000 tokens · Standard · Modality: unspecified

      Region: singapore · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–256000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Alibaba Cloud Model Studio International deployment (Singapore), USD standard real-time price per 1M tokens; China (Beijing) and other regions, batch, context-cache and free-quota terms are out of scope.
      • Model ID qwen3.7-plus.
      • Input tokens per request: 0 < tokens ≤ 256K (K = 1,000; M = 1,000,000). All tokens in a request are billed at the unit price of its tier.
      • Same price in Non-Thinking and Thinking modes.
      • List price; the source states a limited-time 20% discount without an end date, so the discounted amount is not published.

      Input: USD 1.2 / 1000000 tokens · Standard · Modality: unspecified

      Region: singapore · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): 256001–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Alibaba Cloud Model Studio International deployment (Singapore), USD standard real-time price per 1M tokens; China (Beijing) and other regions, batch, context-cache and free-quota terms are out of scope.
      • Model ID qwen3.7-plus.
      • Input tokens per request: 256K < tokens ≤ 1M (K = 1,000; M = 1,000,000). All tokens in a request are billed at the unit price of its tier.
      • List price; the source states a limited-time 20% discount without an end date, so the discounted amount is not published.

      Output: USD 4.8 / 1000000 tokens · Standard · Modality: unspecified

      Region: singapore · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): 256001–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Alibaba Cloud Model Studio International deployment (Singapore), USD standard real-time price per 1M tokens; China (Beijing) and other regions, batch, context-cache and free-quota terms are out of scope.
      • Model ID qwen3.7-plus.
      • Input tokens per request: 256K < tokens ≤ 1M (K = 1,000; M = 1,000,000). All tokens in a request are billed at the unit price of its tier.
      • Same price in Non-Thinking and Thinking modes.
      • List price; the source states a limited-time 20% discount without an end date, so the discounted amount is not published.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: Alibaba Cloud Model Studio

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T09:01:54.333Z · Price source checked: 2026-09-27T09:01:54.333Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: qwen-api-qwen3-8-max: pricing · qwen-api-qwen3-8-max: availability · qwen-api-qwen3-7-plus: pricing · qwen-api-qwen3-7-plus: availability · qwen-api-qwen3-8-flash: pricing · qwen-api-qwen3-8-flash: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T09:01:54.333Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 5b08a91d20324792ff6df505f35700e0d862e2074c3b605e34d8319fc977a9ea; collector qwen-international-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. Alibaba Cloud is the API seller; the Qwen API models are not linked to downloadable checkpoints.

      Alibaba Cloud Model Studio

      Qwen3.8-Flash

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.15 / 1000000 tokens · Standard · Modality: unspecified

      Region: singapore · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Alibaba Cloud Model Studio International deployment (Singapore), USD standard real-time price per 1M tokens; China (Beijing) and other regions, batch, context-cache and free-quota terms are out of scope.
      • Model ID qwen3.8-flash.
      • Input tokens per request: 0 < tokens ≤ 1M (K = 1,000; M = 1,000,000). All tokens in a request are billed at the unit price of its tier.

      Output: USD 0.47 / 1000000 tokens · Standard · Modality: unspecified

      Region: singapore · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Alibaba Cloud Model Studio International deployment (Singapore), USD standard real-time price per 1M tokens; China (Beijing) and other regions, batch, context-cache and free-quota terms are out of scope.
      • Model ID qwen3.8-flash.
      • Input tokens per request: 0 < tokens ≤ 1M (K = 1,000; M = 1,000,000). All tokens in a request are billed at the unit price of its tier.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: Alibaba Cloud Model Studio

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T09:01:54.333Z · Price source checked: 2026-09-27T09:01:54.333Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: qwen-api-qwen3-8-max: pricing · qwen-api-qwen3-8-max: availability · qwen-api-qwen3-7-plus: pricing · qwen-api-qwen3-7-plus: availability · qwen-api-qwen3-8-flash: pricing · qwen-api-qwen3-8-flash: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T09:01:54.333Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 5b08a91d20324792ff6df505f35700e0d862e2074c3b605e34d8319fc977a9ea; collector qwen-international-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. Alibaba Cloud is the API seller; the Qwen API models are not linked to downloadable checkpoints.

      Alibaba Cloud Model Studio

      Qwen3.8-Max

      Input: See detailsOutput: See detailsHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 2 / 1000000 tokens · Standard · Modality: unspecified

      Region: singapore · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Alibaba Cloud Model Studio International deployment (Singapore), USD standard real-time price per 1M tokens; China (Beijing) and other regions, batch, context-cache and free-quota terms are out of scope.
      • Model ID qwen3.8-max.
      • Mode: Non-Thinking and Thinking modes.
      • Input tokens per request: 0 < tokens ≤ 1M (K = 1,000; M = 1,000,000). All tokens in a request are billed at the unit price of its tier.

      Output: USD 6 / 1000000 tokens · Standard · Modality: unspecified

      Region: singapore · Endpoint scope: international · Tier: standard · Billing: Unknown · Prompt bracket (inclusive): Unknown–1000000 tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • Alibaba Cloud Model Studio International deployment (Singapore), USD standard real-time price per 1M tokens; China (Beijing) and other regions, batch, context-cache and free-quota terms are out of scope.
      • Model ID qwen3.8-max.
      • Mode: Non-Thinking and Thinking modes.
      • Input tokens per request: 0 < tokens ≤ 1M (K = 1,000; M = 1,000,000). All tokens in a request are billed at the unit price of its tier.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: Alibaba Cloud Model Studio

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T09:01:54.333Z · Price source checked: 2026-09-27T09:01:54.333Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: qwen-api-qwen3-8-max: pricing · qwen-api-qwen3-8-max: availability · qwen-api-qwen3-7-plus: pricing · qwen-api-qwen3-7-plus: availability · qwen-api-qwen3-8-flash: pricing · qwen-api-qwen3-8-flash: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T09:01:54.333Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: 5b08a91d20324792ff6df505f35700e0d862e2074c3b605e34d8319fc977a9ea; collector qwen-international-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. Alibaba Cloud is the API seller; the Qwen API models are not linked to downloadable checkpoints.

      Checkpoint: Step 3.5 Flash

      StepFun

      Step 3.5 Flash

      No API price listedSelf-hosting costs not included
      Weights unknownDownload not verifiedtextContext: Unknown tokens

      Weight license: Unknown / not verified

      Developer: StepFun

      Checkpoint: Step 3.5 Flash · Revision: Unknown / not verified · Lineage: Unknown / not verified

      Creator model source ↗

      Model metadata observed: 2026-09-27T05:30:00.221Z · Metadata freshness: fresh

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Apache-2.0 · API availability: unknown (dated observation, not a live service check)

      Active parameters: 11B; total weights: 196.81B.

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: step-3-5-flash: identity · step-3-5-flash: specifications · step-3-5-flash: download · step-3-5-flash: weightAccess · step-3-5-flash: gating · developer:stepfun: identity · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:30:00.221Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

      • Detailed architecture table includes 196B backbone plus0.81B head; active count is approximate. Advertised256K context is not converted into an unsupported exact token count.
      • Creator links first-party StepFun API documentation and distinct OpenRouter hosting; no API prices or hosted equivalence are published. Project Apache license establishes repository code terms; separate weight scope remains unknown.
      • Official checkpoint/download reference: https://huggingface.co/stepfun-ai/Step-3.5-Flash. No model weights downloaded; access/download availability remains unverified.
      • Document revision kind git-blob; revision 0d9ab7147b3f5cb80fd6f73a9127afec6b743976; SHA256 8ad3990e1f7b2f9fbbf151e6da8f100adc3c97a8200ac5d400442ca32c3746dd; private manifest ai/v1/onboard-step-3-5-flash-readme-md/2026-09-27/complete.json.gz; manifest SHA256 151c5375f6abaed7949bcd4753305936ecb18304668f5ad26bad3a72813c411c.

      Primary source ↗ · vendor-docs · Claims: step-3-5-flash: codeLicense · Retrieved · Effective: Unknown · Checked: 2026-09-27T05:30:05.324Z · Reviewed: 2026-09-27T05:40:06.050Z · Review after: 2026-10-27T05:40:06.050Z · Freshness: fresh · Confidence: A

      • Detailed architecture table includes 196B backbone plus0.81B head; active count is approximate. Advertised256K context is not converted into an unsupported exact token count.
      • Creator links first-party StepFun API documentation and distinct OpenRouter hosting; no API prices or hosted equivalence are published. Project Apache license establishes repository code terms; separate weight scope remains unknown.
      • Official checkpoint/download reference: https://huggingface.co/stepfun-ai/Step-3.5-Flash. No model weights downloaded; access/download availability remains unverified.
      • Document revision kind git-blob; revision 261eeb9e9f8b2b4b0d119366dda99c6fd7d35c64; SHA256 c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4; private manifest ai/v1/onboard-step-3-5-flash-license/2026-09-27/complete.json.gz; manifest SHA256 0b2753884062762906a72b619a15c4a862655b9ab70a11e625a9bfb6db3fcff4.

      StepFun

      Step-3.5-Flash

      Input: USD 0.10 / 1000000 tokens · Tier unspecifiedOutput: USD 0.30 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.10 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • StepFun international platform (platform.stepfun.ai) direct API, USD per 1M tokens; China-platform pricing is out of scope.
      • Model ID step-3.5-flash.
      • Input price on cache miss.

      Cache read: USD 0.02 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • StepFun international platform (platform.stepfun.ai) direct API, USD per 1M tokens; China-platform pricing is out of scope.
      • Model ID step-3.5-flash.
      • Input that hits the prompt cache.

      Output: USD 0.30 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • StepFun international platform (platform.stepfun.ai) direct API, USD per 1M tokens; China-platform pricing is out of scope.
      • Model ID step-3.5-flash.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: StepFun

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T11:56:18.231Z · Price source checked: 2026-09-27T11:56:18.231Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: stepfun-api-step-5-preview: pricing · stepfun-api-step-5-preview: availability · stepfun-api-step-3-7-flash: pricing · stepfun-api-step-3-7-flash: availability · stepfun-api-step-3-5-flash: pricing · stepfun-api-step-3-5-flash: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T11:56:18.231Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: e53bfd9dc98d997e58aeda7d9030de843588063fda5bbe4569f39c9844c3449a; collector stepfun-international-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. The API models are not linked to downloadable checkpoints.

      StepFun

      Step-3.7-Flash

      Input: USD 0.20 / 1000000 tokens · Tier unspecifiedOutput: USD 1.15 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 0.20 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • StepFun international platform (platform.stepfun.ai) direct API, USD per 1M tokens; China-platform pricing is out of scope.
      • Model ID step-3.7-flash.
      • Input price on cache miss.

      Cache read: USD 0.04 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • StepFun international platform (platform.stepfun.ai) direct API, USD per 1M tokens; China-platform pricing is out of scope.
      • Model ID step-3.7-flash.
      • Input that hits the prompt cache.

      Output: USD 1.15 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • StepFun international platform (platform.stepfun.ai) direct API, USD per 1M tokens; China-platform pricing is out of scope.
      • Model ID step-3.7-flash.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: StepFun

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T11:56:18.231Z · Price source checked: 2026-09-27T11:56:18.231Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: stepfun-api-step-5-preview: pricing · stepfun-api-step-5-preview: availability · stepfun-api-step-3-7-flash: pricing · stepfun-api-step-3-7-flash: availability · stepfun-api-step-3-5-flash: pricing · stepfun-api-step-3-5-flash: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T11:56:18.231Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: e53bfd9dc98d997e58aeda7d9030de843588063fda5bbe4569f39c9844c3449a; collector stepfun-international-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. The API models are not linked to downloadable checkpoints.

      StepFun

      Step-5 Preview

      Input: USD 1.00 / 1000000 tokens · Tier unspecifiedOutput: USD 2.70 / 1000000 tokens · Tier unspecifiedHeadline rates · native source units · conditions below
      Weights unknownAPI listedDownload not verifiedContext: Unknown tokens

      Input: USD 1.00 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • StepFun international platform (platform.stepfun.ai) direct API, USD per 1M tokens; China-platform pricing is out of scope.
      • Model ID step-5-preview.
      • Input price on cache miss; for step-5-preview this includes writing new content to the cache.

      Cache read: USD 0.05 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • StepFun international platform (platform.stepfun.ai) direct API, USD per 1M tokens; China-platform pricing is out of scope.
      • Model ID step-5-preview.
      • Input that hits the prompt cache.

      Output: USD 2.70 / 1000000 tokens · Modality: unspecified

      Region: international · Endpoint scope: international · Tier: Unknown · Billing: Unknown · Prompt bracket (inclusive): Unknown–Unknown tokens · Cache TTL: Unknown seconds · Valid through: Unknown

      • StepFun international platform (platform.stepfun.ai) direct API, USD per 1M tokens; China-platform pricing is out of scope.
      • Model ID step-5-preview.
      • Output tokens include both the model's reasoning process and final answer.

      Weight license: Unknown / not verified

      Developer: Unknown / not verified · API seller: StepFun

      Checkpoint: Unknown / not verified · Lineage: Unknown / not verified

      Official pricing ↗

      Price observed: 2026-09-27T11:56:18.231Z · Price source checked: 2026-09-27T11:56:18.231Z · Price freshness: unknown

      Weight-license reviewed: Unknown. Open weights do not establish permission for every commercial use; consult the linked terms.

      Gating: unknown · Code license: Unknown / not verified · API availability: available (dated observation, not a live service check)

      Source, conditions and dates

      Primary source ↗ · vendor-docs · Claims: stepfun-api-step-5-preview: pricing · stepfun-api-step-5-preview: availability · stepfun-api-step-3-7-flash: pricing · stepfun-api-step-3-7-flash: availability · stepfun-api-step-3-5-flash: pricing · stepfun-api-step-3-5-flash: availability · Retrieved · Effective: Unknown · Checked: 2026-09-27T11:56:18.231Z · Reviewed: Unknown · Review after: Unknown · Freshness: unknown · Confidence: A

      • Raw SHA-256: e53bfd9dc98d997e58aeda7d9030de843588063fda5bbe4569f39c9844c3449a; collector stepfun-international-1.
      • Pricing and listing observation only; no checkpoint, license, context-window or release-date assertions. The API models are not linked to downloadable checkpoints.

      Read before comparing

      The same unit is not the same offer.

      API prices retain their original currency, decimal amount and billing unit. No AI currency conversions are performed. Prompt brackets, input modality, cache reads and writes, geography and service tiers remain separate; an unknown dimension is not a standard-tier assumption. Output charges can include reasoning or thinking tokens. Keep each record's dated source conditions in view; a listed rate is not a promise of current availability.

      Unknown cached-input prices are not zero. Download-only records have no fabricated API price. Open weights do not automatically mean an OSI-approved open-source license; review the license and any gated download terms.

      Q4 and Q8 estimates use 5 and 9 bits per weight respectively, plus 20% loading overhead in decimal GB. These are weight-capacity estimates, not performance benchmarks. KV cache, activations, runtime scratch space, drivers and backend compatibility require additional planning. See the methodology and source register.

      Estimated local deployment

      How model size maps to accelerator memory

      Estimated weights plus 20% loading overhead. Runtime, KV cache, context length, and software can require substantially more memory.

      Estimated 4-bit weight memoryGB, including 20% overhead · logarithmic bar scale

      Method

      Weights are only the starting point.

      The estimate uses parameters × planning bits ÷ 8 × 1.20: 5 planning bits for Q4 and 9 for Q8 to allow for scales, metadata, and mixed-precision tensors. “Comfortable” reserves at least 20% of device memory. “Tight” means estimated weights fit but leave less headroom.

      This is a capacity screen—not a speed benchmark or guarantee that a particular runtime, context length, multimodal projector, or GPU configuration will work.

      Estimated 4-bit model weight fit by available accelerator memory
      ModelParametersQ4 estimateQ8 estimate8 GB12 GB16 GB24 GB32 GB48 GB80 GB141 GB192 GB256 GB512 GB
      DeepSeek R1671B503.3 GB905.9 GBNoNoNoNoNoNoNoNoNoNoTight
      DeepSeek V3671B503.3 GB905.9 GBNoNoNoNoNoNoNoNoNoNoTight
      Gemma 3 1B IT1B0.8 GB1.3 GBComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortable
      Gemma 3 4B IT4B3.0 GB5.4 GBComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortable
      Llama 3.2 11B Vision10.6B7.9 GB14.3 GBTightComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortable
      Llama 3.2 1B Instruct1.23B0.9 GB1.7 GBComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortable
      Llama 3.3 70B Instruct70B52.5 GB94.5 GBNoNoNoNoNoNoComfortableComfortableComfortableComfortableComfortable
      Mistral Small 3.1 24B24B18.0 GB32.4 GBNoNoNoComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortable
      Qwen3 0.6B0.6B0.4 GB0.8 GBComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortableComfortable
      Qwen3 30B-A3B30.5B22.9 GB41.2 GBNoNoNoTightComfortableComfortableComfortableComfortableComfortableComfortableComfortable