WeaveScope

Model pricing

Calculate model cost from tokens, service tiers, prompt size, and custom pricing rules.

WeaveScope calculates model cost when an observation includes token usage but does not provide an explicit cost. Pricing rules apply across the organization.

Model cost is observability data. It is separate from your WeaveScope subscription and ingestion charges.

Where cost appears

You can see model cost in:

  • The trace table

  • Trace headers

  • Model observation details

  • Monitoring cost charts

Use these views to compare spend by trace, model, user, release, tag, or other searchable context.

Managed pricing

WeaveScope includes managed pricing for:

  • OpenAI

  • Gemini

  • Anthropic

  • Kimi/Moonshot

  • xAI

  • Z.ai

  • DeepSeek

  • TypeSafe / Jev

Managed rules are maintained by WeaveScope and cannot be edited. Open Settings → Model pricing to see the models, effective prices, service tiers, prompt-size bands, cache behavior, and provider notes currently available.

Because provider catalogs and rates change frequently, use that page rather than copying prices into application configuration.

OpenAI GPT-6.1 Sol

The managed catalog includes gpt-6.1-sol, matching BeamWeaver's openai:gpt-6.1-sol profile. Standard prices per million tokens are:

Input context Input Cached input Cache write Output
Up to 272K input tokens $2.00 $0.10 $2.50 $10.00
Above 272K input tokens $4.00 $0.20 $5.00 $15.00

The context band applies to the entire request. Batch and Flex use half these rates, while Fast and the legacy priority tier use twice these rates. Reported US or EU regional processing adds 10%. Fast is unavailable with EU data residency. See the model specification and OpenAI pricing .

Managed pricing sync adds this model to existing organizations during the normal release migration process.

TypeSafe / Jev

The managed catalog includes jev-1.13.0, jev-latest, and jev-preview, with the TypeSafe icon and provider label. The aliases currently resolve to Jev 1.13. The verified rate is $0.042 per million input tokens; output tokens are free. See TypeSafe's model documentation for the current rates. Unknown future versions do not inherit another version's price.

BeamWeaver decision traces can supply their calculated cost as metadata.cost.total_cost or outputs.metadata.cost.total_cost. WeaveScope recognizes these values as explicit costs, including zero. If no valid explicit cost is present, the matching managed pricing rule applies. Small nonzero costs retain enough decimal places in trace tables and details to remain visible.

The existing release migration process syncs managed pricing for existing organizations. Local installations can run WeaveScope.Organizations.ensure_default_model_pricing_rules_for_all_organizations!/0 after updating the application. Updating rules does not reprice already ingested observations. Ingestion replay preserves normalization snapshots, including their original costs; historical repricing requires a separate audited backfill.

Kimi K3 cache writes use the provider's 5-minute and 1-hour rates ($3 and $6 per million written tokens). BeamWeaver must export the cache-write token count and the response-header TTL split for WeaveScope to distinguish them. See Kimi's cache billing guide .

DeepSeek V4 Pro retains its own peak and off-peak rates after September 14, 2026. The provider reversed its planned reroute to Flash; the managed pricing sync archives the obsolete post-September-14 Flash-rate rule. See DeepSeek's current pricing .

Provider-hosted tools

When a generation reports billable server-tool usage, WeaveScope adds these charges to the model's token cost:

Provider Usage Added cost
OpenAI Web search; non-reasoning web_search_preview; file search $0.01/search; $0.025/non-reasoning preview search; $0.0025/file-search call
OpenAI Declared image-generation tool with a known image model and reported tool tokens Image-model text input, image input, and image output token rates
Anthropic Web search $0.01/search
Kimi/Moonshot Legacy $web_search tool call $0.005/call
Z.ai Web Search in Chat with returned results $0.01/use
xAI Web search; file search; code interpreter $0.005/search; $0.0025/file-search call; $0.005/code call
xAI X Search since September 21, 2026, 19:00 UTC $0.005/post fetched; $0.01/user profile fetched

Before that X Search change, WeaveScope uses $0.005/X Search call. The trace's Hosted usage details show the added charge and the reported xAI item counts. Kimi's standalone search/fetch REST endpoints are separate requests and need their own observations. For Z.ai, a nonempty web_search result identifies an observed use; the provider response does not expose a separate search count. Search content tokens remain part of model input token pricing. An explicit event cost remains authoritative and is not increased by these estimated tool charges. xAI responses provide usage.cost_in_usd_ticks, an provider-billed per-request amount including token and hosted-tool charges; WeaveScope uses it before falling back to rate estimates and stores the USD amount at its eight-decimal observation precision. These estimates require the provider's usage counters.

OpenAI image-generation tool cost is separate from the mainline model's token cost. WeaveScope uses the tool model declared in the request and the response's tooling.hosted.usage.image_gen token details. If the model or token modality is unavailable, billing.image_generation_unpriced marks the missing estimate. OpenAI does not return cached image-tool input counts in the response, so the uncached input rate is an upper-bound estimate; the billing metadata marks this limit. Explicit provider cost remains authoritative.

Gemini Search and Maps grounding share account-level free allowances. Cache storage, Anthropic code-execution and Managed Agents runtime, and OpenAI container and file-search storage are also billed outside an individual model response. WeaveScope records the available per-call billing evidence and provides account-level pricing functions for usage grouped by the provider's billing account. Kimi's standalone search/fetch endpoints still need their own observations.

Direct OpenAI Images API calls made with BeamWeaver.OpenAI.generate_image/2 emit a generation observation with the requested image model and the API's image/text token details. WeaveScope prices those separately from Responses API image tools and does not store returned image bytes in the trace. Multimodal audio, image, and video charges cannot always be reconstructed from generic text-token counts. Direct embedding calls must emit their own usage observations to appear in trace totals.

Sources: OpenAI pricing , Anthropic web search , Kimi web-search pricing , Z.ai pricing , xAI pricing , xAI cost tracking , and Gemini pricing .

Custom pricing

Add a custom rule when your model is not in the managed catalog.

Field Meaning
Provider Provider label shown in WeaveScope.
Name Display name for the rule.
Model ID Exact identifier emitted in trace metadata.
Input price USD per 1 million input tokens.
Cached input price USD per 1 million cached input tokens.
Output price USD per 1 million output tokens.

The model ID must exactly match the value shown in trace details. A custom rule created from one project can price matching observations in every project in the organization.

How cost is selected

For each observation, WeaveScope:

  1. Uses an explicit event cost, such as cost_usd, cost, or total_cost, when present.

  2. Otherwise finds a matching organization pricing rule.

  3. Applies token, cache, service-tier, prompt-size, and geography information supported by that rule.

  4. Uses 0 if it has neither an explicit cost nor a matching rule.

Token formula

For a matching rule:

uncached_input_tokens =
  input_tokens - cached_input_tokens - cache_creation_tokens

cost =
  uncached_input_tokens * input_price_per_token +
  cached_input_tokens * cached_input_price_per_token +
  cache_creation_tokens * cache_creation_price_per_token +
  output_tokens * output_price_per_token

WeaveScope derives per-token values from the configured prices per million tokens.

If a cached-input or cache-creation price is unavailable, the normal input price applies to those tokens.

Tiered pricing

Managed rules can distinguish service tiers such as:

  • standard

  • batch

  • fast

  • flex

  • priority

When an observation includes service-tier metadata, WeaveScope uses the matching tier when available and otherwise falls back to standard. Cost is based on the tier the provider reports as served, not only the tier requested.

Some managed rules also select prices by input-token count or inference geography. When supported, WeaveScope applies the matching prompt-size band or geography multiplier from the observation metadata.

Why cost can be zero

Cost remains zero when:

  • The event has no explicit cost.

  • Token usage is missing.

  • Provider or model metadata does not match a pricing rule.

  • A custom rule uses a different model ID from the trace.

  • The model is outside the managed catalog and has no custom rule.

Open the model observation in Tracing → Details and copy the exact provider and model values into the custom rule.

Pricing changes and historical traces

WeaveScope records calculated cost when it ingests an observation. Updating a pricing rule affects future observations only.

Existing traces keep their original cost unless an operator explicitly replays or backfills them.

DeepSeek V4.1 Flash uses the API ID deepseek-flash. Its rates took effect at 04:00 UTC on September 10, 2026: per million tokens, off-peak input/cache-hit/output cost $0.15/$0.003/$0.60 and peak cost $0.30/$0.006/$1.20. Peak windows are weekdays 01:00–04:00 and 06:00–10:00 UTC; weekends are off-peak.

The legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names use these rates from the same cutoff. deepseek-v4-pro changes to Flash rates at 04:00 UTC on September 14, 2026. Managed rules select the correct version from the observation timestamp, including older rates for delayed historical ingestion. These dates and rates follow the DeepSeek release announcement and official pricing table .