Portfolio 12 · modality ledger

Multimodal AI API cost estimator

Apply provider-specific media preprocessing, accept official pasted counts where required, and keep every text, image, audio, video, document, cache, and output charge separate.

Decision this answers: Which provider bill is lower for the same counted workload after its native resizing and counting rules—not which model has equivalent quality?

Scenario

Numeric scenario inputs

Official or pasted model-specific count.

measured[1]

Default example resizes 4032×3024 to 1568×1176 before counting.

measured[2]

Keep time- or token-based native counting explicit.

measured[1]

Use the provider's sampled/resized count, not original pixels.

measured[1]

Output and thinking rates remain separate from input.

measured[1]

Gemini 3.5 Flash-Lite standard list snapshot.

published list[1]

Output meter stays distinct.

published list[1]

The complete default scenario is server rendered. Interactions stay in this browser.

Decision summary

Uncached input

$14.55exact-derived

Text, image, audio, video, and document lines.

[1]

Cached input

$0.30exact-derived

Cached meter kept separate.

[1]

Output

$12.50exact-derived

Output and thinking-token meter.

[1]

Mixed-media bill

$27.35exact-derived

All visible modality line items reconciled.

[1][2][3]
Direct-labeled result profile

The tables below are the accessible source of truth; bar lengths never carry identity alone.

Provider counting comparison

Provider counting comparison
CandidateOriginal mediaBillable mediaCounting authorityUnknown policyComparison claimEvidence
GeminiVisibleProvider-counted tokensCountTokens/pricingPaste countBill onlymeasured[1]
OpenAIVisibleResized/tiled dimensionsVision guide/API countDo not priceBill onlymeasured[2]
AnthropicVisibleDocumented image countVision/token countingDo not priceBill onlymeasured[3]

Official surface versus this site

Gemini API pricing

Use the official surface to confirm native prices, counting rules, limits, and product eligibility.

This offline workbench

Official pages define model and modality meters. This page reconciles a complete mixed-media batch in one local ledger while preserving provider preprocessing and unknown counting rules.

Sources and methodology

Every result-affecting reference is visible here without JavaScript and is retained in the JSON export.

[1] Gemini API pricing
Authority
Google
Native identifier
gemini-3.5-flash-lite standard
Region
global
Unit
USD/1M tokens
Evidence
published list
Price kind
list
As of
2026-08-01
Effective from
2026-08-01
Retrieved
2026-08-01
Freshness
fresh
Normalization
The dated $0.30 input, $0.03 cached input, and $2.50 output meters are applied to separate counted modalities.

Open the authoritative source

[2] OpenAI vision guide
Authority
OpenAI
Native identifier
vision input counting
Unit
resized pixels/tiles
Evidence
measured
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
Original and billable dimensions remain visible; unverified local counts are labeled estimates.

Open the authoritative source

[3] Anthropic vision guide
Authority
Anthropic
Native identifier
vision token counting
Unit
image tokens
Evidence
measured
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
Provider-specific media preprocessing is kept distinct and unknown rules cannot yield zero cost.

Open the authoritative source

Limits and non-claims

  • Cross-provider rows compare bills, not capability, quality, latency, or safety.
  • Model-dependent counts should come from the provider counting API and can be pasted locally.
  • Unsupported modalities remain visible and unpriced rather than inheriting a zero meter.