Inference provider · Google

Google Vertex AI

Google Cloud’s model platform. An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

Prices reported by OpenRouter’s public API, snapshot · third-party-reported, not measured by Agent

Where its price sits

11 of 13models where its price is the cheapest, ties included

Models with two or more standard-tier providers. Ties included in “cheapest”; a calculation on reported prices, 3 input : 1 output tokens. Not a score.

At a glance

14

models in the index (of 27)

Listed under Google Vertex

1x

median price vs the cheapest provider of the same model

Calculation · n = 13 models with 2 or more providers

100%

median uptime over the last day, as reported

n = 12 endpoints; reported, not measured

Not reported

latency and throughput

The keyless API returned latency for 0 of 265 endpoints. Gateway delay is not measured: no key in the environment.

Its price vs every other provider

Its cheapest standard-tier endpoint per model, against every other provider of that model. Switch the price type; select another provider to pin it instead. Blended price, position and “vs cheapest” are calculations on the reported prices. A cheaper endpoint can run lower precision or a shorter context.

Reported
Calculation: (3 × input + output) ÷ 4

Pinned Google Vertex: on 13 of 14 rows

Model
  1. Claude Haiku 4.54 providers · all $2.00
  2. Claude Sonnet 55 providers · all $4.00
  3. Claude Sonnet 5.55 providers · all $4.00
  4. Claude Opus 4.85 providers · all $10.00
  5. Claude Opus 55 providers · all $10.00
  6. Claude Opus 5.55 providers · all $8.00
  7. Claude Fable 5.14 providers · all $20.00
  8. gpt-oss-120b20 providers · CoreWeave $0.065 → Cerebras $0.45
  9. Gemini 3.8 Flash2 providers · all $1.50
  10. Gemini 3.5 Flash2 providers · all $3.38
  11. Gemini 3.5 Flash Lite2 providers · all $0.85
  12. Gemini 3.1 Pro Preview2 providers · all $4.50
  13. Llama 4 Maverick3 providers · DigitalOcean $0.3037 → Parasail $0.5125
  14. Llama 3.3 70B Instruct10 providers · DeepInfra $0.155 → Together $1.04
  • one provider (blended: a calculation, hollow)
  • first-party list price

Third-party reported values. 74 listings, 4 series: Blended 3:1, Input, Output, Cache read. Blended 3:1: highest Claude Fable 5.1 · Amazon Bedrock $20.00 (n 4). Lowest gpt-oss-120b · CoreWeave (fp4) $0.065 (n 20). Input: highest Claude Fable 5.1 · Amazon Bedrock $10.00 (n 4). Lowest gpt-oss-120b · DekaLLM (bf16) $0.03 (n 20).

Notesn 2–20 per row

One row per model it serves; Google Vertex AI is pinned in color, every other provider is gray. USD per million tokens, snapshot 2026-10-06.

Prices reported by OpenRouter’s public API (third-party-reported, not measured by Agent). One dot per provider: its cheapest standard-tier endpoint. The diamond is the first-party list price. Spread: highest ÷ lowest price of the row. A parenthesis names the quantization the provider reported. n = providers with a standard-tier endpoint.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Price per provider: Claude Haiku 4.5

Reported by OpenRouter’s public API, snapshot October 6, 2026. Third-party-reported prices, not measured by Agent.

Reported

4 providers with a standard-tier endpoint · all 4 providers report the same price ($2.00 blended) · showing 4 of 4 rows

  1. Inputall 4 at $1.00

  2. Outputall 4 at $5.00

  3. Cache readall 4 at $0.10

  • one provider (reported price, USD per million tokens, log scale per strip)
  • first-party list price

Blended price and “vs cheapest” are calculations on the reported prices (3 input : 1 output tokens), against the cheapest standard-tier provider. A lower price can come with lower precision (fp4, fp8) or a shorter context. Uptime is what the API reported for the last day; the API reported no latency or throughput.

Every value from the study

Each row is one of its prices on its own track, with the other providers of the same chart as muted dots (37 values). Reported prices, not measured. Use Table for the plain values; each row links to the chart it copies.

$1.00
$5.00
$0.10
$2.00
$10.00
$0.20

Reported by OpenRouter’s public API, not measured by AgentMuted dots: the other providers on the same chartUSD per million tokens

Google Vertex AI in Inference provider index: 27 models, 52 providers: 37 values, first Claude Haiku 4.5: price per million tokens by provider (Input) $1.00.

Compare Google Vertex AI

Price rows state the gap and name no winner: a reported price has no interval.

Next

Write-ups that use this data

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.