← Back to all sparks
O

OpenRouter

AI-ASSISTANTS
Velocity10.0

Unified API marketplace for 200+ LLMs from OpenAI, Anthropic, Google, Mistral, Meta and more.

OpenRouter launches Batch API for half-price async inference while building out its decision model catalog.

batch-inferencecost-optimizationdecision-modelsllm-routingasync-inference
◆Current state
OpenRouter is an LLM routing layer expanding in two directions: a new Batch API offering half-price inference for workloads that can tolerate 24-hour turnaround, and a growing catalog of decision models (Jev) that return typed probabilities instead of prose. The Batch API emerged from a two-week beta with 230k+ completed batches at a median of 7 minutes. The platform's content cadence has shifted toward developer tutorials and model comparisons, reflecting a growing education investment alongside catalog additions.
◆Where it's heading
OpenRouter is moving beyond pure model routing toward an inference optimization layer. The Batch API is the clearest signal: a pricing tradeoff that makes cost-sensitive bulk workloads viable on the platform for the first time. The volume of Jev-related content — five entries in a week — suggests a formal push to make typed decision models a first-class primitive alongside generative ones. The platform is positioning as the place to run all inference, synchronous or async, generative or structured.
◆Prediction
Given the Batch API beta scale and the sustained Jev content push, the next likely move is either an SDK or dashboard feature that surfaces per-workload cost-vs-latency tradeoffs and routes automatically between sync and batch — or a more formal tiering of the decision model category in the model browser.

◆Recent moves

  1. 2d ago

    Is Kimi K3 Open Source? Weights, License, and How to Call It

  2. 3d ago

    Best Embedding Models in 2026

  3. 3d ago

    How to Use Jev: Moderation with the Jev API in TypeScript

    End-to-end tutorial for using Jev on marketplace listing moderation, covering prompt structure, probability thresholds, and publish/hold/reject routing logic. Instructional content; no new platform capability.

  4. 4d ago

    Is Jev as Accurate as Frontier Models at Classification?

    Benchmark putting Jev 1.13 against Claude Opus 5 on Banking77 classification: Jev scores 81.0% vs 84.4% at 175ms and $0.11/thousand vs 2.3s and $2.42. Useful comparative data but a blog piece, not a platform change.

  5. 4d ago

    What Is Nemotron 3.5 Lightning

    Explainer on Nemotron 3.5 Lightning — NVIDIA's 30B MoE model with ~3B active parameters per token — covering architecture, endpoint features, and structured output usage via OpenRouter. Catalog documentation, not a platform change.

  6. 4d ago

    Batch API launches: half-price inference for async workloads

    ⚡ SPARK

    The Batch API ships out of beta: send a full workload in one POST, pay roughly half the per-token price, collect results within 24 hours. 230k+ batches completed during the beta with a median of 7 minutes — far shorter than the 24-hour SLA. This is the first time OpenRouter has offered a fundamentally different pricing tier based on latency tolerance, making bulk inference economics competitive with self-hosted options.