apisesdks
CONFIRMED: OpenRouter launches Batch API, cutting prices in half across 70+ models
CONFIRMED: OpenRouter launched its Batch API on September 22, with a 50% discount on input and output token prices across more than 70 models, for tasks that do not need an immediate response and are delivered within 24 hours.

[{"h":"CONFIRMED: batch processing now costs half price on OpenRouter","p":"OpenRouter, the platform that aggregates access to dozens of AI models through a single API and is now controlled by Stripe, officially launched its Batch API on September 22. The feature cuts the per token price, both input and output, in half across more than 70 supported models for tasks that do not require a real time response."},{"h":"How the discount works","p":"The logic is simple: instead of paying full price for an immediate response, developers submit a batch of requests processed asynchronously, with a guarantee that all results arrive within 24 hours. In exchange for that wait, the per token price drops by half for both the text sent and the text the model generates."},{"h":"The numbers from the test period","p":"During a two week test period before the official launch, the platform processed more than 230,000 batch tasks, with a median completion time of seven minutes and 90% of tasks finished within an hour, well under the promised 24 hour ceiling. That suggests that in practice most tasks come back fast, and the one day window functions more as a floor guarantee than a typical wait time."},{"h":"What kind of task this is for","p":"Batch processing is the right choice for work that can wait minutes or a few hours: classifying a large batch of support messages, generating product descriptions in bulk, summarizing a large volume of documents, or running sentiment analysis over conversation history. It is not suited for a chatbot answering a customer in real time, where the user expects an immediate response."},{"h":"Why this matters for tight budgets","p":"For a small or midsize company using AI for behind the scenes work, such as organizing spreadsheets, generating content in bulk, or processing forms, a 50% cut in per token cost changes the math directly: the same monthly budget now covers twice the volume, as long as the task can tolerate not being instant. It is a practical reminder that optimizing AI cost is often less about switching models and more about choosing the right processing mode for each task."},{"h":"MaxAssistant's read","p":"The launch reinforces a trend already showing up at other AI providers: pricing by urgency, not just by model. Worth mapping out, inside your own operation, which AI tasks genuinely need an instant response and which could run in batch overnight for half the cost."},{"h":"Sources","p":"OpenRouter, OpenRouter Batch API, half-price inference by bundling requests: https://openrouter.ai/blog/announcements/batch-api/ | PANews, AI model aggregation platform OpenRouter launches Batch API, token prices halved for most models: https://panews.io/articles/01a0cc16-5bb5-74e7-8803-5263c7a4fc27 | Phemex News, OpenRouter Batch API Launch, 50% Off Tokens for 70+ AI Models: https://phemex.com/news/article/openrouter-launches-batch-api-with-50-token-price-reduction-across-70-models-97535"}]