OpenAI's GPT-6 Astra hits OpenRouter at $10/$50 per million tokens
OpenAI’s new flagship model, GPT-6 Astra, went live on OpenRouter shortly after its September 4, 2026 release. OpenAI positions it for heavy end-to-end work—advanced analysis, software engineering, deep research, and document creation—with an emphasis on long-horizon agentic tasks that drive a computer or browser. It ships with a roughly 1M-token context window (listed at 1,050,000) and up to 128,000 completion tokens, accepts PDFs, images, and text as input, and returns text. Function calling, tool_choice, and JSON-schema structured outputs are all supported.
Pricing is steep at the high end: $10 per million input tokens and $50 per million output, with cache reads at $1/M, cache writes at $12.50/M, and web search billed at $10 per 1,000 calls—so aggressive prompt caching is the obvious lever for controlling cost. On OpenRouter the model is served by two providers, OpenAI and Azure (US), with automatic failover between them and routing modes (Balanced, Nitro, Exacto) that trade off price, speed, and tool-calling accuracy. Early performance figures show about 65 tokens/sec throughput, 3.36s median latency, and 100% three-day uptime, with routing lifting availability from 95.3% on a single provider to roughly 98.4%.
The more telling signal is where the traffic is going. The top consumers of the model are agentic tools rather than chat apps—Nous Research’s open-source Hermes Agent leads at 2.2B tokens, with Anthropic’s Claude Code among the heaviest users at 811M tokens. That usage pattern reinforces OpenAI’s framing of Astra as an agent engine, and underscores how quickly autonomous coding and browsing workloads are becoming the primary demand driver for frontier models.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.