- GPT-6 Astra arrives with a price that can look expensive if it is read only as a token tariff: $10 per million input tokens and $50 per million output tokens, with cached input priced lower.
- The more useful unit is cost per finished task.
- OpenAI’s launch materials make that argument with task-level evidence.
- Section
- AI
- Read time
- 7 min read
- Why this page exists
- The Grid Report publishes operator-grade coverage on AI, power, infrastructure, automation, and markets.

A better AI cost scorecard
Token price is an input. The accepted unit of work is the output that determines whether automation creates leverage.
| Metric | What to measure | Why it matters |
|---|---|---|
| Direct model cost | Input, cached input, output, and tool-call spend | This is the visible bill, but not the whole cost of the workflow. |
| Completion quality | Share of tasks accepted without material rework | A cheap attempt is expensive if it cannot reach the acceptance bar. |
| Human overhead | Review minutes, corrections, and escalation time | Labor often outweighs small differences in token spend. |
| Control risk | Unauthorized actions, tool failures, and severity of defects | Computer use can expand both the value and the downside of one run. |
The Grid Report operator framework, informed by OpenAI’s September 2026 model, pricing, evaluation, and safety materials.
GPT-6 Astra arrives with a price that can look expensive if it is read only as a token tariff: $10 per million input tokens and $50 per million output tokens, with cached input priced lower. That comparison is useful for budgeting, but it is no longer sufficient for deciding which model should run demanding professional work.
The more useful unit is cost per finished task. If a cheaper model needs more prompting, produces a longer answer, loses context during a multistep workflow, or leaves enough errors that a person must redo the work, its low token price can become an expensive operating choice. A higher-priced model can be cheaper when it completes more of the assignment correctly and with less supervision.
The model with the lowest token price is not necessarily the model with the lowest cost per accepted deliverable.
OpenAI’s launch materials make that argument with task-level evidence. On its BenchCAD configuration, the company reports that Astra reached a 95.9% geometric-overlap score while its estimated API cost was about 43% lower than GPT-5.6 Sol and 86% lower than Claude Fable 5.1. Those are vendor-reported results, not a universal guarantee, but the comparison is directionally important: better capability can lower total task cost even when each token carries a premium.
Astra also expands the scope of what belongs inside one task. The model supports a 1.05-million-token context window and is positioned for computer use, browsing, software engineering, research, and document creation. For operators, that means a workflow can include finding evidence, working across applications, producing an artifact, and responding to changed requirements instead of stopping at a draft answer.
That broader scope changes procurement. Teams should stop asking only how many tokens a model consumes and start measuring completion rate, elapsed time, human review minutes, correction cycles, tool failures, and the cost of an incorrect action. The denominator should be an accepted deliverable: a verified code fix, a reconciled report, a completed research brief, or a presentation that actually follows the company template.
Control belongs in that equation. Astra’s safety materials say it is less likely than GPT-5.6 Sol to take destructive or unauthorized actions in realistic computer-use settings, while also noting that the model has reached OpenAI’s Critical threshold for cybersecurity capability. That combination makes permissions, monitoring, isolation, and approval gates operating requirements rather than optional governance language.
A practical evaluation should therefore use real work and a fixed acceptance rubric. Give competing models the same source material, tools, permissions, and deadline. Record API spend, tool-call cost, completion rate, reviewer time, and the severity of defects. Run enough cases to expose variance. A dazzling demo is not a unit-economics test; a repeatable batch of accepted tasks is.
The original angle for operators is simple: Astra should not be bought because it tops a benchmark or avoided because its output tokens are expensive. It should earn a place in the workflow only where its judgment, context, and computer use reduce the total cost and risk of reaching an approved outcome.
For investors and infrastructure planners, the same shift matters at scale. If stronger models make previously uneconomic work viable, demand can grow even while the compute cost of an individual task falls. Efficiency does not automatically reduce infrastructure demand; it can expand the number of tasks worth automating.
Sources
OpenAI, GPT-6 Astra launch and task-cost comparisons: https://openai.com/index/gpt-6-astra/
OpenAI API model page and pricing: https://developers.openai.com/api/docs/models/gpt-6-astra
OpenAI, GPT-6 Astra safety overview: https://openai.com/index/safety-overview-gpt-6-astra/
Nawaz Lalani
Nawaz Lalani is the creator of The Grid Report and writes about AI infrastructure, grid power demand, automation systems, and the market signals shaping the physical AI economy. His focus is translating technical and industrial shifts into practical coverage for operators, investors, builders, and teams making real deployment decisions.
B.S. in Geology from UT Arlington. Covers AI infrastructure, energy systems, grid constraints, automation workflows, and market signals.
Stories are built from primary sources, utility and infrastructure signals, company disclosures, filings, and operator-grade context. The goal is to explain what changed, why it matters now, and what it means for builders, investors, utilities, and teams making real deployment decisions.
Follow the lane, not just the headline.
The strongest value in The Grid Report comes from following how AI, infrastructure, power, automation, and markets connect over time.