‘Expensive AF’: The Week Users Did the Maths on GPT-6 Astra
OpenAI shipped a new flagship. Within days the conversation wasn't about how clever it is — it was about the bill, and whether the smartest model is the one you can actually afford to run.
OpenAI shipped a new flagship this week. GPT-6 Astra arrived on 3 September, at roughly two and a half times the token price of the model it replaces, and the striking thing about the reaction was how quickly it stopped being about intelligence. Within a day or two the people who actually pay for these tools weren’t asking whether Astra was cleverer. They were doing sums. The question that ran through the week wasn’t “is it good” but the older, harder one: is the smartest model the one you can actually afford to run?
This is a Voices piece, so the verdict belongs to the users, not to us. Our job is to open the threads, quote the wording exactly, be fair to the case for the defence, and resist turning a reasonable grumble about price into a scandal. There isn’t a scandal here. There is something more useful: a real-time audit, by paying customers, of whether the frontier is worth the fare.
Quotes sourced from: Hacker News. Every quote below was opened on the live thread, lifted verbatim from the comment, checked against Hacker News’s own record, and listed in the Sources with its handle, platform and date. We quoted only what we could open and read in full — Reddit and X threads we couldn’t reach are not quoted here — we aimed at the products and the pricing decisions rather than the people, and, because fairness is the job, we kept in the commenters who made the case for paying up.
The first feature anyone noticed was the bill
The launch-day verdict came not from a benchmark but from a wallet. forrestthewoods described the most modern of disappointments: “Threw $10 at this to help me prepare for my league’s fantasy auction this weekend. It spend $3.50 and then said ‘this action would cause you to go above your spending limit’.” The fix was to throw more money at it — a $100 Codex Max subscription that bundled Astra — and the conclusion was the week’s unofficial headline: “Sure seems like Astra is expensive AF.”
He was not alone in hitting a wall fast. xfax, trying the model on a small project, reported that it “used up all my limits for the day and had to continue the following day.” And the cost doesn’t just cap what you can do; it changes how you work. copperx, weighing a $24 charge for a single generated site, put his finger on the subtler harm: the price “doesn’t leave much room for error or experimentation.” When each attempt has a meaningful price tag, the cheap, iterative, try-it-and-see workflow that made these tools fun in the first place quietly stops being cheap. This is the meter we described when a flat AI subscription becomes a meter — only now the meter is attached to the flagship everyone is being nudged towards.
Does 2.5x the price buy 2.5x the model?
The sharpest thread of the week was the value argument, and users came armed with numbers. simianwords did the arithmetic out loud, using a third-party intelligence index: “GPT-6 Astra (low): 57 Intelligence Index, $7.70/M tokens… GPT-5.6 Sol (high): 57 Intelligence Index, $3.08/M tokens,” concluding that for the same measured score “Sol costs only 40% as much… while Astra is ~2.5× more expensive.” Whatever you make of any single benchmark, the framing is deadly: if last year’s model hits the same number for a fraction of the price, the premium has to be justified somewhere the index can’t see.
ellessarr named exactly where that justification is supposed to live, and why nobody checks it: “2× price only wins for review if it catches bugs the cheap model drops — nobody runs that test, everyone quotes the benchmark.” It is the whole problem with paying for the top tier in one line. The premium is sold on edge cases caught and disasters averted, but almost no one measures that; they look at the leaderboard, which is precisely the thing we’ve argued means less than you think.
To be fair — and the threads were fair — there is a real case for the expensive model, and gentlewater made it well. Astra and its peers, he argued, are “not the every day workhorse you reach for to do basic tasks… they’re the tool you break out when you need the absolute strongest performance,” the kind where “finding and fixing one or two extra edge cases saves the business a lot of money, even if the cost is high.” That is the honest steel-man: for the hardest problems, the premium can pay for itself many times over. The complaint underneath most of the grumbling isn’t that this is false; it’s that most work isn’t the hardest problem, and the pricing makes reaching for the flagship out of habit an expensive reflex.
The workaround economy: mandated, rationed, dreaded
What happens when a token-hungry model meets a corporate budget is that humans start behaving strangely around it. wookmaster reported the top-down version: “My company literally mandated tokenmaxxing while all the engineers told them this was a bad idea” — a policy of pushing everything through the priciest model, imposed over the objections of the people who’d have to explain the invoice. MisterMunchkin described the inevitable sequel, from the other side of the same ledger: “$10/$50 is incredibly expensive compared to Chinese models which are cents… My company is already massively cutting down on access because they’ve realised most people don’t actually produce any value using it. All the tokenmaxers have ruined it for the rest of us now that accounting have seen the costs.” Enthusiasm meets the finance team, and the finance team wins.
And hanging over all of it is a learned suspicion about where the price goes next. 1saadcodes voiced the fear that these threads return to again and again: even if the economics work today, the worry is “this will end up coming back to bite us, by becoming more expensive once they inevitably nerf it. Every major model provider does that now after all.” It is opinion, not prophecy, and we frame it as such — but it is opinion earned by pattern. We watched the specifics of it when DeepSeek turned its famously cheap tokens into peak-hour pricing: undercut, win the users, then discover what serving them costs and raise the price. Users have learned to read a generous launch as the first act, not the deal.
None of this is new to anyone who has watched a subscription quietly turn into a meter, but the launch of a flagship is when the tension is at its rawest: the marketing promises a leap in capability, and the invoice answers with a leap in cost. We heard the same note only days ago, when users described a day of hitting walls and watching their credits evaporate. What Astra’s arrival sharpened was the question underneath: not just “why did I run out” but “why am I paying a premium to run out faster.” When the best model is also the one you have to use most sparingly, the word “best” starts to wobble — because a tool you ration is a tool you use less, whatever the leaderboard says about it.
It wasn’t only Astra
The week’s discontent wasn’t confined to one launch. christophilus supplied the quality counterpoint to all the AGI talk, reporting that Astra “generated some of the worst Odin code I’ve ever seen” — a reminder that a top benchmark score and a good afternoon’s work are not the same thing. Over on the Claude side, larodi described a visceral fatigue with the output of a tool he uses daily: “with 4 agents doing my stuff on a daily basis, I feel like vomiting at some point, not mere nausea, but disgust. damn Codex seems to fare better at this imho.” His deeper frustration was that the style is untameable: he’s “tried many times to instruct it to not produce this nonsense, but… always finds a way around it.”
That tell-tale house style is its own small grievance. ahepp catalogued the symptoms every heavy user now recognises: “Anything I have claude or codex write carries a ton of distinctive characteristics. Obsession with ‘bit-for-bit identical’, ‘it’s not the X it’s the Y Z’ and so on,” adding that “it’s driving me nuts, I constantly have to prompt it to ‘explain in plain, simple English’.” And the feeling that useful capacity keeps being fenced off surfaced too: lucas_t_a noted that “opus 5 came by default with 200k context and 1M gated behind usage tokens, auto-compaction by default, and it keeps telling you to clear and start from scratch all the time.” More capable on the spec sheet; more conditional in the hand.
What the week actually said
Read together, these aren’t the complaints of people who think AI is bad. They’re the complaints of people who use it all day and have started pricing it like adults. The models are getting stronger; the experience of paying for them is getting more calculated — metered, rationed, and shadowed by the expectation of a rise. A few things worth carrying out of the week:
- Meter the flagship before you marry it. A new top model’s first felt feature is how fast it spends your allowance. Watch cost-per-task in the first hour, while the novelty is still paying the bill.
- Match the model to the job, not the hype. The steel-man for the expensive tier is real — for the hardest tasks. Reaching for it by default is how you end up as the cautionary tale accounting cites.
- Treat the launch price as act one. The recurring fear — cheap now, dear once you’re dependent — is earned by the industry’s own record. Keep a cheaper fallback you actually know how to use.
- A benchmark is not a value proof. “It tops the index” and “it’s worth 2.5× for your work” are different claims. Only one of them shows up on your invoice.
- “More capable” and “more conditional” are arriving together. Bigger context behind a paywall, output you can’t restyle, limits you hit by lunch — the power is real, and so are the strings.
None of this is a case against the new model. The same people totting up the cost are, almost to a person, still paying, still building, still back tomorrow. That is the tension the whole series keeps landing on: the tools are good enough to be worth the money, which is exactly why it’s worth watching the money so closely. The frontier moved again this week. So did the price of standing on it.
Frequently asked questions
What is GPT-6 Astra and why are users complaining about its price?
GPT-6 Astra is OpenAI's new flagship model, which began rolling out on 3 September 2026 at roughly 2.5 times the token price of the GPT-5.6 model it succeeds. On Hacker News, paying users' first reaction was sticker shock: reports of burning through spending caps and five-hour usage limits fast, and a running argument about whether a markedly more expensive model returns markedly more value for everyday work.
Is Astra actually worth 2.5x the price?
That is exactly what users were debating, and the sceptics were loud. One did the arithmetic and noted that for the same score on a third-party 'intelligence index', the older Sol model costs roughly 40% as much. The steel-man, also voiced in the threads, is that a frontier model is not the everyday workhorse but the tool you reach for on the hardest tasks, where catching one extra bug can be worth the premium. The complaint is less 'it's bad' than 'most work doesn't need it, and the price makes casual use expensive.'
What is 'tokenmaxxing'?
It is the informal term users use for pushing as many tokens as possible through the strongest and most expensive model, rather than routing routine work to cheaper ones. In these threads it came up as a grievance: one engineer said their company 'mandated tokenmaxxing' against the engineers' advice, and another blamed 'tokenmaxers' for driving up costs to the point that their employer cut everyone's access once the finance team saw the total.
Were the complaints only about GPT-6 Astra?
No. Alongside the Astra cost debate, users aired cross-tool frustrations: strong distaste for Claude Code's output and an inability to instruct the style away, a broader irritation at the tell-tale 'it's not X, it's Y' voice that models like Claude and Codex produce, and complaints that Anthropic's Opus 5 keeps its largest context window gated behind extra usage. The common thread was tools getting more capable and more conditional at the same time.
How were these quotes verified?
Every quote was lifted verbatim from the live Hacker News comment, checked against Hacker News's own record, and listed in the Sources with the commenter's handle, the platform, the date and a direct permalink. We quoted only comments we could open and read in full; Reddit and X threads we could not reach are not quoted here. We aimed at the products and pricing decisions rather than the people, and kept in the commenters who defended the models.
Sources
- forrestthewoods on Hacker News — “Threw $10 at this… It spend $3.50 and then said ‘this action would cause you to go above your spending limit’… Sure seems like Astra is expensive AF.” (5 Sep 2026) — Hacker News
- MisterMunchkin on Hacker News — “$10/$50 is incredibly expensive compared to Chinese models which are cents… My company is already massively cutting down on access… All the tokenmaxers have ruined it for the rest of us now that accounting have seen the costs.” (5 Sep 2026) — Hacker News
- simianwords on Hacker News — “GPT-6 Astra (low): 57 Intelligence Index, $7.70/M tokens… GPT-5.6 Sol (high): 57 Intelligence Index, $3.08/M tokens… Sol costs only 40% as much… while Astra is ~2.5× more expensive.” (5 Sep 2026) — Hacker News
- ellessarr on Hacker News — “2× price only wins for review if it catches bugs the cheap model drops — nobody runs that test, everyone quotes the benchmark.” (5 Sep 2026) — Hacker News
- 1saadcodes on Hacker News — “I'm still very worried that this will end up coming back to bite us, by becoming more expensive once they inevitably nerf it. Every major model provider does that now after all.” (5 Sep 2026) — Hacker News
- wookmaster on Hacker News — “My company literally mandated tokenmaxxing while all the engineers told them this was a bad idea.” (5 Sep 2026) — Hacker News
- copperx on Hacker News — “I would say that $24 is trivial IF that's the final design. The truth is that the cost doesn't leave much room for error or experimentation.” (5 Sep 2026) — Hacker News
- gentlewater on Hacker News — “Astra and Fable are not the every day workhorse… they're the tool you break out when you need the absolute strongest performance… finding and fixing one or two extra edge cases saves the business a lot of money, even if the cost is high.” (5 Sep 2026) — Hacker News
- xfax on Hacker News — “I used Light mode and it used up all my limits for the day and had to continue the following day.” (Astra) (5 Sep 2026) — Hacker News
- christophilus on Hacker News — “Astra generated some of the worst Odin code I've ever seen. Turns out AGI is indistinguishable from an Oracle subcontractor who hates tech and hates his job.” (5 Sep 2026) — Hacker News
- larodi on Hacker News — “with 4 agents doing my stuff on a daily basis, I feel like vomiting at some point, not mere nausea, but disgust. damn Codex seems to fare better at this imho.” (Claude Code) (6 Sep 2026) — Hacker News
- ahepp on Hacker News — “Anything I have claude or codex write carries a ton of distinctive characteristics. Obsession with ‘bit-for-bit identical’, ‘it's not the X it's the Y Z’… It's driving me nuts.” (6 Sep 2026) — Hacker News
- lucas_t_a on Hacker News — “opus 5 came by default with 200k context and 1M gated behind usage tokens, auto-compaction by default, and it keeps telling you to clear and start from scratch all the time.” (6 Sep 2026) — Hacker News
