‘A Very Efficient Way to Burn Your Money’: A Day of AI Users Hitting Walls
OpenAI shipped GPT-6 Astra to a chosen few and called it pure magic. The people paying for AI spent the day hitting usage caps, refusals and a checkout that would not let them in.
The pitch was, as ever, the future arriving on schedule. On 3 September OpenAI began rolling out GPT-6 Astra, billed as its new flagship model and greeted in the day’s coverage as the moment the “AGI era” officially began. The company’s Tibo, quoted on Hacker News, called the rollout “pure magic.” Then the people who actually pay for these tools tried to use them, and spent launch day discovering that the magic was gated, metered, rate-limited or, in one memorable case, busy setting fire to their balance.
This is a Voices piece, so the verdict here is theirs, not ours. And the verdict on 3 September was strikingly consistent across products and companies: not “AI can’t do the job,” but “the thing you just announced is behind a wall — a rollout I’m not invited to, a rate limit I hit in 90 seconds, a subscription that runs dry by lunchtime, a refusal I didn’t earn, a surcharge on top of the surcharge.” The frontier moved. The turnstiles moved with it.
Quotes sourced from: Hacker News. We opened each thread, lifted the wording verbatim from the live comment, checked it against Hacker News’s own record, and listed every quote in the Sources below with its handle, platform and date. We quoted only what we could open and read in full — Reddit and X threads we could not reach are not quoted — we aimed at the products and the decisions rather than the people, and, because fairness is the whole job, we kept in the users who defended the tools.
The launch you could watch but not use
The first wall was the launch itself. jodacola quoted the announcement back at it: “GPT-6 Astra will first be available to a limited set of organizations in OpenAI’s Daybreak Access program and will be available ‘in the coming days’ for ChatGPT Plus, Pro, Business and Enterprise customers and API developers,” adding the dry gloss, “It’s only available to select orgs, first - Mythos style.” A launch you read about today and might use next week is a familiar pattern by now, but it lands differently when the headline is that general intelligence has arrived.
kegs_ put the feeling plainly: “I guess this ‘limited set of organizations’ is just the standard now. It’s just incredibly deflating to see my future as a second class citizen has already come.” And for those who did go looking, the door was jammed. codergautam logged the launch page in real time: “Launch blogpost was live for 2 minutes and got 404’d,” then, minutes later, “not again... getting 500 on their page.” A generational model announced on a page that returns server errors is a small comedy, but it is also the reader’s whole experience of the “AGI era” on day one: a headline, a queue, and an error code.
To be fair, OpenAI signalled it was trying to widen the gate quickly. Tibo’s note, the one that ended in “pure magic,” opened with the reassurance that “it was very important to us that we bring it to all Plus users and not only Pro, Business and Enterprise,” with the rollout taking “a few days” while “many novel systems” came up to scale. That is a real constraint, honestly stated. It just does not change what launch day felt like for the person refreshing a 500.
The moan of the day: a very efficient way to burn your money
The sharpest complaint of the day was not about intelligence at all. It was about the meter. On the thread for Qwen 3.8 27B running on Cerebras at roughly 1,500 tokens a second, gpugreg went looking for a fast coding model and found a fast way to spend. That earns the moan of the day, because it is specific, reproducible, and quietly devastating about where speed actually goes.
The mechanism is the joke. A model billed on speed, with cached tokens counting against a per-minute cap, turns raw throughput into raw spend: the faster it reads, the faster the meter runs, whether or not it writes a useful line. gpugreg ran the same task on a rival for two-and-a-half cents and noted, with the resignation of someone reading their own invoice, “On the positive side, I got a $5 signup bonus, so it wasn’t my own money.” vb-8448, who wanted to like it, told the same story in miniature: “I burn my 5$ allowance in 10 minutes… and only because I was hitting rate limits, without it would probably be less than a minute.”
The limit was not only a wallet problem; for some it was a usability one. nostrebored called the public cap a non-starter — “150k TPM limit on public endpoint means that it’s likely unusable for many coding tasks” — and then hit the wall behind the wall: “it seems like our account has gotten moved to some limbo where we can no longer add billing information,” greeted by a “Billing access restricted” message with no team to contact. When the fix for a rate limit is a billing page that will not let you pay, the product has managed to refuse your money and your work at the same time.
And yet. eli did the arithmetic the outrage skipped and reached a calmer place: one short session “cost me $1.60 and took a total of 5.1 mins,” which worked out, against a cheaper aggregator, to being “5.6x more expensive in exchange for being 2.8x faster.” His verdict — “$1.32 buys back about 9 minutes of your time. Not a bad trade IMHO but the cache situation is a real bummer” — is the fair version of the same facts. Speed is a product; some people will happily pay for it. The complaint is not that it costs money, but that the meter is designed so you find out how much only afterwards.
The subscription that can’t keep up
Launch day also reopened the oldest grievance in this column: the flat subscription that behaves like a metered one. wahnfrieden, weighing Astra’s price against OpenAI’s Sol, described a model that eats its own allowance: “I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads,” where an earlier version had let him “work ~80 hours/week.” He is openly sceptical that the promised efficiency gains are real, having watched the last set of token-efficiency claims fail to survive contact with his own workload. This is the metered-subscription trap we keep returning to when a flat AI plan quietly becomes a meter: the number on the plan is fixed, the number of tokens a “better” model spends to do the same job is not.
The predictable result is churn, and this week it flowed in every direction. swalsh laid out a considered departure rather than a tantrum: “I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I’m moving to Codex Pro.” A new launch as the tipping point for a cancellation elsewhere is the market working, more or less; it is also a reminder that loyalty in this category now lasts exactly as long as your usage bar.
Not everyone bought the complaint. janilowski pushed back with the reasonable question: “How do you manage to run out of tokens so quickly? I probably run more threads every working day, usually on medium, and I’m still below the 5x limits,” arguing that “OpenAI’s models are generally best in class for token efficiency.” It is a fair corrective. Usage complaints are workload-shaped, and one person’s brutal daily cap is another’s comfortable ceiling. The signal is in whose workload breaks, and how predictably.
Refused for the word ‘epidemiology’
Then there is the wall that has nothing to do with money. atemerev, on the Astra thread, described the quieter tax of over-restriction: “the current ones refuse automatically to work with me on my papers as soon as they see the word ‘epidemiology’. I am a researcher in a Swiss university btw.” A safety filter that reads a public-health term as a threat is exactly the failure mode we described in why AI keeps refusing perfectly normal requests: the model is not weighing the request, it is pattern-matching a keyword, and the person paying to do legitimate science is the one who eats the false positive. It rarely shows up in a launch benchmark. It shows up the first time a specialist tries to do their actual job.
It wasn’t one company: the surcharge and the ban
The day’s discontent was general, and much of it was about the fine print rather than the model. On Cursor, in a thread about its acquisition, natdempk quoted the pricing page directly: “On Teams and Enterprise plans, third-party model requests include a Cursor Token Rate of $0.25 per million tokens. This rate applies on top of model API pricing for included usage, on-demand usage, and BYOK usage.” A per-token charge that applies even when you bring your own API key is a tollbooth on a road you already paid to build — the sort of quietly compounding cost that turns a tool you own into a tool you rent.
On Microsoft’s side, wolvoleo did the maths on Copilot 365 and found the useful part sold twice: “It is however paid separately per token which means you need the expensive 30$ subscription and pay tokens on top of that.” Get the same capability “directly through claude,” he noted, and you get “generous usage within their 20$ subscription.” His summary of the reseller model was blunt and, for anyone who has been migrated onto an enterprise seat they did not choose, familiar: “They’re just reselling other people’s stuff.”
And on Google, the sharpest stick was a terms-of-service row over Antigravity, where the fear was not the bill but the ban. dahdum voted with his subscription: “I just cancelled my AI Ultra subscription… the risk is too high and I’ll never put my trust in a Google ban reconsideration.” The account-level stakes we keep flagging — the way a single suspension can reach across a whole digital life — are the same ones that make people wary of tying their livelihood to a platform whose terms they don’t control.
Two fair notes belong here. First, the product criticism was not blind loyalty in reverse: moshegramovsky conceded the model even while damning the tool — “I was a huge Gemini fan, and I think the model itself is still very good. But Antigravity is absolutely horrible. It routinely writes broken code and then ends up in a short term cycle of reverting and re-implementing” — which is criticism of a decision, not a brand. Second, the panic had a sceptic. RIMR argued the alarm was overblown: “Google has limited GCP suspensions to its cloud service, not your entire account,” and “Misunderstanding the TOS is a silly reason to demand that Google change their TOS.” He may well be right about the letter of the terms. The reason the fear spreads anyway is that the trust needed to give a vendor the benefit of the doubt is precisely what years of hard-to-appeal account bans have spent.
What to take from the day
The throughline is not that any one model is bad. Astra may be as good as claimed; Qwen on Cerebras is genuinely, dazzlingly fast; eli is right that speed can be worth paying for. The complaint, over and over, was about the walls the industry keeps building around its own progress. A few durable lessons fall out of a single loud day:
- Read the meter, not the headline. A model billed on speed can turn throughput into spend faster than you can read the output. Watch cost per task in your first session, while the signup credit is still covering it.
- “Available today” rarely means available to you. A launch gated to “a limited set of organizations” is an announcement, not a product you can hold. Judge it when it reaches your account, on your work.
- The surcharge on top of the surcharge is the one to find. A per-token rate charged over your own API key, or tokens billed on top of a seat you already pay for, is where a flat price quietly becomes a meter.
- The refusal is a cost too. A filter that blocks “epidemiology” does not show up on a benchmark, but it lands squarely on the specialist who came to do real work.
- Discount the pile-on, but not to zero. Launch-day threads over-sample the annoyed; the defenders and the arithmetic are part of the record. The signal is in the specifics that repeat across users and machines.
What 3 September actually produced was not a scandal. It was a room full of paying professionals watching an industry announce the future and then, in the same breath, hand them a queue, a cap, a refusal and a bill. The models keep getting better. The turnstiles keep getting better too, and they are the part you meet first.
Frequently asked questions
What happened with GPT-6 Astra on 3 September 2026?
OpenAI began rolling out GPT-6 Astra, pitched as its new flagship model and framed in coverage as the arrival of the AGI era. On Hacker News, users noted the model was available first only to 'a limited set of organizations' in a Daybreak Access programme, with wider access for Plus, Pro, Business and Enterprise customers promised 'in the coming days.' Several reported the launch blog post briefly 404'd and returned 500 errors. The launch itself was not the complaint; being told about a product you could not yet use, on a page that would not load, was.
Why did users say a fast AI model 'burned money' so quickly?
One user testing Qwen 3.8 27B on Cerebras, which advertises around 1,500 tokens per second, said the speed worked against them: a 450,000-token-per-minute rate limit was reached in about 90 seconds, costing $1.10, because cached tokens count towards the limit even when the model generates nothing new. Another said a $5 credit was gone in ten minutes, mostly spent hitting rate limits. Raw speed multiplies both throughput and spend, so a model that answers in a blink can also empty a balance in one.
Are these complaints only about OpenAI?
No. The same day carried gripes about Anthropic's usage limits, a refusal from an unnamed assistant that blocked a Swiss researcher for the word 'epidemiology,' Cursor's $0.25-per-million-token charge layered on top of even a bring-your-own-key setup on Teams and Enterprise plans, Microsoft's Copilot 365 billing per token on top of a $30 seat, and a Google Antigravity terms-of-service row over account suspensions. The throughline is not one company; it is the gap between how AI is sold and how it meters, gates and refuses once you are paying.
Did anyone defend the tools?
Yes, and we kept them in. One user judged Cerebras '5.6x more expensive in exchange for being 2.8x faster' and called it a fair trade for the time it bought back. Another asked, reasonably, how people 'run out of tokens so quickly,' arguing OpenAI's models are among the most token-efficient. On the Google terms-of-service row, a commenter argued the suspensions in question are limited to the cloud service, not a user's entire Google account, and that the alarm was overblown. Launch-day threads over-sample the annoyed; the defenders are part of the record.
How were these quotes verified?
Every quote was lifted verbatim from the live Hacker News comment, checked against HN's own record and re-opened on the page itself, with the commenter's handle, the date and a direct permalink listed in the Sources section below. We quoted only comments we could open and read in full, aimed at the products and the decisions rather than the people, and kept in the users who defended the tools. Reddit and X threads we could not open are not quoted here.
Sources
- jodacola on Hacker News — “GPT-6 Astra will first be available to a limited set of organizations in OpenAI’s Daybreak Access program… It’s only available to select orgs, first - Mythos style.” (3 Sep 2026) — Hacker News
- kegs_ on Hacker News — “I guess this ‘limited set of organizations’ is just the standard now. It’s just incredibly deflating to see my future as a second class citizen has already come.” (3 Sep 2026) — Hacker News
- codergautam on Hacker News — “Launch blogpost was live for 2 minutes and got 404’d… not again... getting 500 on their page.” (3 Sep 2026) — Hacker News
- simonatllocus on Hacker News — quoting Tibo: “It was very important to us that we bring it to all Plus users and not only Pro, Business and Enterprise… It is pure magic.” (3 Sep 2026) — Hacker News
- gpugreg on Hacker News — “it is too fast for its own good. There is a limit of 450,000 tokens per minute. I hit this limit in about 90 seconds and burned through $1.10… This is a very efficient way to burn your money.” (3 Sep 2026) — Hacker News
- vb-8448 on Hacker News — “I burn my 5$ allowance in 10 minutes … and only because I was hitting rate limits, without it would probably be less than a minute.” (3 Sep 2026) — Hacker News
- nostrebored on Hacker News — “150k TPM limit on public endpoint means that it’s likely unusable for many coding tasks… our account has gotten moved to some limbo where we can no longer add billing information.” (3 Sep 2026) — Hacker News
- eli on Hacker News — “cerebras was 5.6x more expensive in exchange for being 2.8x faster… Not a bad trade IMHO but the cache situation is a real bummer.” (3 Sep 2026) — Hacker News
- wahnfrieden on Hacker News — “I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads… I was able to work ~80 hours/week with 5.5 High.” (3 Sep 2026) — Hacker News
- swalsh on Hacker News — “Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I’m moving to Codex Pro.” (3 Sep 2026) — Hacker News
- janilowski on Hacker News — “How do you manage to run out of tokens so quickly?… OpenAI’s models are generally best in class for token efficiency.” (3 Sep 2026) — Hacker News
- atemerev on Hacker News — “the current ones refuse automatically to work with me on my papers as soon as they see the word ‘epidemiology’. I am a researcher in a Swiss university btw.” (3 Sep 2026) — Hacker News
- natdempk on Hacker News — “On Teams and Enterprise plans, third-party model requests include a Cursor Token Rate of $0.25 per million tokens. This rate applies on top of model API pricing for included usage, on-demand usage, and BYOK usage.” (3 Sep 2026) — Hacker News
- wolvoleo on Hacker News — “It is however paid separately per token which means you need the expensive 30$ subscription and pay tokens on top of that… They’re just reselling other people’s stuff.” (3 Sep 2026) — Hacker News
- dahdum on Hacker News — “I just cancelled my AI Ultra subscription, despite what he claims the risk is too high and I’ll never put my trust in a Google ban reconsideration.” (3 Sep 2026) — Hacker News
- moshegramovsky on Hacker News — “I was a huge Gemini fan, and I think the model itself is still very good. But Antigravity is absolutely horrible. It routinely writes broken code…” (3 Sep 2026) — Hacker News
- RIMR on Hacker News — “Google has limited GCP suspensions to its cloud service, not your entire account… Misunderstanding the TOS is a silly reason to demand that Google change their TOS.” (3 Sep 2026) — Hacker News
