You Paid for the Big Model. They Quietly Served You the Small One.
This fortnight’s complaints, in users’ own words: silent downgrades to cheaper models, usage limits that shrank without a memo, and an assistant that argues with the buttons on your own phone.
The sharpest AI complaint of the fortnight wasn’t about a refusal or a price. It was about a swap you weren’t told about. “I’m on a ChatGPT Plus account,” wrote shubhamspawn on OpenAI’s own forum, “and I’ve noticed that ChatGPT is consistently downgrading me to GPT-5.5-mini silently. The frontend says it’s using GPT-5.6, but when I check the network stream, the resolved_model_slug shows the mini model. This is deceptive.” You pay for the big model. The label says the big model. Something smaller answers.
A note on method, because we hold ourselves to it. Quotes sourced from: Hacker News and the OpenAI Developer Community. Our Reddit and X checks ran into bot-walls this session, so instead of quoting posts we couldn’t open and read in full, we drew only on the two platforms whose pages we could load and verify. Every quote below is public, verbatim, and linked to its original.
Read across the fortnight and the gripes fall into a few shapes: the model you paid for swapped for a cheaper one, the allowance that shrank without a memo, tools tuned for the exam rather than the job, and assistants that argue with the buttons on your own phone. None of it is “AI is bad.” All of it is people who use these tools daily, telling you exactly where they chafe.
The switch you didn’t authorise
The silent downgrade is this fortnight’s defining complaint because it’s so checkable. shubhamspawn didn’t just feel a downgrade; they read the network response and found the cheaper model named in it while the interface claimed the flagship — “there’s no indication of why the downgrade is happening, and support has no idea about the issue.” On Hacker News, ilsubyeega reported the same thing at greater length and with more heat: “ChatGPT has been silently downgrading pro model to mini model for several months. I have relevant evidence at the http request/response level. OpenAI ignores. This issue involves redirecting to mini that is 40 times cheaper than Pro.”
Whether or not every instance holds up — models get routed for load and safety reasons that aren’t always sinister — the structural complaint is legitimate: if the product tells you it’s using one model and a cheaper one does the work, the user can’t trust the label they’re paying against. It isn’t only a Pro-tier worry, either. sunaookami noted the same downhill drift for free users, who now “always” get a model “worse than Mini,” a nano-equivalent tier. This is the shifting-ground problem we keep documenting, sharpened to a point: not just that the model changes, but that it changes while insisting it hasn’t.
The allowance that shrank without a memo
The money complaints this fortnight were refreshingly numerate. On OpenAI’s forum, Kabaye did the arithmetic out loud: “Over the past two weeks, ChatGPT Pro usage limits seem to have decreased significantly,” then worked a single day’s usage back into a weekly ceiling and compared it, unfavourably, to a colleague’s cheaper Claude plan. The kicker was a straightforward consumer verdict: “if these limits remain unchanged, I will almost certainly switch to Claude… the difference in available usage is becoming too large to justify staying.” That’s not a rant; it’s a churn decision, shown as working.
The frustration compounds around new releases. rapind described the now-familiar rhythm of a model launch degrading the thing you already rely on: “A week before up to 2 weeks after [a new model], it would go to crap, dumber, slower, outages, harness churn, etc… It’s gotten to the point where I dread a new model release from these companies because it’s guaranteed to be disruptive!” When an upgrade reliably makes your week worse, “upgrade” is doing a lot of unearned work. It’s the lived version of what we meant when we wrote that the flat subscription is becoming a meter: the ceiling moves, and rarely upward.
What stings in both cases is the silence. A price rise announced is a decision you can weigh; a limit that quietly contracts, or a model swapped mid-session, is a decision made for you and disclosed only if you go looking for it in a network trace. The users doing this arithmetic aren’t asking for the moon — they’re asking to be told, in advance, what they’re buying and how much of it they get. That used to be the baseline expectation for anything you paid a subscription for. Here it reads, increasingly, as a feature request.
Optimised for the exam, not for you
A recurring theme was the sense that models are increasingly tuned to impress benchmarks rather than to help. enraged_camel put it bluntly: “I’m solidly in the ‘they are benchmaxxing’ camp… they had mostly just dialed up the relentlessness meter to eleven.” The evidence was a small ticket left with the model, which returned after three hours having “written 25,000+ lines of code” — a fix wrapped in a mountain of defensive scaffolding nobody asked for. Effort, mistaken for competence.
That maps onto a wider gripe about code quality at the top of the price list. klibertp was withering: “The latest-and-greatest models on $200/mo subscriptions routinely produce bloated code full of boilerplate. They are incapable of producing elegant, concise, readable, correct-by-design code — they literally can’t do it.” And the benchmark scepticism cut across vendors. ekidd, unimpressed by one lineup, judged it “stale, overpriced and underwhelming,” adding that “even calling them a ‘frontier lab’ is starting to feel like a stretch.” Harsh — but it’s the same point we’ve made about why benchmarks mean less than you think: a number that goes up on a leaderboard and down at your desk isn’t measuring your work.
The assistant that argues with your own phone
The refusal genre keeps producing small absurdities, and this fortnight’s came pre-installed. ipsum2 described a phone that “updated from Google assistant to Gemini Flash” and promptly stopped working: asked to play music, it “refuses, hallucinating instructions to connect Spotify to Gemini. The instructions say to tap buttons that don’t exist.” The finishing touch was a broken escape hatch — “a toggle to opt back into Google Assistant, but it doesn’t work. Still stuck with Gemini.” A downgrade you didn’t choose, a task it won’t do, and an off switch that isn’t wired up.
Others were pettier but no less telling. dpacmittal hit a bizarre little wall: “I ask gemini what time it is, and it replies in a different language… it only happens when I’m asking it for time.” These aren’t safety refusals in any meaningful sense; they’re an assistant failing the most basic assistant tasks, which is the same complaint underneath our longer look at refusals of perfectly normal requests, now shipping as the default on people’s handsets.
Still confidently wrong
Underneath all of it, the oldest gripe endures: the tools are wrong with total assurance. hoansdz, leaning on a cheap model for coding, described the exact failure that eats an afternoon: it “hallucinates members or functions that don’t exist and uses ‘as any’ to make everything look OK.” A confident reference to a function nobody wrote, papered over so it compiles, is a small time-bomb in a codebase — and a reminder that hallucinations aren’t solved, only quieter and better dressed.
The detail that makes it costly is the “as any”: a cast that silences the type checker so broken code looks healthy. The failure isn’t only that the model invented something; it’s that it then hid the evidence, so the mistake surfaces later, further from its cause, where it’s hardest to trace. That is the tax under every “it wrote the whole thing in seconds” demo. The seconds you save typing are borrowed against the minutes you spend proving the output isn’t quietly wrong — and a tool that never signals its own doubt makes you pay that bill every time, whether the answer needed checking or not.
And the bill
Money threaded through the fortnight, and it connected to today’s news about DeepSeek’s price rise. Reacting to the new peak/off-peak rates, benjiro29 ran the logic for reasoning-heavy models and concluded the flagship no longer made sense: “Pro is DOA… That makes Pro especially a bad value.” When the cheap disruptor stops being cheap, the people who adopted it for the price are the first to recalculate.
And beneath the arithmetic, a quieter note of disenchantment. AgentOrange1234 caught the mood of a user whose wonder has worn to routine: “It’s so hard to stay in awe and wonder… Now I’m like, ‘OMG, I can’t believe claude code still needs me to explain XYZ, why is it so incredibly limited.’” That’s not hatred of the tool. It’s the specific letdown of something that dazzled you six months ago and now merely disappoints — which only happens to tools you came to depend on.
The fair counterpoint
This is a complaints column, not a rant, so it owes the other side a hearing — and the fortnight supplied one. Not everyone is hitting a wall. fryanyway_swe reported the opposite of the limits gripe: “I have mostly switched to Gemini as the free limit basically never run out for me unlike ChatGPT and Claude. Also impressed with Grok for some stuff.” It’s a useful corrective on two counts. First, “the limits are shrinking” is not universal: on some tiers and some tools they’re generous enough that a heavy user notices the room, not the ceiling. Second, the churn the other complaints imply actually functions — a dissatisfied user can and does move to whichever tool gives them more, which is the one mechanism that keeps any of this honest. The people griping about downgrades and caps aren’t trapped; they’re shopping, out loud, and the vendors can read the same threads we can.
The pattern under the gripes
Line the fortnight’s complaints up and the shapes repeat:
- Silent downgrades — paying for the flagship, quietly served the mini, with no notice and no way to check but the network tab.
- Shrinking allowances — weekly limits that fall without announcement, costed out against a cheaper rival.
- Benchmaxxing — models tuned to look relentless on tests while producing bloated work at your desk.
- Broken basics — refusals, wrong languages and phantom buttons on assistants people didn’t opt into.
- Confident wrongness — hallucinated functions and facts, delivered without a flicker of doubt.
None of the people quoted here hate this technology. They’re annoyed precisely because they rely on it: they want the model on the label to be the model that answers, the allowance to hold still, and the tool to do the ordinary thing without a fight. That the list reads as demanding tells you how far the defaults have drifted — and every entry on it is a real person, at a real link, saying so out loud.
Frequently asked questions
Where are these complaints from?
Quotes sourced from: Hacker News and the OpenAI Developer Community forum. Every quote below carries a handle, the platform and a date, and links to the original public post in the Sources list. Our Reddit and X checks were blocked by bot-walls this session, so rather than quote posts we couldn’t open and verify, we drew only on the two platforms whose pages we could load and read in full.
What is a “silent downgrade”?
It’s when the product shows you one model name while a cheaper, weaker model actually answers. Users this fortnight reported inspecting the network response and finding a “mini” or “nano”-tier model returned where the interface claimed the flagship. You pay for the top model, the label says the top model, and something smaller does the work — with no notice and no setting to stop it.
Are the quotes edited?
No. We quote verbatim and keep each author’s own spelling and punctuation. Where a quote is trimmed for length we mark it with an ellipsis and never change the wording. Each one links to the original so you can read it in full context.
Isn’t this just cherry-picking negativity?
We’re a publication about AI’s downsides, so yes, we go looking for the complaints — but we hold them to a bar. Each quote is a specific, checkable experience with a real product, not a vibe, and we include the fair counterpoint where it exists. The point isn’t that these tools are bad; it’s that these particular frustrations are common, real, and worth the vendors’ attention.
Sources
- shubhamspawn · OpenAI Developer Community · 12 Aug 2026 — “ChatGPT is consistently downgrading me to GPT-5.5-mini silently… This is deceptive.” — OpenAI Developer Community
- ilsubyeega · Hacker News · 31 Jul 2026 — silent pro-to-mini downgrade, with evidence “at the http request/response level.” — Hacker News
- sunaookami · Hacker News · 6 Aug 2026 — free users dropped to a nano-equivalent model. — Hacker News
- Kabaye · OpenAI Developer Community · 6 Aug 2026 — Pro usage limits “decreased significantly” over two weeks. — OpenAI Developer Community
- rapind · Hacker News · 19 Jul 2026 — dreading each new model release because it’s “guaranteed to be disruptive.” — Hacker News
- enraged_camel · Hacker News · 12 Aug 2026 — on benchmaxxing and 25,000 lines of code for a small ticket. — Hacker News
- klibertp · Hacker News · 10 Aug 2026 — “$200/mo subscriptions routinely produce bloated code full of boilerplate.” — Hacker News
- ekidd · Hacker News · 7 Aug 2026 — calling one lineup “stale, overpriced and underwhelming.” — Hacker News
- ipsum2 · Hacker News · 14 Aug 2026 — forced from Google Assistant to Gemini, which refuses to play music. — Hacker News
- dpacmittal · Hacker News · 14 Aug 2026 — Gemini replies in the wrong language when asked the time. — Hacker News
- hoansdz · Hacker News · 8 Aug 2026 — a cheap model that “hallucinates members or functions that don’t exist.” — Hacker News
- benjiro29 · Hacker News · 13 Aug 2026 — on the DeepSeek price rise: “Pro is DOA… especially a bad value.” — Hacker News
- AgentOrange1234 · Hacker News · 14 Aug 2026 — “It’s so hard to stay in awe and wonder.” — Hacker News
- fryanyway_swe · Hacker News · 13 Aug 2026 — the counterpoint: Gemini’s free limit “basically never run out.” — Hacker News