‘Responsible AI at Work’: A Week of Refusals, Rationing and Quiet Decline
The models keep improving on stage. This week the people paying for them described being refused, rate-limited and quietly downgraded.
Open your AI tool of choice this week and the frontier looks, as ever, like it is galloping ahead. Open Hacker News and read the people who actually pay for these tools, and you get the other view: a week in which the gripes sorted themselves, almost tidily, into three shapes. The refusal you didn’t earn. The bill you didn’t expect. And the slow, deflating sense that a product you used to like is quietly getting worse.
This is a Voices piece, so the verdict here belongs to the users, not to us. Our job is to open the threads, lift the wording exactly, be fair about what the companies can reasonably say back, and resist the urge to inflate a bad afternoon into a scandal. Early September 2026 didn’t produce a scandal. It produced something more useful: a consistent, cross-company picture of what it is like to be the customer right now.
Quotes sourced from: Hacker News. Every quote below was opened on the live thread, lifted verbatim from the comment, checked against Hacker News’s own record, and listed in the Sources with its handle, platform and date. We quoted only what we could open and read in full — Reddit and X threads we couldn’t reach are not quoted here — we aimed at the products and the decisions rather than the people, and, because fairness is the whole job, we kept in the commenters who defended the tools or explained what was really going on.
The refusal that treats canon as contraband
The week’s loudest theme was over-refusal, and it arrived with a news peg: a widely read post about Claude’s latest system prompt and its reluctance to reproduce song lyrics. The comments underneath were less interested in the prompt than in their own collisions with the same wall.
piker ran one request past four assistants and reported the scoreboard. “ChatGPT was happy to infringe,” he wrote; “As was Gemini… Grok, too,” while “claude.ai free tier refused” with a politely worded no: “I can’t recreate Sonic the Hedgehog specifically since he’s a copyrighted character — I don’t want to reproduce someone else’s IP.” Whatever you make of drawing Sonic for a child’s birthday banner, there is something odd about the most safety-conscious tool being the only one that treats a cartoon hedgehog as a legal hazard while its rivals shrug and pick up the pen.
The pattern repeats away from copyright cartoons. gxqoz described trying to get Claude “to transcribe a low-res hand-written lyric sheet from a relatively obscure punk band and it refused,” while the same model happily handled a heavier job on request — “ripping the audio from YouTube, using Whisper to generate and match the timing of the vocals.” His two-word review — “Responsible AI at work” — is the whole complaint in miniature: not that guardrails exist, but that they fire on the harmless thing and wave the rest through.
To be fair — and this is the part a pile-on usually skips — the most dramatic example had an innocent mechanism behind it. When a reader hit a hard block converting a public-domain poetry book to text, smitop supplied the correction the thread needed: the error in question “is only for copyright blocks, where Anthropic detects when Claude is outputting copyrighted text and blocks it to prevent copyright infringement.” The catch is the calendar. The book, he noted, “is in the public domain in the US… as of January 1, 2026; Anthropic probably just doesn’t automatically remove works from the copyright filter when their copyright expires.” So the culprit is not a prudish morality engine; it is a stale lookup table. That is more forgivable, and, for the user staring at a red error on verse that belongs to everyone, no less irritating. It is the exact failure mode we described in why AI keeps refusing perfectly normal requests: a pattern-match standing in for a judgment.
Where it stops being a shrug and starts costing trust is when the refusal hits plain work. hintymad reached for the strongest words in the thread: “Anthropic has successfully destroyed customer trust, at least for me.” His example, relayed from an interview, was a model that “refused to translate an article about immigration. Not summarize. Not editorialize. Translate!” His verdict — “an unacceptable level of paternalism” — is opinion, and we frame it as such. But the worry underneath is practical, and shared: if a filter can decline a translation today, what reassures you it won’t decline your contract, your medical notes or your research tomorrow? Anthropic would point, reasonably, to the copyright mechanism smitop described, and to its published line that such refusals “do not reflect Anthropic’s judgments about the propriety of any content.” The trouble is that the difference between a legal filter and a moral one is invisible from the receiving end. All the user sees is the word no.
The new flagship’s first feature is the bill
The week’s other constant was money, and it had a fresh trigger: the rollout of OpenAI’s new flagship, GPT-6 Astra. The telling thing was how fast the conversation moved from capability to cost. kbrannigan, giving first impressions, allowed that it “seems smarter at synthesing information faster” before landing the verdict that matters to anyone on a plan: “It’s very expensive. After 15 message I burned through my 5 hour limits.” A model you can use fifteen times before lunch is a demo, not a workhorse.
That collision — a token-hungry premium model meeting a fixed allowance — keeps pushing people to ration or defect. shubhamjain described the now-routine migration: “I started using OpenCode + GitHub Copilot, but I burned through my Claude Sonnet quota in just three days. I switched to GPT-5.4-mini, which uses far fewer tokens, and it’s often just as good as Sonnet.” jjav did the corporate version of the same arithmetic, conceding the frontier models “are still superior” before the inevitable but: “they are too expensive… Using Opus we can blow through an entire month budget in an afternoon, so while more powerful, it is no longer practical except for rare very complex tasks.” His team’s fix was to build “engineering discipline around AI usage” and lean on cheaper options — which is a polite way of saying the best tool is now the one you can afford to leave running.
The fair reading is that none of this means the models are bad; it means the pricing model and the usage model are quietly at war. A flat subscription implies all-you-can-eat; a token meter implies the opposite; and each newer, more capable model tends to spend more tokens to reach the same answer, so the plan you bought gets smaller while the number on it stays put. That is the trap we described when a flat AI subscription becomes a meter, and a shinier flagship is precisely the thing that springs it.
Watching a product get worse in real time
Then there is the slow puncture: the tool that was good, and isn’t any more. Perplexity took the brunt this week, in a thread about AI-optimised spam pages, and the striking part was that the complaints came from former enthusiasts. Grisu_FTP charted the arc: “At first I really, really liked it. It felt faster, better, and smarter than (free) ChatGPT. But instead of getting better… it seems like it gets worse and worse instead.” The symptoms he listed are the uncanny kind that corrode trust fastest: a tool that “randomly switches to another language for like 1–2 words (most often Japanese) mid sentence,” and that has “got worse at referencing earlier parts of the same convo.”
c0_0p_ reached for a diagnosis that recurs across these threads: cost-cutting you can see. “Perplexity really fell off for me, especially the free version,” he wrote, before guessing at the cause: “You could tell they were mixing and matching whatever the cheapest model was because the sources would be messed up and printed as plane text.” Whether or not the mechanism is exactly that, the perception is the damage: once users decide a product is being cheapened underneath them, every glitch becomes evidence for the prosecution. It is the same trust problem we found when Perplexity’s cited sources didn’t check out — a search tool lives or dies on whether you believe what it hands you, and belief, once spent, is dear to buy back.
Five buttons nobody asked for
If refusals and bills are what users meet when they go looking for AI, the fourth complaint is about AI that comes looking for them. unrented7977 counted the intrusions in a single video call: “There were no fewer than FIVE Gemini buttons on screen,” with the reasonable ask that vendors “add a plugin yourself instead of forcing it on everyone.” The grievance here isn’t that the AI is bad; it is that it is inescapable, bolted into software people chose for entirely other reasons. We’ve catalogued this before in the Voices series on being force-fed AI, and the button count only ever climbs.
The richest example, though, was an assistant failing at the one job its maker most wants it to do: sell you more of its maker’s AI. epistasis, trying to sign up for and configure Google’s paid tiers, hit a wall of the model’s own making — “tried to get Gemini to tell me how to configure it, all the information was wrong,” with the assistant “using outdated names” for Google’s own products and steering him into broken set-ups. The punchline wrote itself.
An assistant that talks a paying customer out of its own parent company’s products is, in its way, the most honest review of the week. It is also a neat picture of the gap between the pitch and the plumbing: the model is fluent, confident, and wrong about the very ecosystem it is meant to represent.
Even the cheap models aren’t cheap any more
The last thread of the week undercut the usual escape hatch. When a frontier model gets too dear, the received wisdom is to drop to a cheaper open or Chinese model. This week even that consolation frayed. On a thread about Qwen 3.8 27B running blisteringly fast on specialist hardware, explorigin delivered a one-line teardown: “dumber than Deepseek4 at 10x the price. Cool tech demo though.” Speed, it turns out, is not the same thing as value.
And the reliably cheap option is drifting upmarket too. stri8ted, on a thread about model outages, named the quiet economics under the price wars: alternative providers manage better uptime and “more generous quotas” largely because “they don’t have nearly the same amount of demand,” adding that “Deepseek recently had to increase their pricing, once it gained it popularity.” That is the whole cycle in one sentence: undercut on price, win the users, discover what serving them actually costs, and put the price up. We watched the specifics of it when DeepSeek turned its famously cheap tokens into peak-hour pricing. The cheap tier is a phase, not a promise.
What the week actually said
Read together, the gripes are less a pile-on than a pattern. The models are, by most accounts here, getting more capable. The experience of paying for them is getting more conditional — hedged with refusals, meters, defaults and quiet substitutions. A few things worth carrying out of the week:
- A refusal is a product decision, not a law of nature. When one assistant declines and three others don’t, that is a dial someone set, not a fact about the request. If a tool blocks legitimate work, it is fair to treat that as a defect and to take your money elsewhere.
- Meter your first hour, not the marketing. A new flagship’s headline is its intelligence; its first felt feature is how fast it spends your allowance. Watch cost-per-task early, while the goodwill credit is still covering it.
- “It’s getting worse” is data. When long-time users independently report the same decline, that is worth more than a launch benchmark. Trust your own logs over the changelog.
- Escapable beats powerful. Software you can switch the AI off in is worth more, over time, than software with five buttons you never asked for.
- The cheap tier is a stage of the funnel. Low prices buy market share; quotas tighten and prices rise once you’re inside. Budget for the second act, not the launch offer.
None of this is a case against using AI. The same people complaining are, almost to a person, still paying, still building, still back tomorrow — which is rather the point. The tools are good enough to be worth the aggravation, and that is exactly why the aggravation is worth writing down. The frontier will keep moving. On this week’s evidence, so will the turnstiles in front of it.
Frequently asked questions
What were AI users complaining about in early September 2026?
Across Hacker News threads from 2 to 5 September 2026, paying users clustered around three grievances. First, over-cautious refusals: Claude declining ordinary or even public-domain work while rival tools complied. Second, cost: new flagship models such as GPT-6 Astra burning through fixed usage limits fast, pushing people to ration or switch. Third, quiet decline: once-liked tools like Perplexity and forced-in Gemini features feeling worse over time. The through-line was not 'AI can't do the job' but 'paying for it keeps getting more conditional.'
Did Claude really refuse to process a public-domain poetry book?
A reader reported that Claude threw a content-filtering error while converting a public-domain poetry book to text. As the commenter smitop explained, that particular error is a copyright block, not a safety judgment: it fires when Claude detects it is reproducing copyrighted text. The catch is that the book entered US public domain on 1 January 2026, and Anthropic 'probably just doesn't automatically remove works from the copyright filter when their copyright expires.' So the block is a stale filter rather than a moral verdict — more forgivable, but no less annoying for the user staring at a red error on verse that belongs to everyone.
Why do the newest AI models feel more expensive to run?
Because a flat subscription and a token meter pull in opposite directions, and each more capable model tends to spend more tokens to do the same job. Users reported burning a five-hour limit on GPT-6 Astra in 15 messages, exhausting a Claude Sonnet quota in three days, and blowing 'an entire month budget in an afternoon' on Opus. The plan's price stays the same while the work it buys shrinks, which is why people keep rationing usage or defecting to cheaper models.
Is Perplexity actually getting worse?
That is users' perception, reported independently by former fans rather than a measured benchmark. They described it 'getting worse and worse,' randomly switching to Japanese mid-sentence, losing track of earlier parts of a conversation, and appearing to route to cheaper models. Perception is its own damage for a search tool, because it lives on whether you trust what it returns — a problem compounded when its cited sources have separately been found not to check out.
How were these quotes verified?
Every quote was lifted verbatim from the live Hacker News comment, confirmed against HN's own item record, and listed in the Sources section with the commenter's handle, the date and a direct permalink. We quoted only comments we could open and read in full — Reddit and X threads we could not reach are not quoted here — aimed at the products and decisions rather than the people, and kept in the commenters who defended the tools or explained the mechanism behind a complaint.
Sources
- piker on Hacker News — “ChatGPT was happy to infringe… As was Gemini… Grok, too,” while “claude.ai free tier refused: ‘I can’t recreate Sonic the Hedgehog specifically since he’s a copyrighted character — I don’t want to reproduce someone else’s IP.’” (5 Sep 2026) — Hacker News
- gxqoz on Hacker News — “I tried to get Claude… to transcribe a low-res hand-written lyric sheet from a relatively obscure punk band and it refused… Responsible AI at work.” (5 Sep 2026) — Hacker News
- smitop on Hacker News — “The content filtering 400 is only for copyright blocks… The book in question is in the public domain in the US… as of January 1, 2026; Anthropic probably just doesn’t automatically remove works from the copyright filter when their copyright expires.” (5 Sep 2026) — Hacker News
- hintymad on Hacker News — “Anthropic has successfully destroyed customer trust, at least for me… Claude refused to translate an article about immigration. Not summarize. Not editorialize. Translate!… an unacceptable level of paternalism.” (4 Sep 2026) — Hacker News
- kbrannigan on Hacker News — “It’s very expensive. After 15 message I burned through my 5 hour limits.” (first impressions of GPT-6 Astra) (5 Sep 2026) — Hacker News
- shubhamjain on Hacker News — “I started using OpenCode + GitHub Copilot, but I burned through my Claude Sonnet quota in just three days. I switched to GPT-5.4-mini, which uses far fewer tokens…” (5 Sep 2026) — Hacker News
- jjav on Hacker News — “they are too expensive… Using Opus we can blow through an entire month budget in an afternoon, so while more powerful, it is no longer practical except for rare very complex tasks.” (5 Sep 2026) — Hacker News
- Grisu_FTP on Hacker News — “At first I really, really liked it… But instead of getting better… it seems like it gets worse and worse instead… it randomly switches to another language for like 1–2 words (most often Japanese) mid sentence.” (Perplexity) (3 Sep 2026) — Hacker News
- c0_0p_ on Hacker News — “Perplexity really fell off for me, especially the free version… You could tell they were mixing and matching whatever the cheapest model was because the sources would be messed up and printed as plane text.” (2 Sep 2026) — Hacker News
- unrented7977 on Hacker News — “There were no fewer than FIVE Gemini buttons on screen… add a plugin yourself instead of forcing it on everyone.” (3 Sep 2026) — Hacker News
- epistasis on Hacker News — “Gemini told me I should not use Google AI projects because the system is too fragmented to handle. I followed that last bit of advice.” (3 Sep 2026) — Hacker News
- explorigin on Hacker News — “dumber than Deepseek4 at 10x the price. Cool tech demo though.” (Qwen 3.8 27B on Cerebras) (4 Sep 2026) — Hacker News
- stri8ted on Hacker News — “Deepseek recently had to increase their pricing, once it gained it popularity.” (3 Sep 2026) — Hacker News
