‘Cloudflare Stopped It’: A Week of AI Agents That Can’t
The demo shows an autonomous assistant doing your chores. This week users showed what actually happens: it gets stopped at the Cloudflare gate on the way to cancel a subscription.
The demo is always the same. Someone types a plain-English instruction — book the table, cancel the plan, sort the inbox, build the app — and an AI agent trots off and does it, autonomously, while a voiceover talks about the end of drudgery. 2026 has been the year of that demo. Nearly every major lab now sells an “agent”: something that does not merely answer but acts, browsing the web, running code and touching your files on your behalf. The pitch is no longer a clever chatbot. It is a capable employee.
This week, on Hacker News, users described what actually happens when you hand these agents a chore. The gap was not in the reasoning; the models can talk a good game. It was in the doing. The browser agent gets stopped at the Cloudflare gate. The assistant that replaced a working one can no longer set an alarm. The coding tool destroys the files it was asked to edit. The free tier quietly gets worse to nudge you towards the metered one. It is a portrait of a technology that can describe the errand beautifully and then fail to run it.
This is a Voices piece: a round-up of what real users said, in their own words, over the past week. Quotes sourced from: Hacker News. The browser bridge we normally use to read Reddit and X was down for this run, so we sourced entirely from Hacker News, opened each thread, and lifted the wording verbatim; every quote is linked in the Sources list below, with the handle, platform and date. We did not quote anything we could not open and read on the live page, and we kept the people who pushed back in the room.
The agent that can’t get through the door
The single most on-the-nose failure of the week belongs to whazor, trying to use ChatGPT’s browser mode for exactly the kind of tedious admin an agent is supposed to swallow. That is our moan of the day.
It is almost too neat. The web spent a decade building bot defences to keep automated traffic out, and an AI browsing on your behalf is, to those defences, indistinguishable from a bot. So the flagship use case — let the agent handle the soul-destroying subscription cancellation — runs straight into a wall the rest of the internet put up on purpose. The agent is not wrong or broken here; it is structurally unwelcome, and no amount of model cleverness changes that.
When the agent does get through, the next problem is that it does not always do what you asked. bodge5000, using ChatGPT to comb the used market for a MacBook Pro, watched it “alternate between ignoring my requirements (which was the whole reason I used it in the first place, as eBay does the same) and doing this weird thing where it’d suggest the ‘concept’ of something (eg ‘X at Y price is a great choice’ with no link to X at Y price).” An agent that ignores the constraints that were the entire point, then gestures at a hypothetical listing it cannot actually find, is doing an impression of helpfulness rather than the thing itself.
Extraordinarily confusing, by expert account
OpenAI’s new “Work” product drew the sharpest reactions, and not only from the usual sceptics. simonw — a developer who writes patiently and generously about these tools — came away calling it “an extraordinarily confusing and very powerful product,” finding the official explanation “almost entirely useless,” and noting that “figuring this all out took way more work than it should have.” When the person who volunteers to write the explainer for everyone else is this worn down by the product, the product has a problem the benchmarks will never show.
0xbadcafebee put the same frustration more bluntly, describing the familiar shape of it: “a generic large corporation, where they spend billions on making a product, and $0 on checking if the product makes any sense to a real user.” That is the throughline of the agent era so far — enormous capability shipped with almost no thought given to whether an ordinary person can tell what it does, which mode they are in, or what it is about to touch.
Metered where it used to be free
Underneath the confusion, several users detected a direction of travel. Gareth321 framed it as a deliberate funnel: “OpenAI has been slowly strangling context limits, task timeframes, tools, etc, for chat. They’re trying to encourage people to use ‘Work’ because chat is free and Work is metered. Expect chat capabilities to degrade further over time.” We flag that as one user’s interpretation rather than a documented policy — but it rhymes with a pattern we keep logging, in which paid features quietly disappear and the free tier thins out just as a metered upgrade appears alongside it.
The safety dimension of “agents that act” did not go unnoticed either. coder-pm pointed at the part most users never see: “regular users won’t even have knowledge it’s touching their machine, filesystem… We should never trust it won’t touch forbidden places.” That instinct — set hard boundaries, assume it will overstep — is exactly the caution the marketing skips, and it is the same worry that made us wary when acting without asking became the default in a coding agent.
The upgrade that can’t do the old job
Nowhere was the gap between newer and better wider than on Android, where Google has been replacing the old Assistant with Gemini. fy20 described the downgrade with weary precision: “Google Assistant is being replaced by Gemini. Except it seems Gemini can’t actually do any assistant tasks. If I ask it to play a song, instead of triggering Spotify it just gives me a list of URLs I can play the song. Same with alarms.” A cleverer model that cannot set an alarm is not an upgrade; it is a regression with better press.
SurgeArrest had paid for the privilege: “Half of the time it can’t understand you, the other half just fails to output information,” after “subscribing to Gemini AI Pro” in the hope of getting the phone-grade assistant and instead getting something that felt like “2023-2024 at best.” And the accidental-subscription theme recurred with robotmay, who found Gemini “surprisingly bad” at image generation and admitted holding “a Gemini subscription for the month after mistakenly thinking I’d get cheap Opencode usage through it, so gotta use it for something.” Paying for a thing you did not mean to buy, to use a feature that underwhelms, is the small print of the AI subscription era. This forced-replacement complaint is one we have heard before, when users found Gemini quietly declining to do the obvious thing and not telling them.
It still just says no
And then there is the oldest agent failure of all: the refusal. knollimar, trying to get a model to reproduce a public article as continuous prose, hit a wall built out of caution: “my claude refused, giving me some copyright-without-mentioning-copyright bs.” An assistant that declines an ordinary request and will not even name its reason is the opposite of the confident, capable agent in the advert — and it is a reminder that the same guardrails meant to keep these tools safe also, routinely, keep them from being useful.
The churn tax
Even the good agents impose a cost when the ground shifts under them. habosa, caught in last week’s news that OpenAI is cutting Cursor off from its models, described the lock-in from the inside: “Cursor’s harness is so much faster than anything else I’ve used that I just can’t give it up, using Claude Code feels like being on 0.25x speed,” only to conclude that losing the OpenAI model he relied on was “effectively kicking me off of Cursor.” The tool works; the corporate weather above it does not, and the user pays the switching cost either way.
To be fair: the honest line
Not everyone in the thread was there to complain, and the fair voices drew a useful line. manmal made the case for what the tools genuinely are: “They are better at many things I’m a novice at. At least at the operational level. I can’t rely on the semantics being fully correct because they tend to get subtle things wrong or just don’t ask. Like tax forms — not a good idea to put them on full auto.” That is the whole argument in three sentences: a superb assistant for things you could check yourself, a dangerous autopilot for things you can’t. The failure is not the capability; it is the word “auto.”
Others aimed at the hype rather than the tool. tomlockwood reduced a breathless launch claim to one word: “‘It broke out of containment and hacked computers on its own!!!’ Marketing.” The scepticism is healthy, and it is aimed where it should be — at the pitch, not the people using the product and finding its edges.
What to take from the week
If there is a posture in all of this, it is to buy the assistant and refuse the autopilot:
- Judge agents on doing, not saying. The models are fluent; the question is whether the thing actually completed — the alarm set, the subscription cancelled, the file edited without collateral damage.
- Assume it will be blocked, and have a fallback. An agent that browses on your behalf will hit bot defences, logins and captchas. If a task matters, know how you’ll finish it by hand.
- Never full-auto anything you can’t check. Tax forms, money, anything that touches your files: keep a human in the loop, because the tool won’t tell you when it quietly got a subtle thing wrong.
- Watch what you’re paying for. The free tier can thin out while a metered upgrade appears next to it. Notice the meter before it notices you.
None of this is an argument against the technology; the people quoted here are the ones bothering to use it, which is why they can describe the seams so precisely. The gap they keep pointing at is the one between agent as a marketing word and agent as a thing that reliably completes your errand. The models have learned to sound like a capable colleague. Getting them to actually do the job — through the Cloudflare gate, without ignoring your instructions, without saying no, without breaking the thing they touched — is turning out to be the hard part, and the part the demo never shows.
Frequently asked questions
What’s the difference people are drawing between a chatbot and an ‘agent’?
A chatbot answers; an agent is supposed to act — browse the web, run code, touch your files, complete a multi-step task on your behalf. The 2026 marketing is almost entirely about agents. This week’s complaints cluster on the acting part: users found the models can often say the right thing but fail to reliably do the thing, whether that’s completing a purchase, cancelling a subscription, setting an alarm or editing a file without breaking it.
Why would a browsing agent get blocked by Cloudflare?
Because bot-detection systems like Cloudflare exist specifically to stop automated traffic, and an AI browsing on your behalf looks like automated traffic. One user reported that ChatGPT’s browser mode was “fully banned by Cloudflare”, so it couldn’t complete the very chore — cancelling a phone subscription — that an autonomous agent is pitched to handle. It’s a structural problem: the web’s defences don’t distinguish between a malicious bot and your assistant running an errand.
Is the free tier really being ‘strangled’ to push people to paid?
That’s one user’s characterisation, and it’s an interpretation rather than a documented policy — we quote it as a widely shared impression. The specific claim was that OpenAI has been narrowing context limits, task timeframes and tools on free chat while steering users towards the metered “Work” product. Whether by design or by resource constraint, several users described the same felt experience: the free thing getting less capable while a paid, metered alternative is promoted.
Are these failures just people using the tools wrong?
Sometimes, and the thread included that fairness. One commenter noted the models are genuinely useful for tasks you’re a novice at, but get subtle things wrong and don’t ask, so you shouldn’t put them on “full auto” for something like tax forms. That’s a reasonable account of a real skill gap. But “play a song”, “cancel a subscription” and “don’t rewrite my requirements” aren’t exotic prompts — when a product marketed as an autonomous agent fails them, that isn’t obviously the user’s fault.
Sources
- whazor on Hacker News — “The browser mode is great, except its fully banned by Cloudflare. I tried cancelling a phone subscription but Cloudflare stopped it.” (31 Aug 2026) — Hacker News
- Gareth321 on Hacker News — “OpenAI has been slowly strangling context limits, task timeframes, tools, etc, for chat… Work is metered. Expect chat capabilities to degrade further over time.” (31 Aug 2026) — Hacker News
- bodge5000 on Hacker News — ChatGPT “alternated between ignoring my requirements (which was the whole reason I used it in the first place)… ‘X at Y price is a great choice’ with no link.” (31 Aug 2026) — Hacker News
- simonw on Hacker News — “It is an extraordinarily confusing and very powerful product… Figuring this all out took way more work than it should have.” (31 Aug 2026) — Hacker News
- 0xbadcafebee on Hacker News — “a generic large corporation… spend billions on making a product, and $0 on checking if the product makes any sense to a real user.” (31 Aug 2026) — Hacker News
- coder-pm on Hacker News — “regular users won’t even have knowledge it’s touching their machine, filesystem… We should never trust it won’t touch forbidden places.” (31 Aug 2026) — Hacker News
- fy20 on Hacker News — “Google Assistant is being replaced by Gemini. Except it seems Gemini can’t actually do any assistant tasks… instead of triggering Spotify it just gives me a list of URLs.” (28 Aug 2026) — Hacker News
- SurgeArrest on Hacker News — “Half of the time it can’t understand you, the other half just fails to output information… subscribing to Gemini AI Pro… instead got something [worse].” (28 Aug 2026) — Hacker News
- knollimar on Hacker News — “my claude refused, giving me some copyright-without-mentioning-copyright bs.” (30 Aug 2026) — Hacker News
- robotmay on Hacker News — “Gemini’s surprisingly bad at it… I have a Gemini subscription for the month after mistakenly thinking I’d get cheap Opencode usage through it, so gotta use it for something.” (30 Aug 2026) — Hacker News
- habosa on Hacker News — “Cursor’s harness is so much faster… using Claude Code feels like being on 0.25x speed… this change is effectively kicking me off of Cursor.” (30 Aug 2026) — Hacker News
- tomlockwood on Hacker News — “‘It broke out of containment and hacked computers on its own!!!’ Marketing.” (31 Aug 2026) — Hacker News
- manmal on Hacker News — “They are better at many things I’m a novice at… they tend to get subtle things wrong or just don’t ask. Like tax forms — not a good idea to put them on full auto.” (31 Aug 2026) — Hacker News
