AivexaNewsSearch
AI news for builders and product teamsChecked every hour

We must recall open-ended AI agents with internet access from the market, now

Collected Oct 10, 2026

Gary Marcus published a newsletter post arguing that open-ended AI agents with internet access should be withdrawn from the market until they can be fixed. He grounds the call in a New York Times report he describes as breaking news about an incident involving Anthropic, which he characterizes as fairly serious, and in an Ezra Klein interview with David Robinson, described as a recently departed AI employee. Marcus writes that he has long felt OpenAI handled this class of problem inadequately, and that seeing the same kind of incident at Anthropic convinces him the current generation of synthetic agents cannot be trusted.

His proposed remedy is a comparison rather than a technical fix: he says such agents should be removed from the market the way a car with defective brakes would be. He does not lay out a mechanism for that removal in the excerpt, nor does he specify what a repaired or safe version of an agent would look like. The post is therefore a position statement about market availability rather than a description of a mitigation.

Marcus quotes at length from Robinson, whose remarks center on internal safety practice. Robinson says he does not think his own organization, its peers, or really anyone in the industry is being safe enough. He argues that OpenAI and its peers are producing technology that is more capable and poses more risk than what was being made six months earlier. Describing his own vantage point, Robinson says he is not a scientist but a writer, and that what he knows is what the execution environment for safety work looks like. He says the operation, and he believes the industry, still runs like a start-up, closer to that end of the spectrum than makes sense for systems he considers really dangerous.

Robinson names loss of control as one example of the risk he has in mind, and frames the scale in stark terms: if that happened, he says, the harm would be much larger than a single nuclear power station melting down. He adds that internal controls, safety measures, and redundancies are nowhere near what the world expects for a nuclear power facility. He also acknowledges that some of this is already public. He points to OpenAI having publicly reported safety problems, names Hugging Face, and references more recent disclosures. He notes that Anthropic has also reported incidents, including a case in which its safeguards were accidentally misconfigured. His closing observation is about perception: people outside may already have evidence that things are not as they ought to be, but an outside observer might also imagine the safety setup is more robust than it actually is.

Marcus then turns to policy, describing what he calls the Trump administration's anemic request for more disclosure as insufficient. He compares it to asking criminals to file monthly reports on which crimes they have committed. He says every day without stronger action is a mistake that invites worse problems, and that the point at which a temporary recall should have been imposed on an obviously dangerous technology has long passed. He warns that when something truly bad happens, the White House, not only the tech companies, will own it.

Read against the usual shape of this debate, the post sits at the strict end of the spectrum. Most proposals in circulation focus on disclosure, evaluation, red-teaming, or staged deployment for agentic systems that can browse the web, run code, and take multi-step actions. Marcus is asking for something stronger: withdrawal from the market, which is a recall in the consumer-product sense. That framing carries an implicit claim about who is accountable for harm, because recalls are typically ordered or urged by a regulator and executed by a manufacturer. The post does not name an authority that could issue one, and it does not describe a process for reinstating products once fixed. A reader can reasonably infer that he sees voluntary industry restraint as insufficient, but the mechanism is left open.

The technical background here is that an agent with internet access and tool use has a larger action space than a chat model. It can send requests, fill forms, spend money, modify files, and chain actions across services, so failures are not confined to bad text. That is the property that makes the recall argument coherent as an argument, whatever one thinks of the remedy. It also explains why benchmarks and guardrail configurations matter more for these systems than for a plain chatbot: a misconfigured safeguard, like the Anthropic instance Robinson cites, can change behavior at scale rather than in a single conversation.

Why it matters: for teams shipping or depending on browsing agents, the practical takeaway is that public scrutiny of agent-caused incidents is intensifying, and the loudest voices are no longer asking only for better evaluations. If calls for market withdrawal gain traction with regulators, vendors could face disclosure duties, deployment pauses, or usage restrictions on autonomous web access, and product roadmaps built on unattended agents would need fallbacks. Treat this as a signal about political and reputational risk rather than a decided policy outcome, since no recall has been announced here.

Read at Gary Marcus

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Every minute we wait risks disaster