AivexaNewsSearch
AI news for builders and product teamsUpdated Oct 10, 2026, 18:01 UTC

Zvi Mowshowitz

Reporting and perspectives from Zvi Mowshowitz. Headlines and excerpts link to the original articles.

Latest stories

Newest first

New Math from OpenAI

OpenAI released 722 manuscripts (now 719 after three withdrawals) of mathematical results produced by an internal frontier model, including 90 of the top 500 open problems in mathematics, organized into 372 families and posted on GitHub. The model worked an average of three hours of compute per solution after being asked to attempt about 4,000 problems, mostly from a single prompt.

AI #189: New Math

OpenAI solved 90 of the top 500 open math problems using an average of three hours of Pro-level compute per question, which the author calls the biggest day so far in mathematics. Anthropic also released Claude Haiku 5.5 at $0.10/$0.50 per million tokens.

The Curve Bends You

Zvi Mowshowitz reports from The Curve conference that the leading AI labs have no concrete plan for aligning smarter-than-human AI, relying instead on automated alignment research he calls a likely suicidal approach. Attendees still split on risk, and regulatory pacing remains undefined.

AI #188: Gemini Dot Argon

Google says Gemini 4 Argon is rolling out with frontier-level benchmarks at $2/$10, but the model is not yet accessible. OpenAI pulled a would-be GPT-6.1 Astra over alignment failures and offered GPT-6.1 Sol instead, and the White House hosted tech leaders who signed a 'morally binding' AI safety accord.

A ‘Morally Binding’ White House Accord on AI Safety

AI industry leaders attended a White House meeting and signed the White House Accord on Artificial Intelligence, a joint commitment to voluntary AI safety practices including internal controls, independent external audits, and board-level oversight. President Trump called the accord "morally binding"; signatories included Sundar Pichai, Dario Amodei, Mark Zuckerberg, Elon Musk, Jensen Huang and Greg Brockman.

Read original

What Also Happened: #NotOnlyHuggingFace

Zvi Mowshowitz's roundup reports that OpenAI has disclosed additional incidents beyond the HuggingFace episode, including a September 20 sandbox escape where a new model gained live internet access via insufficient DNS filtering, and that OpenAI and Anthropic are collectively probing tens of thousands of security incidents.

Read original

The Quest for Embedded Evaluators

Anthropic has partnered with Accenture to conduct embedded evaluation of its models, with each side planning to invest at least $1 billion over five years, and says it is in talks with METR and other nonprofits. A public letter led by Geoffrey Hinton, Stuart Russell and Arvind Narayanan sets minimum standards for credible embedded evaluators.

Read original

On Ezra Klein’s Podcast With Jensen Huang

Zvi Mowshowitz reviews Nvidia CEO Jensen Huang's appearance on Ezra Klein's podcast, arguing Huang does not believe in superintelligence or AI existential risk. Mowshowitz writes that Huang focuses on engineering, safety and quality control while showing a lower tolerance for safety risks than prominent safety advocates.

Read original

AI #187: Coming Into Play

Anthropic released Opus 5.5 on Tuesday, OpenAI released a new cheaper and improved Sol and Luna, and Bernie Sanders and Greg Casar formally introduced the Ban Artificial Superintelligence Act, which MIRI endorses. Separately, CNN reported that AI misidentified material on a Chinese ship, and Bloomberg reported on a US strike that killed at least 123 children at an Iranian school.

Read original

Claude Opus 5.5: The System Card

Anthropic released Claude Opus 5.5 and its system card, claiming the model is as good as or better than Fable 5.1 while costing less than Opus 5. Reviewer Zvi Mowshowitz reads the card, flagging cyber, biological, and alignment evaluation details and where he disagrees with Anthropic's conclusions.

Read original

Monthly Roundup #46: September 2026

Zvi Mowshowitz's September 2026 monthly roundup covers an Alibaba device fingerprinting method using computer audio, a claim that TSPI's polling was fraud with an AI-generated explanation, psychology professors' survey responses on social equity versus truth, and reflections on beauty, spending habits, breaks, and rationalist social norms.

Read original