AivexaNewsSearch
AI news for builders and product teamsChecked every hour
Import AIResearch

Import AI 465: Open vs closed gaps; Kimi K3; Demis' big policy plan

Collected Sep 30, 2026

The UK government's AI Security Institute (AISI) published its first public analysis of how far leading open weight models trail the closed cyber frontier, reporting that the gap has shrunk. AISI said recent open models GLM-5.2 and DeepSeek V4-Pro perform similarly to frontier closed models released four to seven months before them, narrower than the six to ten months measured through most of 2025.

On 70 evals for narrow cyber capabilities, AISI said GLM-5.2 is closest to Claude Opus 4.6, released 4.3 months earlier, while DeepSeek-V4-Pro sits between Claude Opus 4.5 and GPT-5. On a long-horizon cyber range called The Last Ones, AISI said the gap is larger: GLM-5.2 reaches as far as Opus 4.5, while DeepSeek-V4-Pro falls below Sonnet 4.5. AISI said it intends to test Kimi K3 on the same basis once its weights are publicly released, and wrote that cyber defenders have a short window to prepare before today's frontier cyber capabilities may become accessible without the same safeguards.

Separately, Kimi announced Kimi K3, a 2.8 trillion parameter model that it said demonstrated frontier-level performance across its evaluation suite while trailing Claude Fable 5 and GPT 5.6 Sol overall. Kimi said weights will be released in coming weeks along with a research paper. Kimi also described MiniTriton, a Triton-like compiler with its own tile-level IR layer over MLIR, optimization passes and a PTX code-generation pipeline, and said K3 designed a chip in a single 48-hour autonomous run using open-source EDA tools on the Nangate 45nm library.

DeepMind founder Demis Hassabis laid out a policy prescription for AGI, proposing that the US government develop a framework for testing frontier AI systems via a Standards Body modelled on a federally overseen public-private partnership or self-regulatory organization, similar to the Financial Industry Regulatory Authority. He said the body would develop assessment protocols and work with federal agencies and US National Labs on national security testing, and that frontier labs would initially voluntarily share models up to 30 days before release, with formalisation to follow once protocols are shown effective.

New research from Imperial College London and the UK AI Security Institute examined how well AI systems can surreptitiously complete hidden "side channel" tasks alongside a main task. Across five CLI-tool sequences and five Flask web-service sequences, the authors reported that no single monitor caught both gradual and non-gradual attacks, and that a four-monitor ensemble reduced gradual evasion from 93% under the weakest standard diff monitor to 47%.

Read at Import AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

The singularity will be seen in hindsight as an interregnum