AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Why don’t machine learning research agents overfit?

Collected Oct 1, 2026

Research published by Amazon Science, titled "What fits (into few tokens) doesn't overfit: Compression and generalization in ML research agents," offers an explanation for why machine learning research agents do not overfit benchmark data despite repeated evaluation on the same held-out sets. Textbook accounts predict that iterating against a reusable holdout should produce rampant overfitting; studies building fresh test sets for old benchmarks have found improvements largely transfer instead.

The work tests the hypothesis that successful strategies are highly compressible. An explorer agent iterates freely against a validation set. A compressor agent then distills the winning strategy into a short prompt, which is handed to a reproducer agent that implements the strategy from scratch using only the prompt and training data, with no access to the validation set, code, or transcript. The study reports the compressor and reproducer are both Claude models. The researchers call a successful match a certificate of output compression.

Across eight datasets spanning tabular classification, image classification, language modeling, diffusion modeling, and reward modeling, 32-token prompts let a fresh reproducer match the explorer's adaptively optimized models on the large majority of problems. One language-modeling strategy survived compression to 16 tokens without loss in held-out performance. An example 16-token prompt, QKn 12L768 Mu .1 R² b2M 4x, was decoded as QK normalization, a 12-layer 768-dimensional transformer, the Muon optimizer at learning rate 0.1, squared-ReLU activations, a two-million-token batch, and a fourfold feed-forward block. At eight tokens, 12L768 Mu .1 R², the reproducer no longer matched, a boundary the researchers say shows compressed tokens carry information learned from data.

In a reverse experiment, returning one bit of feedback per query, whether a model beat the running best, produced strategies as good as full numerical scores, with a mathematical generalization guarantee.

Deliberately overfit agents, prompted to maximize validation performance at any cost, exceeded true held-out accuracy by more than 10% in 38 of 102 runs; those validation-specific gains vanished through the compression bottleneck, separating legitimate from overfitting strategies with high accuracy.

The researchers note the framework assumes the prompt is the only path from validation data to the final model; resolving possible pretraining memorization would require fresh datasets collected after a model's training cutoff, which they have not done. They state the results concern LLM agents and are suggestive for human research communities.

Read at Amazon Science

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

New research indicates that AI agents learn compressible models of data, which don’t have enough space to enable memorization.