AivexaNewsSearch
AI news for builders and product teamsChecked every hour

How Much Memory Does Your Agent Actually Need?

Collected Oct 1, 2026

A Hugging Face blog post describes experiments with ALTK-Evolve, a method that lets an agent learn from its own past trajectories by distilling reusable guidelines and injecting them back at inference time, with no weight updates and no human annotation. The post reports that the benefit depends on how much memory is delivered, describing the amount as a dose to calibrate to the model rather than a feature to switch on.

The evaluation covered eight models, from a 30B dense model to frontier proprietary systems, on AppWorld, described as 585 multi-step tasks (168 test_normal and 417 test_challenge) across 9 simulated apps. Tasks were scored by Task Goal Completion (TGC) and the stricter Scenario Goal Completion (SGC). Guidelines were mined once from AppWorld's training split only, and the post states no test-split data was used to build the set.

The post reports three patterns. Strong models with headroom want the full guideline set; DeepSeek-V3.2 (671B MoE) gained +9.5 percentage points in TGC. Smaller or weaker models do best with a compact core plus per-task retrieval; gpt-oss-120b (117B MoE) gained +16.1pp TGC at only +5% tokens, while the full set gained less and cost about 50% more tokens. Already-saturated models showed no measurable gain; GLM-5 (745B MoE) was placed in this pattern, which the post says describes what was observed, not a proven cause.

The post says parameter count alone does not determine the pattern, listing benchmark headroom, context-window size, architecture, guideline quality and task distribution as factors it says remain under study. It also reports that GPT-5.5 and Opus, near the TGC ceiling, gained +7.2 and +7.1pp SGC respectively, and that DeepSeek ran roughly the same number of ReAct steps with and without memory (about 18-19 on average). It notes prompt caching as a production efficiency lever and says controlled experiments isolating context-window size have not yet been run.

Read at Hugging Face Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt