Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence
General Science
Releases, research, and ideas for developers and product teams. Updated every hour.
General Science
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. We face a new epoch in computing. Hardware is changing rapidly — not just faster GPUs, but a growing range of chips from different vendors, each with its own architecture and often tailored to specific AI workloads. Software is changing just as fast, and AI coding tools now generate in minutes what took months of effort a few years
Epoch and METR released MirrorCode, a benchmark testing whether AI systems can reimplement programs from CLI access alone; some tasks were solved, but 8 of 25 targets were never fully solved. Anthropic, Sunday, and OpenAI also reported robotics and safety findings.
Overview of ABBEL compared to traditional recursive summarization. Beliefs replace the full interaction history as the agent’s working context, and belief grading improves performance by supervising the contents of each belief state.. As task horizons grow, LLM contexts can’t scale forever. Self-summarization enables concise, interpretable contexts, but at a significant performance cost, especially for human assistance domains where high quality data is scarce, e.g., collaborative code generation. We address this w
General Science
Machine Intelligence
Google commits $40M in AI tokens and credits for the Genesis Mission
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
The UK AI Security Institute found the cybersecurity capability gap between open-weight and closed frontier models has narrowed, with GLM-5.2 and DeepSeek V4-Pro performing like closed models released four to seven months earlier. Kimi also announced Kimi K3, a 2.8 trillion parameter model, while Demis Hassabis proposed a FINRA-style standards body for frontier AI testing.
An article from Ahead of AI explains how large language models can be trained to support multiple reasoning-effort modes, citing OpenAI's GPT-5.6 family, gpt-oss, Qwen3, and Thinking Machine Labs' Inkling as examples.
Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.
Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience and AI models.
Algorithms & Theory
Google and AIM launched ATL Saathi, a Gemini-powered AI tool empowering Indian educators in robotics labs.
Generative AI
Algorithms & Theory
... government of the people, by the people, for the people ... — Abraham Lincoln, Gettysburg Address (1863) The cost of AI is dropping rapidly. GPT-4-class capabilities cost roughly $30 per million tokens in early 2023; today the same runs under $1 , and some providers are pushing costs below $0.10 . Across benchmarks, inference prices have fallen between 9x and 900x per year , with a median decline near 50x. Even frontier models are getting dramatically cheaper each generation, with open-source models follo