AI news for builders and product teamsUpdated Oct 11, 2026, 00:01 UTC
Developer tools news
The latest Developer tools stories across our sources, prepared from the publishers’ own reporting.
In this topic
Newest first
Mistral AI has released Workflows in public preview, an orchestration layer for enterprise AI built on Temporal's durable execution engine. It offers durable execution, observability, human-in-the-loop approvals, and role-based access control within Mistral's Studio, with a Python SDK v3.0 available.

Mistral AI published an engineering blog post describing Spaces, its internal CLI for scaffolding projects, running dev environments, and deploying to staging. The post explains how the tool was adapted for use by coding agents, through flags, explicit state, introspectable plugins, and generated context.json and AGENTS.md files.

Ollama has released a preview build powered by Apple's MLX framework on Apple Silicon, citing large speedups on all Apple Silicon devices and support for NVIDIA's NVFP4 format. The preview, Ollama 0.19, accelerates the Qwen3.5-35B-A3B model and requires a Mac with more than 32GB of unified memory.

Ollama 0.17 can install and configure OpenClaw, a personal AI assistant, with the single command ollama launch openclaw --model kimi-k2.5:cloud. Setup requires Ollama 0.17 or later, Node.js, and a Mac or Linux system, with Windows via WSL.

Ollama announced support for subagents and built-in web search in Claude Code, requiring no MCP servers or API keys. Subagents run tasks in parallel, and web search is integrated into Ollama's Anthropic compatibility layer, working with any model on Ollama's cloud.

OpenClaw is a personal AI assistant that links messaging platforms to local AI coding agents through a centralized gateway running on the user's own devices. Ollama published installation and launch instructions, including an ollama launch openclaw command and a recommended context length of at least 64k tokens.

Mistral AI released Mistral Vibe 2.0, an upgrade to its terminal-native coding agent, powered by the Devstral 2 model family. It adds custom subagents, multi-choice clarifications, slash-command skills, unified agent modes, and automatic updates, and is available on Le Chat Pro and Team plans.

Ollama released a new command, ollama launch, that sets up and runs coding tools such as Claude Code, OpenCode and Codex with local or cloud models without environment variables or config files. It requires Ollama v0.15+ and recommends at least 64000 tokens of context length.

Mistral AI engineers describe investigating a memory leak in vLLM that appeared during pre-production testing of disaggregated serving with Mistral Medium 3.1 and graph compilation, causing system memory to grow 400 MB per minute and eventually reach an out-of-memory state.

Ollama has added experimental image generation support on macOS, with Windows and Linux planned. The feature runs models like Alibaba's Z-Image Turbo and Black Forest Labs' FLUX.2 Klein locally via the ollama run command.

Ollama v0.14.0 and later are now compatible with the Anthropic Messages API, allowing tools like Claude Code to run with open-source models locally or via ollama.com. The release supports tool calling, streaming, vision, and other features.

Open models can now be used with OpenAI's Codex CLI through Ollama, according to an Ollama post dated January 15, 2026. Codex can read, modify, and execute code in the working directory using models such as gpt-oss:20b, gpt-oss:120b, or other open-weight alternatives.

A first-person account describes using Claude Code, powered by Opus 4.5, to autonomously build and deploy a working website selling a 500-prompt set over about an hour and fourteen minutes. It attributes recent AI coding gains to more autonomous self-correction plus an agentic tool harness.

Ollama released a web search API on September 24, 2025, alongside a web fetch API. A free tier is available for individuals, with higher rate limits via Ollama's cloud, plus REST support and Python and JavaScript library integrations.

Ollama released a new model scheduling system that measures exact memory requirements before running a model, aiming to cut out-of-memory crashes and improve GPU utilization, including on multi-GPU and mismatched GPU systems. It is enabled by default for models on Ollama's new engine.

Ollama announced on September 19, 2025 that cloud models are now in preview, allowing users to run larger models on datacenter-grade hardware. The cloud offering integrates with existing local tools and Ollama's OpenAI-compatible API, and Ollama says its cloud does not retain user data.

Ollama announced Secure Minions, a protocol built by Stanford's Hazy Research lab that encrypts communication between local Ollama models and frontier cloud models using NVIDIA H100 confidential computing, with under 1% added latency on long prompts.

Ollama added the ability to enable or disable model thinking, separating reasoning from output when enabled. DeepSeek R1 and Qwen 3 support the feature, available through CLI flags, interactive commands, a scripting option, and a new think parameter in the generate and chat APIs.

Chip Huyen outlines six common pitfalls in building generative AI applications, drawing on public case studies and personal experience, including using gen AI when unnecessary, confusing bad product with bad AI, starting too complex, over-indexing on early success, forgoing human evaluation, and crowdsourcing use cases.

Chip Huyen outlines the common architecture of generative AI platforms, starting from a minimal query-to-model setup and progressively adding context construction, RAG retrieval approaches, guardrails for input and output, query rewriting, and agentic actions. The post describes components and considerations, not model evaluation, prompt engineering, finetuning, or RAG chunking strategies.