AI news for builders and product teamsUpdated Oct 10, 2026, 23:01 UTC
Developer tools news
The latest Developer tools stories across our sources, prepared from the publishers’ own reporting.
In this topic
Newest first
Factory used self-hosted LangSmith to export traces to AWS CloudWatch, link feedback to individual LLM calls, and automate prompt optimization, reporting a 2x improvement in iteration speed. Factory also reports about a 20% reduction in open-to-merge time and 3x less code churn on Droid-impacted code in the first 90 days.

LangSmith LLM Gateway is now in public beta for Plus and Enterprise plans, offering centralized runtime controls for production agents including spend caps, rate limits, model fallbacks, and PII redaction. It supports multiple model providers and is BYOK-first, with PII redaction limited to Enterprise users.

LangChain announced a self-improving evaluator feature in LangSmith that stores human corrections to LLM-as-a-Judge outputs as few-shot examples, which are then fed back into future prompts to align evaluations with human preferences without manual prompt engineering.

LangChain released a stable LangGraph v0.1 framework for building agentic and multi-agent applications and announced LangGraph Cloud, infrastructure for scalable, fault-tolerant agent deployment, now in closed beta. A note says LangGraph Platform was renamed LangSmith Deployment as of October 2025.

Goodfire, a San Francisco AI lab founded in 2024, has made its Silico mechanistic interpretability platform generally available to the public and announced a grant program offering US $1 million in free Silico usage for academic and nonprofit interpretability researchers.

Sentence Transformers v6.0 introduces a fourth model type, MultiVectorEncoder, for ColBERT-style late interaction retrieval, along with a full training approach. A Hugging Face blog post documents finetuning multi-vector models, including a medical retrieval model trained in 14.5 hours on one RTX 3090.

Anthropic has added remote control updates to Claude Code, letting users start a new session from their phone, with dropped laptop-to-phone connections recovering automatically and faster session loading on iOS. The newsletter also notes Claude Academy, free Anthropic courses and guides for using Claude.

Gradio introduced gr.Workflow, a tool that turns AI pipelines into a drag-and-drop graph of typed nodes where each step is runnable and every intermediate result is visible. Each workflow also serves as a REST API and can be deployed to Hugging Face Spaces with one command.

Ollama announced that Claude Desktop can now be configured to work with Ollama as a third-party gateway provider, allowing developers to run open models inside Claude. The integration supports both local and cloud models, with telemetry disabled by default and a Zero Data Retention policy.

Hugging Face detailed how Papers with Code uses Hugging Face Jobs, Storage Buckets, and Inference Endpoints to build a hybrid search system combining PostgreSQL full-text search with pgvector semantic search, covering more than 110,000 papers from arXiv and Daily Papers.

Sentence Transformers v6.0 adds a fourth model type, MultiVectorEncoder, for ColBERT-style late interaction retrieval, per a Hugging Face blog post. The post says PyLate and Stanford-NLP ColBERT checkpoints load directly, colpali-engine visual document retrieval models can be used after per-repository configuration, and it describes MaxSim scoring plus index-size tradeoffs.

Together AI describes running A/B experiments at the endpoint level, splitting live traffic among one control and up to 20 variants with fixed percentages. It covers ramping via member updates, measuring platform and product metrics per deployment, and ending tests by promoting a rollout or deleting the experiment.

A tutorial project describes building a simple AI text detector by fine-tuning a DistilBERT classifier that returns a 0-100 score, then using it as a verifier to train a small language model to produce text that avoids detection.

Hugging Face published a walkthrough of a streaming data loop using Strands Robots, LeRobot, and Hugging Face Storage Buckets, covering recording robot demonstrations, training by streaming datasets directly from the Hub, and deploying checkpoints back to hardware.

Amazon Science released Turnstile, a Rust proxy that sits between an agent harness and a model backend to record exact token IDs, log probabilities, loss masks, and weight-version boundaries during generation for reinforcement learning. Validations included a text-only coding agent and a multimodal computer-use agent, whose harnesses were left unchanged.

Ollama 0.31 makes Gemma 4 significantly faster on Apple Silicon using multi-token prediction powered by MLX, reporting up to 90% faster token generation on the Aider polyglot benchmark. The speedup is enabled by default and does not change the model's output.

Ollama updated its MLX engine for Apple Silicon, adding support for NVIDIA's NVFP4 quantization format and optimizations it says deliver up to 20% faster output, higher-quality responses, and lower memory use. The release also introduces a snapshot system for agent workloads.

Ollama 0.30 adds GGUF model compatibility through llama.cpp, up to 20% faster performance on NVIDIA hardware, and Vulkan enabled by default for broader GPU support on AMD and Intel devices.

OpenJarvis v1.0, an open-source framework for building personal AI agents that run on local hardware, is now available with built-in support for Ollama. It is built by Stanford's Hazy Research and Scaling Intelligence labs as part of their Intelligence Per Watt research into efficient local AI.

Mistral AI released Connectors in Studio, available in Public Preview, letting developers use built-in and custom MCP connectors via API/SDK across model and agent calls. The release adds direct tool calling and human-in-the-loop approval flows.