AivexaNewsSearch
AI news for builders and product teamsChecked every hour

[AINews] TypeSafe/Jev at >$100M ARR, $7.5B valuation 3 weeks after launch

Collected Oct 10, 2026

TypeSafe's Jev, the decision-model API that returns typed answers rather than free text, has reportedly crossed $100M in ARR and reached a $7.5B valuation about three weeks after launch. Latent Space's AINews recap relays that TypeSafe announced its "Series AI" round, and that Sequoia "leaked" the ARR figure in the first week. The same recap notes widespread cloning of the Jev API, criticism and accusations of astroturfing around the launch, and a quoted post from a16z observing that it has been three weeks since launch. The news lands amid a burst of competing "decision model" releases that treat Jev as the reference point.

Decision models are a new product category that returns typed answers in a single forward pass: probabilities, a pick from a list, or a score against levels. OpenAI shipped a Decisions API with three request types covering probability of a condition being true, list selection, and level scoring. It accepts text and images, runs on GPT-6 Luna, costs $0.10 per million input tokens with no output charge, and OpenAI describes it as up to 10x faster. Microsoft released Decision-1, aimed at LLM judges and screening scientific hypotheses; one early evaluator found decision models still struggle with consistency and complex decisions. Perplexity's pplx-decider-v1.1-27b claims top Decision Bench accuracy at 94.5% across 1,071 cases at $0.017 per 1K decisions. Cloudflare's clef-omni accepts audio, video, image and text, with clef-flash cheaper than Jev and clef overall about 2x faster; weights are on Hugging Face. Liquid's d1 is now on Vercel AI Gateway with vision support for classify, route and score tasks.

Serving and routing layers are adapting quickly. vLLM's Semantic Router Decision 2.0 answers multiple questions about one input in one pass with per-option probabilities, and LangSmith uses Jev as a judge returning separate typed answers for difficulty and correctness on every trace. Unsloth released a free notebook that turns Qwen3.5-4B into a decision model on 8GB of VRAM, and a walkthrough on Qwen3.5-0.8B reports accuracy rising from 37% to 65% in 60 steps, about 10 minutes on 4GB. The category matters for harnesses because many agent steps are yes/no calls rather than generation: LangChain says routing each task to the cheapest adequate model cut median Open SWE cost per task by 64%. Related research from Apple and CMU, Selection-based Structured Reasoning, applies the same idea inside agents. That method scores six natural-language strategies by length-normalized log-likelihood in one batched forward pass sharing the KV cache, cutting per-turn reasoning latency by more than 90% while end-to-end latency per question falls only 28-54%. On Qwen3-VL-4B with GRPO, average success was 61.37% versus 61.25% for a TAPO+GSPO baseline.

Multi-agent tooling also moved. Claude's Managed Agents dynamic workflows entered public beta: a lead agent writes a phased plan, fans out to up to 1,000 agents per run, then merges results, enabled with multiagent_20261001. Anthropic advises starting with scoped tasks because token use can be high. Claude Code Projects admitted all waitlisted Pro and Max users, runs tasks as parallel threads, and sessions can now run locally. Opus 5.5 fast mode rolled out but bills against usage credits outside subscriptions. Vals AI benchmarked GPT-6 Sol and Opus 5.5 on Vibe Code Bench alone and as teams: teams cost 1.8-5.1x more, and only Sol at medium effort improved significantly, by 7.3 points. Sol delegated in parallel along architectural lines, while Opus ran sequential waves reaching about 6.8 subagents and roughly 1,140 subagent tool calls per app at max effort with no significant gain. Prime Agent rewrote itself in Rust over two weeks using more than 2,000 agents, 10K+ sandboxes and 200B+ GLM-5.3 tokens, reaching usable input about 13x faster with 83% less startup memory. Codex added a Windows sandbox built on Microsoft Execution Containers, and Composer next-message predictions in beta for Pro users only; some users criticized that gating, and there were complaints of daylong outages. DHH said GPT-6.1 Sol made Codex his primary tool over Claude. Devins can now spawn trees of managed Devins so wall time tracks the slowest branch rather than the sum, and accept personal ChatGPT plans; Grok Bot got its own email address for sign-ups and scheduling.

Model releases were dense. Qwen-Image-2.1-Turbo is an open-weights accelerated checkpoint of the 7B Qwen-Image-2.1 doing 8-step 2K generation and natural-language editing, loading through Diffusers QwenImage21Pipeline alongside Pro and Turbo APIs. StepFun's Step 5 Preview is a 600B-total, 27B-active sparse MoE with 1M context and vision, scoring 33.89 on the Hermes Index, matching GPT-6 Luna, free on Nous Portal for a week, number one on OpenRouter Trending, with open weights due October 15 and max output corrected to 64K tokens. Upstage Solar Mini 4 is a 35B MoE with 3B active, 524K context and 208 tok/s, with an AAII score of 24 that is best at 3B active and within a point of Nemotron 3 Ultra, free in Cline. Gemini 4 Argon was reported at 77.9% on DeepSWE v1.1 versus Opus 5.5's 74.2%, shipping first to 650+ Fairwind Program defenders at $2/$10 per million tokens; reasoning-effort selectors appeared in Antigravity and Logan Kilpatrick said Argon is coming. Business Insider reported an unconfirmed internal Carbon checkpoint approaching Opus 5.5 on coding. HeyGen Voice tops the Artificial Analysis Controlled Voice TTS arena at 1,201 Elo, $30 per 1M characters and 40 chars/s. Whistle is a 16.9MB on-device STT model said to rival Whisper base. In multi-turn image editing, Ideogram 4.5 and FLUX 3 edit locally leaving 95%+ of the image untouched on small edits, while GPT Image 2.5 Sunburst re-renders most of the frame each turn, keeping about 20% unchanged and drifting, and Nano Banana 2.1 gradually darkens. Roboflow's OCR benchmark covers 48 models with GPT-6 Astra leading text localization; Datalab's OmniParseBench has 16K tests across 90 languages and its own model does not rank first. Claude Haiku 5.5 ranks #30 on WebDev at $0.10/$0.50, matching GPT-6 Luna's price while scoring 6 points higher; Mistral Large 4 sits at #43 on Agent Arena; ARC-AGI-3 set a new high score of 59.17%.

Infrastructure and training updates: vLLM reports more than 7.8x GB200 throughput on MiniMax M3 at matched interactivity on AgentX, described as early results, using locality-aware MoE with CUDA 13.4 locality domains worth up to 1.2x faster MoE decode. SGLang reports up to 20% faster FP8 MLA at 128K context and a 5.9% end-to-end gain from MoE tail fusion removing 276 launches per decode step. A preview InferenceX submission cited by SemiAnalysis shows 3.2x profit per gigawatt and up to 10x performance per dollar versus GB300. TRL v1.15 enables the fused LM head by default, avoiding materializing the full logits tensor: on Gemma 3 1B, GRPO sequence length rises from 28K to 114K and DPO from 10K to 59K, peak memory at 8K falls 52-82%, and training is up to about 11% faster. Datology Curation Studio claims a 6x compute multiplier on 39 open datasets for a 30B MoE and cites Thomson-1, trained for $450K, beating GPT-5.6 Sol head-to-head. Tinker cut prices up to 70%, priced long context the same as short, and added GLM-5.3-Flash and DeepSeek-v4.1-Flash. ByteDance Seed found DeepSeek retrieval depends on where a token lands relative to the compression stride, persisting without RoPE or learned gates and tracking stride length, with V4.1's stride of 2 reducing but not eliminating the effect. Meta's agent plasticity work measures held-out gain per learning dollar and finds the best performers are not the most efficient learners; MIMESIS is a 9B user simulator.

Why it matters: TypeSafe has shown that a narrow, opinionated API format can become a category and a fast-growing business, which changes expectations for founders building developer tools and for API teams deciding whether to expose structured outputs. For developers, the practical shift is that routing, judging and classification steps can be billed and benchmarked like generation, so harnesses can cut costs by sending cheap yes/no calls to specialized models rather than a frontier model. The open question is durability: Jev is the benchmark everyone targets, but OpenAI, Microsoft, Perplexity, Cloudflare and open-weight options are all shipping comparable endpoints, so pricing and accuracy leadership, not format novelty, will likely decide who keeps the category.

Read at Latent Space

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Wow.