How a global fintech scaled coding agent traffic with Dedicated Model Inference
Together AI published an account of how a global fintech ran its coding assistant on GLM 5.2 through Together's Dedicated Model Inference (DMI), a shift it describes as moving from managed capacity planning to self-serve endpoint control.
The customer ships financial products to millions of users across dozens of markets and uses AI coding agents in its engineering workflow. Its traffic is spiky, concentrated in engineering hours, and grows as more teams adopt agents. Together said the workload runs at ISL p50/p90/p95 of roughly 81K/163K/178K tokens and RPS p50/p90/p95 of 3/6/7, a peak-load, relatively low-TPS pattern where concurrency matters more than raw throughput.
Together said its earlier dedicated offering required the customer's platform team to file requests and size clusters, and that capacity planning stopped working as adoption became unpredictable. The customer asked for self-service provisioning, programmatic observability, sustained performance under concentrated load, and the ability to swap models without renegotiation.
On DMI, Together said the customer configures endpoints via API, UI, or CLI, including creation, sizing, scaling policy, and configuration changes. Together reported that after a migration reshaped the customer's GLM 5.2 endpoint toward fewer, larger replicas, requests queued one to three minutes and decode throughput fell to roughly 5 tokens per second; the fix was a same-day live configuration change restoring a tuned cache-session-aware routing policy and widening the max-inflight-per-worker threshold, with zero downtime.
Together said its API Support team used a metrics API to trace a single 192-second slow request, finding it queued behind a 2.3M-token pending-prefill backlog rather than its own 250K-token prompt.
The production endpoint ran GLM 5.2 at 256K context on 56 B200s (14 replicas by 4 B200s). Together said the customer moved from GLM 5.1 to GLM 5.2 on the same account and evaluated a 1M-context configuration but turned it down to preserve concurrency headroom, staying at 256K/512K. Together also said the customer's team is scoping a dedicated GLM 5.1 node in a new region for a chat-style customer-support NLP workload.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Inside a global bank's shift to self-serve dedicated inference: how Together's DMI gave engineering teams direct control over scaling, models, and testing.