How Postman runs Agent Mode for 40 million developers on Amazon Bedrock

Postman and AWS have published an engineering account of how Postman's Agent Mode, an AI-native way to work across API testing, documentation, discovery and implementation, runs in production on Amazon Bedrock for what the post describes as 40 million developers. Agent Mode operates directly against the Postman application: in the example given, it opens a pull request and proposes next steps without the user navigating the interface. Postman has evolved over 11 years, and the post frames the central difficulty as integrating an agent into a mature product whose APIs, user experience and product knowledge were all shaped around interface-driven assumptions. An agent, as the write-up puts it, reasons over data rather than navigating a screen.
The architecture has three parts that all resolve to one runtime action: an inference call to a foundation model. Client-side tools live in the Postman application and represent the final actions the agent can take, such as opening requests, modifying settings, running collections and inspecting authentication. Server-side tools cover functions such as web search and agent-loop management, though most tools operate on the Postman application. Generic agent instructions define system-level behavior, including how proactive Agent Mode should be, how it communicates uncertainty, and its baseline product knowledge. A third component is a Retrieval Augmented Generation knowledge base, seeded for the initial rollout from Postman's Learning Center as concise feature-specific articles and selected at runtime based on the incoming query and available context. Selecting a mock server, for instance, automatically injects the related article. Guardrails apply too: Agent Mode requires user approval before actions that modify application state, and Postman uses Amazon Bedrock Guardrails to redact personally identifiable information before it reaches the underlying large language model, a control enterprise admins can enable in Agent Mode's guardrail settings.
The most concrete engineering finding concerns tool sprawl. Early versions favored highly atomic tools such as opening a request, updating one field, or fetching a piece of metadata. That bought correctness and control but produced long tool-call sequences where every step had to return to the model before the next began, so workflows users mentally grouped as one operation felt slow. More importantly, Postman's testing found that tool-selection errors increased once the visible toolset exceeded approximately 40 tools; the agent could call nonexistent tools, pass incorrect arguments despite valid schemas, or pick tools that looked semantically reasonable but were wrong in context. Larger or newer models reduced this behavior without eliminating it. The current design has a root agent query a vector database of tool embeddings, narrow more than 170 tools to roughly 15 relevant to the request, and hand those to a context-isolated sub-agent so the model sees only what the task needs. In practice, this suggests the practical ceiling on agent tooling is a context-budget problem rather than a menu-size problem, and that retrieval over tool descriptions is a reasonable mitigation for catalogs that keep growing.
A subtler issue was coupling between client APIs and interface state. Tools that modified requests needed certain elements open, while other tools opened new tabs as side effects, forcing the agent to open a request tab just to read it. Postman is decoupling tools from tabs, and its Native Git feature uses the approach heavily; Agent Mode can now send requests in the background without an open tab, though approval is still required. For products such as the API Catalog, Postman collapsed multiple narrow views into a single query tool. Because the underlying ClickHouse tables have known schemas, the agent can generate complex queries with joins and WHERE clauses, which sharply reduces the number of distinct tools needed for an analysis question. The illustrative query aggregates total requests, error counts, error-rate percentage, average latency and p95 latency over a seven-day window, grouped by service, filtered to p95 under 100 milliseconds and non-zero traffic, and ordered by error rate. The engineering work thus shifts from building one tool per question to modeling the data well once. A likely trade-off is that schema-aware reads move complexity into data modeling and query correctness, which is a different and less familiar skillset for many product teams, but it scales better than enumerating read tools.
Postman also reports that its initial assumption was wrong: missing or incomplete context caused more failures than missing capabilities. Context here means the agent's understanding of where the user is, which entities are active, and what state has been established. Rather than serializing the existing interface data model, which was shaped for rendering and data transfer rather than reasoning, Postman built dedicated context handlers that distill each entity into what the agent needs. Two forms of context feed the agent: broad, shallow background context gathered automatically and minified for the prompt, and deep, focused selected context chosen by the user and routed through a per-entity-type handler. As more objects gained handlers, truncation became the next problem, because fields such as request descriptions, OpenAPI specifications and request payloads contain open-ended user data that can crowd the context window. Postman is exploring a filesystem-backed approach so each handler does not need custom truncation and expansion logic.
On Amazon Bedrock, four capabilities carry the most weight. Model flexibility lets Postman route each workload across supported Anthropic Claude models, using faster models for high-volume, latency-sensitive interactions and larger ones for complex reasoning, with switching treated mainly as a configuration change. Cross-Region inference lets requests route among destination Regions defined by an inference profile; at runtime the application passes the profile ID or ARN as the modelId in Converse or InvokeModel, and the profile, IAM and service control policies, and quotas must permit every destination Region Bedrock might choose. Geographic profiles keep routing within a defined geography such as the United States or the European Union, while global profiles can use supported Regions worldwide for extra burst throughput. Data residency matters for enterprise buyers: Postman has configured zero data retention with data_retention_mode set to none for supported Agent Mode models, and availability and behavior are model-dependent.
Prompt caching addresses cost and latency, since a production agent resends a large stable prefix each turn: system instructions, agent behavior, a core tool set, selected knowledge and conversation context. The near-immutable core uses a one-hour cache checkpoint and more variable context a five-minute checkpoint that refreshes on a hit; Bedrock requires the longer-lived checkpoint to appear first. Teams can verify behavior through the cacheReadInputTokens and cacheWriteInputTokens usage fields and measure time to first token.
Why it matters: Teams building production agents can borrow concrete patterns here, especially budgeting tools like tokens, giving agents schema-aware reads instead of many narrow tools, and caching stable prompt prefixes with tiered TTLs. Postman's reported roughly 40-tool error threshold and its narrowing from more than 170 tools to about 15 are useful reference points, though they come from one product's testing and should be treated as indicative rather than universal. Enterprise buyers, in inference, should also note that geographic inference profiles and model-dependent retention settings are the levers that determine where prompts are processed and whether they are stored.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Building an AI agent that works in a demo is a different problem from running one for 40 million developers. Postman and AWS share the architectural patterns behind Agent Mode: controlling tool sprawl, exposing schema-based reads, and treating context as the real bottleneck, plus how it runs on Amazon Bedrock at scale.