How Snyk Turned an Internal Support Agent into a Customer Feature
Snyk Assist began as internal tooling for Snyk's own customer support team and, after roughly a year and two intermediate stages, moved into the core Snyk product as a panel in the top bar of every page for every paying customer on 1 September 2026. Snyk is an AI security platform that finds and fixes vulnerabilities in code, open source dependencies, containers and cloud configuration, and works inside the IDE, pull request and pipeline. The agent itself is conversational: customers ask questions in plain language and it can look up open issues, check a package for known vulnerabilities, open a support case from the conversation, or log a feature request when the workflow someone needs is not yet supported.
The motivation was volume. Snyk's support team handles thousands of cases every couple of weeks, and each one had to be read, classified by product, severity and owner, and routed onward. Answers to customer questions were spread across product documentation, support articles, release notes, learning content and account data, so finding the right page often meant reading several documents or filing a ticket and waiting. Because Snyk sells security software, the team set a high bar before putting an agent in front of paying users: it had to answer correctly, refuse what it should refuse, and never return anything the signed-in user was not already allowed to see.
The rollout happened in three phases. In phase one the agent ran for about a year as an internal tool for the support team, functioning both as case triage and as a virtual agent, with Snyk staff as the only users. That period surfaced edge cases that did not appear in offline testing, and feedback arrived in hours rather than release cycles. In April 2026 it reached customers in the support portal, where it could greet users by name, understand account context, run a health check on login, and raise a ticket at any time of day. The move on 1 September 2026 into the core product coincided with new navigation. Notably, the agent runtime stayed the same across all three phases; only the surface in front of it changed.
Architecturally, a single LangGraph agent sits behind several front doors: a Slack app, a web app and direct API access, all routed through the same secured framework. Snyk says it chose LangChain for its community, ecosystem and position as a standard for building agents, and deliberately kept one runtime so a small team could fix bugs, ship features and respond to incidents from one place. Four design decisions are described as making this practical in production.
First, tools are typed functions registered per user and attached at request time based on the signed-in user's actual permissions, so the agent can only reach data the user could already see. Most tools are separate microservices, which lets them scale independently and gain capability without changing the agent loop. Second, middleware handles context management, guardrails and model fallbacks at defined points in the agent lifecycle, so adding or reordering behaviour is a one-line change to a list rather than a rewrite. Third, conversation state is persisted by a checkpointer in PostgreSQL and keyed by session, which gives multi-turn memory, scaling across pods and conversation resumption without extra wiring. Fourth, the agent loop is treated as an internal platform: because an agent is just a model, tools, a prompt, middleware and a checkpointer, a shared factory returns the compiled graph and each team supplies its own configuration, so most new workflows are a config change rather than a new architecture.
Evaluation is the part of the story with the most transferable lessons. Every model call, tool call and decision point has been traced in LangSmith since the first day of development, and those traces became the basis for shipping decisions. Two layers of evaluation are in use. Offline evaluation runs test question sets: one made of real questions with known good answers, graded by a second model to catch quality regressions from prompt or model changes, and one automated red-teaming exercise where people try to trick the agent into ignoring instructions or handing over secrets. Continuous integration acts as a gate, running the real agent against those suites on every pull request and blocking on thresholds committed to the repository. Online evaluation grades every production run through a scheduled job, checking whether the question was about Snyk and whether the response answered it. Those grades produce a measured deflection rate rather than an estimate, and each question is categorised by product area, topic, language ecosystem and error type, giving product managers a live picture of where customers struggle. Bad traces can be investigated and converted into datasets from inside the IDE through the LangSmith MCP server.
One quote in the post comes from Bailey Millns, an AI engineer at Snyk, who says the hard part of building agents is not what the model can do but proving the agent actually works, and that changes which do not clear the evaluation bar do not ship. Matt Jarvis, Snyk's director of AI engineering, frames LangChain as a batteries-included abstraction that let the team focus on capability instead of rebuilding orchestration, tool-calling, streaming and state. Jada Ross, also an AI engineer, cites the size of the community and ecosystem as the reason for the framework choice.
The reported results since the April 2026 customer launch are more than 60,000 queries handled across more than 500 customer accounts, with over 85% of sessions resolved without a support ticket, which Snyk says saved the support team hundreds of hours, and more than 250 cases auto-detected by the agent and escalated straight to the right team. The post also points to Snyk's documentation for Snyk Assist.
Why it matters: this is a concrete example, with numbers, of the internal-tool-to-customer-feature path that many platform teams are now considering, and of permission-aware tool registration as a way to keep an agent inside a user's existing access. Teams building similar agents should note that the reported cost of the approach is not the model but the evaluation harness, since traces, CI gates and online grading are what let this team ship continuously. The main open questions are whether deflection and grading quality hold as surfaces multiply, and how much of the one-runtime, middleware-driven design depends on the specific framework chosen. That second point is inference on my part, not a claim Snyk makes.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Discover how Snyk transformed an internal support agent into Snyk Assist, a customer-facing AI feature powered by LangChain, LangGraph, and LangSmith.