Building A Generative AI Platform
Chip Huyen published an overview of the architecture companies commonly use to deploy generative AI applications, describing shared platform components and how they are implemented. The post states that the architecture is kept general, and that certain applications may deviate from it; components can be skipped if a system works well without them.
The starting point described is a simple flow in which an application receives a query, sends it to a model, and returns the generated response, with no guardrails, augmented context, or optimization. The "Model API" component covers both third-party APIs (for example OpenAI, Google, Anthropic) and self-hosted APIs. From there, components are added as needs arise, and evaluation is described as necessary at every step. The post also states it does not cover model evaluation, application evaluation, prompt engineering, finetuning, data annotation guidelines, or RAG chunking strategies, which it says are covered in an upcoming book, AI Engineering.
The first expansion described is context construction: giving a model access to external data sources and tools. The post cites Lewis et al., 2020, for the finding that relevant context can help models generate more detailed responses while reducing hallucinations. It describes RAG as a generator plus a retriever, and notes that external documents typically need to be split into chunks determined by the model's maximum context length and application latency requirements. Two retrieval approaches are outlined: term-based retrieval, including keyword search, BM25, and Elasticsearch, and embedding-based retrieval using embedding models such as BERT, sentence-transformers, and proprietary models from OpenAI or Google. Vector search is described as nearest-neighbor search using approximate nearest neighbor algorithms including FAISS, Google's ScaNN, Spotify's ANNOY, and hnswlib. Production systems typically combine approaches; combining term-based and embedding-based retrieval is called hybrid search, with sequential reranking and ensemble patterns described.
For structured data, the post describes text-to-SQL, SQL execution, and generation. Agentic RAGs can use web search tools such as the Google or Bing API. Read-only actions retrieve information without changing state, while write actions change state and pose more risks.
Guardrails are divided into input and output guardrails. Input guardrails address leaking private information to external APIs and model jailbreaking; the post cites Samsung employees leaking proprietary information into ChatGPT, which it says led Samsung to ban ChatGPT in May 2023. Output guardrails evaluate generation quality and specify policies for failure modes including empty responses, malformatted responses, toxic responses, hallucinations, and responses containing sensitive information. Hallucination detection is described as an active research area with solutions such as SelfCheckGPT (Manakul et al., 2023) and SAFE (Wei et al., 2024).
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
After studying how companies deploy generative AI applications, I noticed many similarities in their platforms. This post outlines the common components of a generative AI platform, what they do, and how they are implemented. I try my best to keep the architecture general, but certain applications might deviate. This is what the overall architecture looks like. This is a pretty complex system. This post will start from the simplest architecture and progressively add more components. In its simplest form, your appli