Common pitfalls when building generative AI applications
Chip Huyen published a note listing common pitfalls she has observed when building applications with foundation models, citing public case studies and personal experience.
The first pitfall is using generative AI when it is not needed. She describes a team that pitched using an LLM to schedule a household's energy-intensive activities against hourly electricity prices; their experiments showed it could cut a household's electricity bill by 30%. Huyen asked how it compared with simply scheduling such activities when electricity is cheapest, such as doing laundry and charging a car after 10pm. She said she suspects greedy scheduling can be quite effective, and that other cheaper, more reliable optimization solutions exist, such as linear programming. She also cites a big company wanting to use generative AI for network traffic anomaly detection, another wanting to predict customer call volume, and a hospital wanting to detect patient malnutrition.
The second pitfall is confusing a bad product with bad AI. In two teams she looked into, the issue was product, not AI. She gives an example of a meeting-transcript summarization app whose users did not care about the summary itself but only wanted action items specific to them. LinkedIn's skill-fit chatbot found users wanted helpful rather than merely correct responses, and Intuit's tax chatbot initially got lukewarm feedback because users hated typing until suggested questions were added, according to Nhung Ho, VP of AI at Intuit.
The remaining pitfalls are starting too complex, over-indexing on early success, forgoing human evaluation, and crowdsourcing use cases. Huyen cites LinkedIn taking one month to achieve 80% of the desired experience and an additional four months to surpass 95%, and says a startup building AI sales assistants for ecommerce told her getting from 0 to 80% took as long as from 80% to 90%. She cites Ding et al. (2023), UltraChat: the journey from 0 to 60 is easy, whereas progressing from 60 to 100 becomes exceedingly challenging. She recommends daily human evaluation of 30 to 1000 examples to correlate human and AI judgments, understand usage, and detect behavioral patterns.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
As we’re still in the early days of building applications with foundation models, it’s normal to make mistakes. This is a quick note with examples of some of the most common pitfalls that I’ve seen, both from public case studies and from my personal experience. Because these pitfalls are common, if you’ve worked on any AI product, you’ve probably seen them before. 1. Use generative AI when you don't need generative AI Every time there’s a new technology, I can hear the collective sigh of senior engineers everywhere