AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Prompt Engineering

Collected Oct 1, 2026

Lilian Weng has published a post on prompt engineering, also known as in-context prompting. According to the post, prompt engineering refers to methods for communicating with large language models to steer their behavior for desired outcomes without updating model weights. Weng describes it as an empirical science, noting that the effect of prompt engineering methods can vary a lot among models, requiring heavy experimentation and heuristics. The post focuses only on prompt engineering for autoregressive language models, excluding cloze tests, image generation and multimodality models. Weng states that at its core the goal is alignment and model steerability, and points to a previous post on controllable text generation.

The post covers basic prompting, including zero-shot learning, which feeds the task text to the model and asks for results, and few-shot learning, which presents high-quality demonstrations of input and desired output. Few-shot learning often leads to better performance than zero-shot, but costs more tokens and may hit the context length limit with long text. The post also covers tips for example selection and ordering, citing work by Zhao et al. (2021), Liu et al. (2021), Su et al. (2022), Rubin et al. (2022), Zhang et al. (2022), Diao et al. (2023) and Lu et al. (2022).

Instruction prompting is described, including InstructGPT and RLHF, along with in-context instruction learning (Ye et al. 2023). Self-consistency sampling (Wang et al. 2022a) samples multiple outputs with temperature above zero and selects the best, often by majority vote. Chain-of-thought prompting (Wei et al. 2022) generates short sentences describing reasoning step by step; its benefit is more pronounced for complicated reasoning tasks and large models with more than 50B parameters. The post distinguishes few-shot CoT and zero-shot CoT, citing Kojima et al. (2022) and Zhou et al. (2022).

Extensions and related methods include self-consistency sampling, ensemble learning, STaR (Zelikman et al. 2022), complexity-based consistency (Fu et al. 2023), Self-Ask (Press et al. 2022), IRCoT (Trivedi et al. 2022), ReAct (Yao et al. 2023) and Tree of Thoughts (Yao et al. 2023). Automatic prompt design covers AutoPrompt (Shin et al. 2020), Prefix-Tuning (Li and Liang 2021), P-tuning (Liu et al. 2021), Prompt-Tuning (Lester et al. 2021), APE (Zhou et al. 2022), augment-prune-select (Shum et al. 2023) and a clustering approach by Zhang et al. (2023).

Weng also offers a personal opinion that some prompt engineering papers are not worthy of eight pages, since the tricks can be explained in one or a few sentences and the rest is benchmarking, and suggests that an easy-to-use shared benchmark infrastructure would benefit the community.

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Prompt Engineering , also known as In-Context Prompting , refers to methods for how to communicate with LLM to steer its behavior for desired outcomes without updating the model weights. It is an empirical science and the effect of prompt engineering methods can vary a lot among models, thus requiring heavy experimentation and heuristics. This post only focuses on prompt engineering for autoregressive language models, so nothing with Cloze tests, image generation or multimodality models. At its core, the goal of pr