Training a coding model to paint watercolours with TRL and OpenEnv
A Hugging Face blog post describes an open engineering reproduction of a project in which a language model paints watercolours by writing JavaScript through p5.brush, a library described as adding natural drawing tools to p5.js. The post says Surya Narreddi posted a video on 23 August of watercolours painted by a language model, which had over 1.5M views at the time of writing. It states the original idea is Narreddi's and that his earlier blog explained training for close-up flowers, without open artifacts yet, and that a full technical report is coming.
The author says the reproduction uses TRL and OpenEnv, with the reference pool dataset, RL environment, training scripts and trained models all open, running end to end on Hugging Face: training on Jobs, the RL environment and scorer model as Spaces, the pairwise judge through Inference Providers, and artifacts on the Hub.
The reward combines HPSv3, described as an open 7B preference model, and a pairwise judge, Qwen3-VL-30B-A3B-Instruct, called through HF Inference Providers. The pairwise judge compares a candidate painting with four references randomly selected from a hand-rated pool, in both presentation orders, and scores the share of comparisons the candidate wins.
The pool consists of 178 paintings divided into two tiers, love and okay, all generated by models. Four open-weight models wrote p5.brush sketches from openly licensed hibiscus photos from iNaturalist, with a vision model giving written feedback over three refinement iterations before the author rated each render.
Three runs were trained, differing only in the reward weight split between the two model judges. The post says all three learned, both judge runs launched for 200 steps and stopped at 110 with reward still climbing slowly, and that hps-only was run to validate learning. It reports that early learning mostly reduced bad paintings, that the pairwise judge term climbed in the two runs using it, and that no group collapsed to identical rewards. The post also notes the policy did not follow the prompt's instruction of fifteen to thirty filled shapes, with means between 7 and 9.
It lists fixes needed in TRL's GRPOTrainer, including switching LoRA targeting to all-linear for the mixture-of-experts model.
Based on reporting from the original publisher. Visit the source for full context and later updates.