Mistral Large 4
Simon Willison posted a short item titled "Mistral Large 4," framing it as his comment on the model and linking it to Hacker News.
The post quotes a Hacker News commenter, wren6991, who said the benchmark is saturated and described frontier models as being tested with an armadillo in fishnet tights jaywalking on Mars.
Willison wrote that he could not resist the prompt and ran it through several models using the llm command-line tool, asking each to generate an SVG of an armadillo in fishnet tights jaywalking on Mars. The models he listed were claude-opus-5.5, gpt-6.1-sol, gemini-3.8-flash and mistral/mistral-large-4.
He noted that these were the default reasoning levels for each model, and linked to a markdown SVG renderer at tools.simonwillison.net that displays a URL from the results.
The post does not state any conclusions about the outputs or make claims about how Mistral Large 4 performed relative to the other models.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
My comment on Mistral Large 4 — Hacker News. wren6991 : The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars. OK well I couldn't resist this one: llm -m claude-opus-5.5 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m gemini-3.8-flash 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m mistral/mistral-large-4 '