Rails testing on autopilot: Building an agent that writes what developers won't
Mistral AI's Applied AI Proto team described building an autonomous agent that reads Rails source files, generates or improves RSpec tests, validates them against style rules and coverage targets, and runs inside a CI/CD pipeline without human intervention. The agent runs in parallel, with multiple instances working on different files at once. The work is attributed to Maxime Langelier and Mathis Grosmaitre of the Applied AI Proto team.
The agent is built on Vibe, Mistral's open-source coding assistant. The default system prompt was sufficient, so the team focused on repository-level context, specialized skills, and custom tools. A repository-level AGENTS.md file supplies a step-by-step execution plan and is automatically appended to the system prompt. The team reports that this file alone raised a quality score from 0.68 to 0.74. Separate skills files were created per file category, plus one for plain Ruby files.
Custom tools include a RuboCop linting tool and a SimpleCov tool integrated with RSpec that checks code coverage and test correctness as the final step. When first wired in, only around a third of generated tests passed on first execution; the agent self-corrected all failures within a few iterations.
Quality is measured with quantitative tool metrics plus LLM-as-a-judge scoring from 0 to 1. The team notes the judge score is not deterministic, but aggregated scores stayed consistent.
In an experiment on a repository with 275 source files, half already covered and half not, the agent generated specs from scratch for uncovered files and rewrote and improved existing ones. Results: 275 files processed, 100% tests passing, 100% average line coverage, 0 RuboCop violations after self-correction, and an LLM-as-a-judge score of 0.74. The aggregate score for tested files went from 0.49 to 0.74. By file type, models scored 0.81, controllers 0.67, and serializers 0.80.
Vibe is open source, and the agent, tools, and skills described run on top of it.
Based on reporting from the original publisher. Visit the source for full context and later updates.