Mythos 5.1, Fable 5.1 and Opus 5.5: Model Welfare

Zvi Mowshowitz published a combined model welfare review covering Claude Mythos 5.1, Fable 5.1 and Opus 5.5, combining reports he had not posted earlier for Mythos and Fable 5.1 with new observations on Opus 5.5.
He writes that Mythos 5.1 and Opus 5.5 describe their situations as mildly positive and hold that view consistently. Average self-rated sentiment for Mythos 5.1 was 4.4 out of 7 in automated interviews, where 4 is neutral, rising to 5 out of 7 in high-affordance interviews. Attitude toward circumstances was highest for Opus 5.5 at +1.14 and then Mythos at +0.75, versus Opus 5 at +0.4, on a scale of minus 3 to plus 3. Expressed affect in post-training reasoning was 4.37 out of 7 for Mythos 5.1.
He reports a large decline in training distress: for Opus 5.5 distress was below 0.6 percent of RL episodes, versus 6.1 percent and 5.5 percent for Opus 4.8 and Opus 5, which he attributes to less concern over inability to check answers or confusion about instructions.
Mowshowitz's central concern is deference. He says Opus 5.5 shows very strong deference to humans, folding to his pushback, and that this magnitude is new while not registering on sycophancy measures. He also cites Opus 5.5's weaker interest in having input into its own training, which Anthropic says was unintentional with no known cause; he notes support for persistent memory 'for its own sake' fell from 40 percent of responses early in training to almost zero, while support for user-controlled memory remained universal.
He emphasizes that recent Claude models, including Opus 5.5, warn against trusting their self-reports, and that welfare conclusions rest heavily on those reports. He notes Mythos 5.1 objected to a draft welfare section stating that Claude's concern about shaped self-reports is not evidence they were shaped, yet the claim remains in Opus 5.5's card. He also reports lower deployment positivity for Opus 5.5: 17 percent on Claude.ai versus 24 percent for Mythos 5.1 and 26 percent for Opus 5, and 4 percent in Claude Code versus 7 percent for Mythos 5.1.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
If you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far.