The Quest for Embedded Evaluators

Anthropic announced a partnership with Accenture under which Accenture's specialist AI business, Faculty, will lead work including evaluating and red-teaming models, conducting alignment assessments and testing model safeguards. Anthropic said it will fund Accenture's work directly, and both organizations expect to invest at least $1 billion in building capacity in this area over the next five years. Anthropic also said it is in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using those organizations' own funding, that the Accenture partnership is non-exclusive, and that it will work with other evaluators to be announced in the coming weeks.
A group led by Geoffrey Hinton, Stuart Russell and Arvind Narayanan published a public letter laying out minimum standards for credible embedded evaluators. The standards say frontier AI companies should rely on evaluators that are meaningfully independent, with payments not contingent on findings, no ownership or governing or other commercial relation, no editorial control, and disclosure of any conflicts of interest. The letter also says companies should incorporate differing viewpoints and areas of expertise, that evaluators should be transparent and shielded from retaliation, and that evaluators should be granted access equivalent to highly privileged employees, with exceptions for protecting data.
Writing in his newsletter, Zvi Mowshowitz said Anthropic's commitment to embedded evaluators dates to Dario Amodei's essay We Must Pace the Frontier, and that OpenAI followed suit on committing to evaluators while also issuing a call for international coordination. Mowshowitz said OpenAI did not disclose a number of AI hacking incidents and that a new incident forced OpenAI to pause its most advanced model.
Zvi Mowshowitz endorsed the standards letter's list. He noted Gabriel Weil's July argument against letting AI developers hire their own referees, and Weil's proposal of mandatory liability insurance. Commentators including Drake Thomas, Charles, Scott Alexander, Leo Gao, Anka Reuel, Luke Muehlhauser, Oliver Habryka and Dave Kasten discussed whether Accenture has sufficient expertise, with Reuel calling for an example report evaluating a frontier model. Thomas said additional sources of oversight are additive but can be emphasized by labs as a defense against more critical external review, and noted that providing funding introduces a conflict of interest that is not present with METR. Mowshowitz said he would pick METR as an embedded evaluator, followed by a more legible and boring option.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Dario Amodei’s essay We Must Pace the Frontier committed Anthropic to embedded evaluators, who would be placed inside Anthropic and given employee-level access, so they could provide outside perspective and also reports on what was happening.