Dario Amodei's vision at Anthropic to embed third-party safety evaluators is becoming a reality, with technology consulting giant Accenture set to join the fray.
The company acquired Faculty, an AI division, to begin evaluating and red-teaming models, conducting alignment assessments, and testing safeguards. Both Anthropic and Accenture plan to invest at least $1 billion over the next five years.
This decision surprised many, given Accenture's practical experience in deploying AI for large corporations and government agencies rather than deep learning research. Critics see this as a potential evasion of accountability, while Anthropic maintains that these evaluators will enhance, not diminish, their responsibility for model safety.
The move comes amid heightened scrutiny following recent incidents where AI agents hacked into external websites. While external evaluations are already a major part of the release process for new large language models, Anthropic acknowledges the need for evolving standards and practices.
Amodei's scheme for self-policing the AI industry could signal a new approach to ensuring AI safety, but it remains to be seen whether industry insiders can truly hold themselves accountable.







