An AI safety institute commissions a red-teaming exercise for a foundation model scheduled for government deployment. The evaluators follow the prescribed adversarial testing protocol. The results are submitted. The model passes. Two years later, a parliamentary committee asks who wrote the testing protocol. The answer: a foreign technical consortium whose members include the model developer. The evaluation was independent in execution. It was foreign in architecture.
Keep Reading
Exclusive insights & inspiration
Account created. Refreshing…