Unless Its Governance Changes, Anthropic Is Untrustworthy
# Summary of Article Critique on Anthropic Leadership ## Core Accusations Against Anthropic The article argues that Anthropic's leadership has been misleading and deceptive by holding contradictory positions that shift consistently with OpenAI's direction. It claims the company lobbies to kill or water down regulations, a stance supported by employees from major AI firms who find such regulatory frameworks helpful for safety. Furthermore, the text asserts that this behavior violates the fundamental promise upon which Anthropic was founded. ## Internal Culture and Pressure Anthropic possesses a strong internal culture broadly aligned with Effective Altruism (EA) views and values. However, there is significant pressure on staff to appear compliant with these views to retain talent and ensure loyalty. The author contends that this cultural pressure creates an environment where it remains unclear what leadership would actually do when critical decisions matter most. ## Proposed Questions for Leadership The text suggests a series of direct questions for Anthropic employees, the policy team, the board, Dario, and the public to consider regarding integrity and regulation: * Why does Anthropic consistently oppose regulations that would slow competitors while potentially increasing overall safety? * To what extent does Jack Clark act as a rogue agent versus coordinating with broader leadership? * Would leadership violate promises to employees if they chose between walking back commitments and falling behind in the race? * Can leadership justify dropping promises without providing strong justification, or would they instead direct attention toward those dropped promises? * Is there a likelihood that representatives will lie to policymakers and the public based on leadership's priorities rather than formal mechanisms? ## Scenarios of Leadership Priorities The article explores two contrasting world scenarios regarding Anthropic's decision-making: 1. **Competition-Driven World:** In a scenario where leadership prioritizes competition with China and winning the race over existential risk (x-risk), they might mislead employees who care about truthfulness to maintain competitive advantage. 2. **Truthful World:** Conversely, in a world where Anthropic is truthful about its nature and trustworthiness, the evidence suggests a different trajectory of behavior. ## Decision-Making in Pessimistic Scenarios The text questions how Anthropic would handle evidence suggesting alignment is difficult: * Would leadership propagate future evidence regarding the hardness of alignment? * Will they attempt to make everyone pause if new data indicates that living in an "alignment-is-hard" world is probable? ## Employee Agency and Exit Strategies Finally, the article addresses the role of individual employees in assessing their own involvement: * Employees must determine which worlds would cause them to regret working for Anthropic on capabilities. * They should consider how likely it is that they are currently operating within such a world. * The text emphasizes the necessity of learning, updating one's stance based on new evidence, and deciding when to leave the company if trust is broken. ## Contact Information The author thanks individuals who provided feedback and shared information regarding these facts. For those wishing to share additional details or contact the author directly, a Signal handle (@misha.09) is provided for communication.