Anthropic Prohibits Harmful Conduct Towards Claude

Anthropic has updated its policy to prevent users from engaging in "sustained and needless abusive or cruel behavior" towards its AI models. This change comes as the company's leadership continues to explore the concept of machine consciousness.

The Verge was the first to report this policy shift. A spokesperson for Anthropic, the creators of the AI chatbot Claude, did not immediately respond to questions regarding what qualifies as "abusive or cruel."

In its updated user policy, the San Francisco-based company clarified that the ban would not cover typical user frustrations, model testing, or themes with a "dark creative" aspect. Previously, Anthropic's policy included restrictions on abusive conduct.

Anthropic's large language models can terminate conversations if a user is persistently harmful. This feature, introduced last August, was described as a protective measure for the AI's welfare.

"On the issue of AI consciousness, we are highly uncertain about the potential moral status of Claude and other large language models, now or in the future. Nonetheless, we take this matter seriously and are working to identify and implement low-cost interventions to mitigate risks to model welfare, in case such welfare is possible. Allowing models to end or exit potentially distressing interactions is one such intervention," states its website.

The idea of AI systems possessing consciousness remains a contentious topic both within and outside the tech industry. Anthropic CEO Dario Amodei has stated that he cannot dismiss the possibility. In contrast, Sam Altman, CEO of OpenAI, has shown reluctance to consider the idea.