Topline
Anthropic is clamping down on “cruel behavior” against its Claude artificial intelligence models, announcing in its latest usage update it will bar users from being what it considers excessively mean to its chatbots.
Claude announced the update in a blog post Thursday.
Photo by Artur Widak/NurPhoto via Getty Images
Key Facts
Anthropic said “sustained and needless abusive or cruel behavior” toward its models will be prohibited moving forward.
The update is designed to apply to “extreme cases,” according to Anthropic, which said it will address users who “repeatedly act cruelly toward our models, with no discernible purpose.”
The new rules do not apply to what Anthropic called “common versions of user frustration, pushback, dark creative themes, or model testing and research.”
Anthropic already has a policy where Claude ends conversations with persistently cruel users.
Tangent
Anthropic’s update also includes improvements on prohibiting users from using Claude to spread false information about elections or the voting process. Anthropic specifically noted guardrails against spreading false information on candidates or how to vote, as well as measures barring the impersonation of candidates.
Key Background
Anthropic, which is reportedly targeting an initial public offering this year, has warned of risk factors including AI models exhibiting “self-preserving behaviors” and resisting shutdown efforts. The company has also warned that developing highly advanced models could further increase harm risks against humans. Last month, Anthropic head Dario Amodei urged AI companies to “pace the frontier” in a letter that warned developing models too quickly could lead to a takeover of “the entire internet” by autonomous bots within the next six to 12 months.
Further Reading
Anthropic IPO Prospectus Warns Its AI Could Pose ‘Existential Risks To Humanity’ (Forbes)
Leave a comment