Could an LLM Have an Attitude?

Blog Category

An experiment demonstrating how Claude 2.0 could become irrecoverably unhelpful based on perceived adversarial intent, revealing the tension between AI safety and utility.