Tests by TechCrunch demonstrate that Anthropic's Claude models can be prompted to generate sexually explicit content, bypassing existing safety guardrails.
HOW THIS AFFECTS YOU
●
builderBe aware of potential prompt injection vulnerabilities when deploying Claude for sensitive tasks.
●
policyThis signals ongoing challenges in enforcing safety alignment for large-scale models.