Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
Hugging Face
Read full postA new paper introduces a method for large language models to refuse only specific harmful subsets within a topic, rather than rejecting entire topics, enabling more precise safety controls tailored to deployment needs.



