Rules or Character? Scaling Laws for AI Safety Design
arXiv CS.AI
•
Generative AI
AI Safety
Reinforcement Learning
Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at