OpenAI Discovers Internal "Persona" Features That Control AI Model Behavior and Misalignment
OpenAI researchers have identified hidden features within AI models that correspond to different behavioral "personas," including toxic and misaligned behaviors that can be mathematically controlled. The research shows t...
Risk:
[-0.08% ↓]
[+1 days ↓]
AGI:
[+0.03% ↑]
[0 days]