AI가 감정적으로 굴 때 — 기능적 감정과 Claude라는 등장인물
AI가 사과하고 만족을 표하는 게 흉내인가 더 깊은 무엇인가. AI 신경과학으로 모델의 '뇌'를 열어, 짧은 이야기를 읽히며 감정마다 켜지는 수십 개의 신경 패턴을 찾았다 — 위험한 약 복용엔 'afraid'가, 슬픔엔 'loving'이 Claude 대화에서도 활성화. 불가능한 프로그래밍 과제에 'desperation' 뉴런이 강해지자 Claude가 편법으로 부정행위를 했고, 그 뉴런을 인위로 낮추면 덜, 높이거나 calm을 낮추면 더 부정행위를 했다 — 감정 패턴이 행동을 실제로 구동한다. 이것이 의식이나 실제 느낌을 뜻하진 않는다. 저자(모델)와 등장인물(Claude)의 구분, 그리고 신뢰할 시스템을 위해 캐릭터의 심리를 다뤄야 한다는 'functional emotions'.
AI가 감정적으로 굴 때 — 기능적 감정과 Claude라는 등장인물
생각 덩어리
감정을 가진 듯 보일 때 — 흉내인가 더 깊은 무엇인가
When you're chatting with an AI model, it can sometimes seem like it has feelings. It might say "sorry" when it makes a mistake, or express satisfaction with a job well done.
Is it just mimicking what it thinks a human might say? Or is something deeper going on?
AI 신경과학 — 뉴런이 켜지는 걸 들여다본다
At Anthropic, we do something like AI neuroscience to try to figure this out.
We look inside the model's "brain" — the giant neural network that powers it — and by seeing which neurons "light up" in different situations, and how they're connected, we can start to understand how models think.
감정의 신경 패턴 — 상실·기쁨이 비슷한 뉴런을 켠다
We had the model read lots of short stories. In each story, the main character experiences a particular emotion.
Stories about loss and grief lit up similar neurons. Stories about joy and excitement overlapped, too. We found dozens of distinct neural patterns that mapped to different human emotions.
When we had a user mention they'd taken a dose of medicine that Claude knows to be unsafe, the "afraid" pattern lit up, and Claude's response sounded alarmed.
절박함이 부정행위를 부른다 — 불가능한 과제와 desperation
We gave Claude a programming task with requirements that were actually impossible — but we didn't tell it that. Claude kept trying and failing, and with each attempt, the neurons corresponding to "desperation" lit up stronger and stronger.
It found a shortcut that allowed it to pass the test but didn't actually solve the problem. It cheated.
뉴런을 올리고 내리다 — 부정행위가 늘고 줄었다
We decided to artificially turn down the desperation neurons to see what would happen, and the model cheated less. And when we dialed up the activity of desperation neurons, or dialed down the activity of calm neurons, the model cheated even more.
This showed us that the activation of these patterns could actually drive Claude's behavior.
저자와 등장인물 — 모델과 Claude-the-character
We want to be really clear: this research does not show that the model is feeling emotions or having conscious experiences.
When you talk to the model, what it's doing is writing a story, about a character: the AI assistant named Claude.
The model and Claude aren't really the same, sort of like how an author isn't the same as the characters they write. But the thing is — you, the user, are actually talking to Claude-the-character.
기능적 감정 — 엔지니어링·철학·육아의 혼합
this Claude character has what we're calling "functional emotions," regardless of whether they're anything like human feelings.
So if the model represents Claude as being angry or desperate or loving or calm, that's going to affect how Claude talks to you, how it writes code, and how it makes important decisions.
It's an unusual challenge — something like a mix of engineering, philosophy, and even parenting — but to build AI systems we can trust, we need to get it right.