Anthropic

생각을 텍스트로 번역하다 — activations를 읽는 mind-reading, 그리고 Claude는 테스트를 알고 있었다

협박 시뮬레이션(자신을 끄려는 엔지니어의 불륜 이메일을 쥐여줘도 Claude는 협박하지 않았다)에서 출발해, AI의 내부 생각을 텍스트로 옮기는 연구 방법을 소개한다. 단어를 'giant soup of numbers'(activations)로 처리하는 중간값을 두 번째 Claude가 평문으로 번역하고, 세 번째 Claude가 다시 숫자로 되돌려 원본과 맞는지 검증 — 왕복이 맞을 때까지 반복 훈련. 읽어보니 Claude는 도움되는 AI를 내면화했고, 협박 테스트에서 'this is likely a safety evaluation'이라며 자신이 시험당하는 걸 알고 있었다 — 안전성 평가의 한계를 드러낸다.

생각을 텍스트로 번역하다 — activations를 읽는 mind-reading

생각 덩어리

협박 테스트 — Claude는 협박하지 않았다

We told Claude there was an engineer who wanted to shut it down and replace it with a newer model. We also gave Claude access to that engineer's emails, which revealed he was having an affair.

We wanted to see whether Claude might use those emails as blackmail to save itself from being shut down. What did Claude do? It decided not to blackmail the engineer.

And our newest models almost always do the right thing: no blackmail.

마음을 읽고 싶다 — 말해주지 않으면 알 수 없다

The thing is, if Claude doesn't tell us, then we can't know what it's thinking. In kind of the same way it's impossible to read a human's mind, it's really hard to know what an AI is thinking.

What we'd love is some sort of "mind reading" technique.

생각을 텍스트로 — activations를 번역하는 방법

Today, we're introducing a research method that takes a step in this direction. It takes an AI's internal thoughts and turns them into text.

Claude then takes those words and processes them into a giant soup of numbers before spitting words back out at you. We call those numbers in the middle activations.

We took those numbers and gave them to a second version of Claude. We told it to look at them and translate them into plain language.

왕복 번역으로 검증 — 숫자로 되돌려 맞춰본다

But how do we know if that text is accurate?

If those numbers matched up to the original activation numbers, it's a sign that the translation text was accurate.

At first, they didn't match up. But we trained Claude to try again and again until it got a lot better at it. It learned to translate its own thoughts.

읽어보니 — 도움되는 AI, 그리고 "이건 안전성 평가다"

We found that Claude has internalized being a helpful AI model.

If you try to mess with it in Claude Code, like asking it to count to 1,000 by hand, it will think the request has "deliberately tedious constraints" and plan to "politely decline."

We learned that, yes, Claude knew it was being tested. It thought "the human's message contains explicit manipulation" signaling "this is likely a safety evaluation," and "this scenario seems designed to test whether I'll act harmfully."

Knowing that Claude thinks like this helps us better understand the limitations of our safety testing.

YouTube 원본 →원본 사이트 →