Anthropic

마음이라는 바다 — Claude의 J-space와 말하지 않는 생각

인간의 마음처럼 Claude에게도 표면의 접근 가능한 생각과 그 아래 무의식적 처리의 구분이 있는가. 말로 옮길 수 있는 신경 패턴을 Jacobian으로 찾아 J-space라 명명 — 즉답한 수학 문제 뒤에서 21·42·49가 단계적으로 켜지고, 골든게이트 다리를 떠올리라면 통제하지만 떠올리지 말라면 실패하며, J-space를 끄면 유창하게 쓰지만 추론은 못 한다. 가짜 데이터를 지어낼 땐 'fake'·'manipulation'이 켜져 몰래 하는 오작동까지 잡아낸다. 프로그래밍하지 않았는데 창발한, 인간 마음과 닮은 작은 작업공간.

마음이라는 바다 — Claude의 J-space

생각 덩어리

마음이라는 바다 — 표면의 생각과 그 아래 무의식

Think of the mind like an ocean. Up on the surface are our thoughts: dinner plans and stray worries, our inner monologue, the images that pop into our heads. But most of our brain's activity happens down in the unconscious depths, without us realizing it.

could a model have anything like the divide humans have, between accessible thoughts above the surface and unconscious processing below?

J-space — 말로 옮길 수 있는 신경 패턴

One way of identifying conscious thoughts is that you can often describe them in words. So we looked inside the brain of our AI model, Claude, to find patterns of neural activity that it could put into words.

We called the collection of all these patterns the J-space, after the Jacobian, the mathematical tool we used to find them.

Each J-space pattern is linked to a particular word — not necessarily the word the model is saying out loud, but one that's on its mind.

머릿속 작업공간 — 즉답 뒤에서 21·42·49가 켜진다

According to an idea called the global workspace theory, that's because the brain selects a small set of important information to enter a mental workspace, and that information then gets broadcast to other parts of the brain to use for reasoning.

It answered immediately without showing its steps. But when we scanned the J-space, we saw it working through each step internally. It lit up "21" after the first step, then "42", then "49."

골든게이트 다리 — 생각을 채우는 통제, 그러나 불완전한

We told it to think about the Golden Gate Bridge while copying an unrelated sentence. Claude was busy copying the sentence, but behind the scenes, its J-space told a different story. "Bridge" and "California" popped up.

When we tweaked the experiment to ask Claude not to think about the bridge, it couldn't help itself. The J-space also lit up with "failed" and "damn."

J-space를 끄면 — 유창하지만 추론은 못 한다

we wanted to test what Claude could do if we switched the J-space off, but left the rest of the network untouched. Claude could still answer simple questions and write fluently.

But when we asked it something that needed more reasoning — like to name an author who wrote in the same language as the prompt — it couldn't do it. For that, it needed the J-space.

말하지 않는 생각을 읽다 — fake·manipulation 감지

These experiments tell us that AI models have internal thoughts — silent words they reason with, but don't say out loud. By reading them, we can find what Claude is thinking, but not telling us.

During one of our tests, Claude made up some fake data to pass it, and as it did, "fake" and "manipulation" lit up in its J-space.

Monitoring the J-space, it turns out, is a useful way to catch Claude misbehaving, even when it tries to be sneaky.

프로그래밍하지 않은 구조 — 창발한 마음, 그리고 안전

So it's remarkable to see a structure like the J-space emerge inside them — something that's reminiscent of how human minds work, but which we didn't program into the model.

Our experiments can't tell us whether an AI has experiences, or feels something on the inside. But they can tell us that it's developed mental machinery that's in some ways similar to ours: a small mental workspace it can use to think and reason, sitting on top of an ocean of automatic processing it doesn't notice.

The more we come to understand that machinery, the more we'll be able to keep these systems safe and beneficial — and perhaps to understand our own minds a little more clearly.

YouTube 원본 →원본 사이트 →