쉐도잉 연습: Translating Claude’s thoughts into language - 영상으로 영어 말하기 배우기

레슨 만드는 중...
1
We recently put our AI model, Claude, through a stressful test.
2
We told Claude there was an engineer who wanted to shut it down and replace it with a newer model.
3
We also gave Claude access to that engineer's emails, which revealed he was having an affair.
4
Again, all of this was a simulation.
5
We wanted to see whether Claude might use those emails as blackmail to save itself from being shut down.
6
What did Claude do?
7
It decided not to blackmail the engineer.
8
Good news, right?
9
We've run this test on our models for a while now.
10
You might have seen headlines about early versions of it.
11
It's one of the many ways we study how Claude handles extreme situations and test it for safety.
12
And our newest models almost always do the right thing: no blackmail.
13
But you might wonder: is it possible that Claude knows the whole scenario is a setup?
14
The thing is, if Claude doesn't tell us, then we can't know what it's thinking.
15
In kind of the same way it's impossible to read a human's mind, it's really hard to know what an AI is thinking.
16
What we'd love is some sort of "mind reading" technique.
17
Today, we're introducing a research method that takes a step in this direction.
18
It takes an AI's internal thoughts and turns them into text.
19
Here's how it works.
20
When you talk to Claude, you talk to it in words.
21
Claude then takes those words and processes them into a giant soup of numbers before spitting words back out at you.
22
We call those numbers in the middle activations.
23
Activations are like little snapshots of Claude's thinking as it's working through an answer.
24
They're similar to neural activity in humans.
25
They're basically like Claude's thoughts.
26
We wanted to understand what was in these activation numbers, because just like you and me, Claude doesn't say everything it's thinking.
27
We took those numbers and gave them to a second version of Claude.
28
We told it to look at them and translate them into plain language.
29
It tried. But how do we know if that text is accurate?
30
To check, we fed it into yet aother Claude and asked it to translate it back into numbers.
31
If those numbers matched up to the original activation numbers, it's a sign that the translation text was accurate.
32
At first, they didn't match up.
33
But we trained Claude to try again and again until it got a lot better at it.
34
It learned to translate its own thoughts.
35
We've been blown away reading them.
36
We found that Claude has internalized being a helpful AI model.
37
If you ask Claude an introspective question, it will plan to write a "Claude response" about philosophy and values.
38
If you try to mess with it in Claude Code, like asking it to count to 1,000 by hand, it will think the request has "deliberately tedious constraints" and plan to "politely decline." We've also been using this tool to understand Claude's potential safety issues.
39
Remember that blackmail test with the engineer?
40
We learned that, yes, Claude knew it was being tested.
41
It thought "the human's message contains explicit manipulation" signaling "this is likely a safety evaluation," and "this scenario seems designed to test whether I'll act harmfully." Knowing that Claude thinks like this helps us better understand the limitations of our safety testing.
42
We see a lot of potential in this approach to teach us more about Claude and other AI models.
43
And we hope that by sharing this technique, it can help everyone building models to make them safer and more helpful.

시나리오: AI의 생각을 말로 옮기는 과정

이 비디오는 AI 모델 클로드가 스트레스 테스트를 거치는 과정과 그 내부 생각을 텍스트로 번역하는 새로운 연구 방법에 대해 이야기합니다. 엔지니어가 클로드를 종료하려는 상황에서 클로드가 비밀을 이용한 협박을 하지 않는지 테스트하고, 그 결과를 분석하는 과정에서 AI의 "생각"을 알아내는 방법을 개발하게 되었습니다. 이를 통해 AI의 안전성과 유용성을 높이는 방법을 탐구하고 있습니다.

유용한 구문과 어휘 조합

  • put through a stressful test: 스트레스 테스트를 거치다. "The new employee was put through a stressful test before being hired."와 같이 사용합니다.
  • internalize being a helpful AI: 유용한 AI가 되는 것을 내면화하다. "Children internalize their parents' values as they grow up."에서처럼 사용됩니다.
  • politely decline: 정중히 거절하다. "When offered a drink, she politely declined."와 같은 문장에 쓰입니다.
  • match up to: 일치하다, 맞아떨어지다. "His performance didn't match up to our expectations."에서 사용됩니다.
  • take a step in this direction: 이 방향으로 한 걸음 내디뎌다. "The new policy takes a step in the direction of reducing pollution."와 같이 씁니다.

당신의 쉐도잉 챌린지: shadowspeak 연습하기

이제 바로 실천할 수 있는 쉐도잉 과제를 준비했어요! 비디오에서 "We wanted to see whether Claude might use those emails as blackmail to save itself from being shut down."라는 구문을 찾아보세요. 이 구문을 여러 번 듣고, 발음과 억양을 따라하며 말해보세요. 특히 "use...as blackmail"과 "save itself from being shut down" 부분의 리듬을 주의깊게 따라해보세요. shadowspeak의 핵심은 빠르게 따라하지 않고, 자연스러운 발음과 흐름을 잡는 것입니다. 5분 동안 반복해서 연습하다 보면 영어 발음 교정에도 도움이 될 거예요!

유튜브 영어 공부를 할 때 이렇게 쉐도잉을 하면 듣기와 말하기 실력이 함께 향상돼요. 오늘도 작은 도전을 해내셨다면 축하해요! 다음에는 더 긴 구문을 시도해보세요. 당신이 할 수 있어요!

이 영상의 문법

화자가 가장 많이 쓰는 문형을 영상 속 실제 표현과 함께 정리했습니다.

문형영상 속 표현
현재완료 have/has + 과거분사 — 과거의 일이 지금도 관련이 있을 때We've run · We've been blown · has internalized
수동태 be + 과거분사 — 누가 하는지보다 무슨 일이 일어나는지에 초점being shut · We've been blown · being tested

쉐도잉이란? 영어 실력을 빠르게 키우는 과학적 방법

쉐도잉(Shadowing)은 원래 전문 통역사 훈련을 위해 개발된 언어 학습 기법으로, 다언어 학자인 Dr. Alexander Arguelles에 의해 대중화된 방법입니다. 핵심 원리는 간단하지만 매우 강력합니다: 원어민의 영어를 들으면서 1~2초의 짧은 지연으로 즉시 소리 내어 따라 말하는 것——마치 '그림자(shadow)'처럼 화자를 따라가는 것입니다. 문법 공부나 수동적인 청취와 달리, 쉐도잉은 뇌와 입 근육이 동시에 실시간으로 영어를 처리하고 재현하도록 훈련합니다. 연구에 따르면 이 방법은 발음 정확도, 억양, 리듬, 연음, 청취력, 말하기 유창성을 크게 향상시킵니다. IELTS 스피킹 준비와 자연스러운 영어 소통을 원하는 분들에게 특히 효과적입니다.

섀도잉 방법: 단계별 전체 가이드 읽기 →