아웃풋맥싱 — AI 경쟁은 GPU를 더 사는 게 아니라 있는 걸 최대로 쓰는 일
swyx · Anjney Midha
AMP의 Anjney Midha가 말하는 '아웃풋맥싱'. AI 스케일링의 진짜 질문은 GPU를 더 사는 게 아니라 이미 가진 걸 최대로 쓰는 것 — Google에선 95% 가동률이 곧 장애였고 최고 수준 MFU는 60~70%, 낭비는 자본에서 클러스터 관리자까지의 정렬 실패로 스케일에서 복리로 커진다. FLOPs를 메가와트처럼 흐르게 하는 독립 시스템 운영자(ISO) 그리드·중단가능수요·크레딧 우선순위, DeepMind의 6개월 엠바고가 만드는 연구 사재기의 시장 실패, 1.2GW($40B)·스파이크용 6GW, 14년째 붙든 임종 예측(메디케어 30%)과 규제 병목, API 경계마다 새는 정렬 손실, MatX가 NVIDIA 레퍼런스 아키텍처를 택한 이유, '준비된 마음에 행운이 온다'로 본 Anthropic의 코딩·P0·깨지기 쉬운 문화.
The Professor of Outputmaxxing
생각 덩어리
95% 가동률이면 곧 장애 — Google이 세운 기준선
There’s no excuse, right? I think 95% at Google, which is where my co-founder, Seb, came from, he built the Borg, PBorg/GQM scheduler at Google, and there I think 95% was considered an outage, so 96% node utilization is, should be standard.
MFU should be, I would say the best in class today is somewhere between 60 and 70%.
정렬 문제 — 자본에서 클러스터 관리자까지의 거리가 낭비를 복리로 키운다
Fundamentally it’s an alignment question, which is are the people who are funding the cluster and then deploying the cluster actually aligned?
the number of people in the chain, the supply chain between, the capital and all the way to whoever’s managing the cluster and then whoever’s measuring what the output is, are just so many, degrees of separation away
they initialize the plan, which is kind of like North Star with a team that wants to do good, but then they’re, required to scale so fast instead of iteratively that the wastage just compounds really fast at scale.
상식엔 프리미엄 — AI 스케일링이 상식의 값을 올린다
We have a lot of new capabilities, but that doesn’t mean just abandon common sense. Common sense should always be in fashion.
AI scaling should be putting a premium on the value of common sense and infrastructure because the margin of error now is so much lower and the costs of wastage are so much higher.
책임 있는 인프라 — 빠르게, 그러나 무너지지 않게
Move fast with stable infrastructure. I think now we need to move fast with, responsible infrastructure. People are going to ask where the impact is.
커뮤니티 반발 — 전기료를 낮춰주면 파트너가 된다
let’s call it, $4 an hour. If you’re having to bring up a new data center in a new community, why not just say we’re going to charge 4.50 an hour, and that marginal impact or that marginal increase, we just literally take that and give it to the local community as cash?
Up to 20% of all data centers this year in the US, my understanding is are at risk.
imagine I think if you said there’s a new AI deal. If we’re bringing up a data center in your community, we’re actually going to reduce the cost of your electricity bill. Okay, now we’re talking.
neocloud라는 마케팅 용어 — 20년 된 사업자를 신뢰한다
I think this whole idea of neoclouds being somehow this new category is a lot of marketing speak. There are really good, reliable, trusted data center providers in America who’ve been around 20 plus years.
Are they sponsoring happy hours at NeurIPS? No.
They can run LAN, power, and shell. They have credit histories.
FLOPs를 메가와트처럼 — AMP는 풀스택 통합의 반대
The goal is to try to make FLOPs flow like megawatts, and that is very hard to do today for many reasons. There’s stranded pools of compute all over the place and there’s no fungibility.
we’re actually the opposite of a full stack integration like approach.
독립 시스템 운영자 — 자산을 소유하지 않는 그리드
We see ourselves as what’s called an independent system operator.
If you study like the history of grids, the most enduring ones were those that never owned their own assets.
each of you is guaranteed some base load, but then you kind of schedule your spikes to drive a peak utilization across the town.
중단 가능한 수요 — 크레딧으로 우선순위를 매기는 입찰
the big innovation that was not discovered, but kind of implemented in the space, this infra space maybe three, four years ago at Google was the idea of interruptible demand, right?
It’s a dynamic prioritization Basically. And jobs can get interrupted based on somebody else who’s saying, “what? I have 10 tokens, 10 credits I want to spend on this job.”
this is a thing that has been tried, internally within Google, and it led to Google missing GPT.
연구 사재기 — DeepMind의 6개월 엠바고와 역선택
What’s worse is the paper is actually not even being published anymore ‘cause there’s a six-month embargo inside of DeepMind, right?
So the stuff that gets published is the stuff that’s not good enough.
There’s an adverse selection problem, basically.
기가와트 규모 — 1.2GW는 아무것도 아니다
We only have 1.2 gigawatts of compute. That’s nothing. That’s about $40 billion of cloud spend.
the steady state would be that we have a base load pool Of 1.2 gigawatts at all times Of base load capacity. For spike capacity, right now my estimate is we need roughly six gigawatts over the next four years for all our teams to feel like they were able to keep moving the frontier
임종 예측 — 14년째 머릿속을 떠나지 않은 문제
over 30% of all Medicare, Medicaid spend, at least at that time, was spent on end of life care.
Could you have an AI system make a recommendation that is orders of magnitude more precise about how much time you have left once you’ve been diagnosed with a terminal condition than a human?
The problem remains then and now is regulatory, because you actually can’t shift the burden of the wrong clinical diagnoses from the physician to the AI system.
아웃풋맥싱 — 낭비 반대의 공학
The, from an engineering perspective, it’s very simple. It’s output maxing. It’s the, it’s the department of output maxing.
that doesn’t mean you just like throw 500 GB300, 500,000 GB300s at your suboptimal model scaling and you waste a bunch of compute.
One of the reasons Anthropic has had extraordinary sort of velocity is ‘cause they picked the transform architecture and said, “This is simple. Let’s double down on it,” right?
확장할수록 새는 정렬 — API 경계마다 손실이 생긴다
the more you try to scale, the more division of labor happens, the more specialization happens, and at each step you add abstractions. And wherever there’s an API interface, there’s like loss. There’s communication loss.
Is there a way to actually scale up and scale out Without losing any alignment, without lossy transmission?
You either have to standardize On protocols or API specs that allow lossless communication, or you can come up with a whole new capability that unlocks so much abundance, the standardization doesn’t matter
NVIDIA 레퍼런스 아키텍처 — MatX가 데이터센터를 재발명하지 않은 이유
when they decided to pick the standard For their data center, they picked the NVIDIA reference architecture. So the MatX chips Just plug in to any site that has an NVIDIA bring up planned.
You just can’t fight on every front.
So Jensen’s actually enabled someone like Rainer to build a chip company like MatX, and I don’t see them as competitive.
신뢰 경계 — 공동 설계의 진짜 병목
The primary bottleneck for them is trust boundary. To do co-design well, you need visibility into the next model generation as soon as possible ‘cause it takes two years to tape out.
So when you’re inside the trust boundary of Google, then your systems co-design loop is super tight. When you leave as a founder, one of the biggest risks you take is now you’re outside the trust boundary.
연구자를 CEO로 과소평가 — 과학자는 정신의 스타 선수
Being a CEO, nominally speaking, is not that hard. Being a good CEO is hard. Being a great CEO actually requires a level of performance that scientists who have already published at the top of their field have accomplished.
you are a star athlete. Like, you are an athlete of the mind, and you perform at the highest levels.
To be a great CEO, you basically have to be willing to be confrontational up and down the stack.
리드 vs 승리 — 이기는 게 아니라 앞서는 것
No, I think you want to lead. Yes, so you want to push the frontier. You want to push the SOTA. You want to do something that hasn’t been done before. You want to capture value, but you don’t want to capture so much value that, people think you’re unaligned with your mission or trying to do what’s best for the world.
it’s very important to see the distinction between a heuristic and an axiom.
준비된 마음에 행운이 온다 — Anthropic이 코딩을 뚫은 방식
Anthropic has been the most prepared company for four years. And so then when the right, context data comes in, the right developers start sending in, the right context diffs, Sure, you could say you got lucky, but if you ask me, they’re pr-pretty damn prepared with paranoia for like four years.
Luck favors the prepared mind.
문화는 모트가 아니라 정원 — 매일 돌보지 않으면 바랜다
the thing about culture is it’s very fragile.
“Culture is not a set of beliefs, it’s a set of actions.”
It’s a very brittle, fragile thing that requires daily tending to like a garden.
P0는 첫날부터 코딩 — 결핍은 버그가 아니라 기능이었다
P zero from day one was coding. The reason, the mechanism system there was if we crack coding, Then we will crack AGI.
in hindsight, was a feature, not a bug for Anthropic. The number of people who said no, the number of people who said, “Sorry, we’re all doing investors in OpenAI,” that is competitive difference.
teams who can raise too much money too fast, too early, who don’t have to define what the P zero is, because that’s the only thing when you have scarce resources you got to You got to invest in, Those cultures end up being the most fragile and brittle, and they almost don’t even make it to take off.
선교사이자 용병 — 돈이 척도가 되면 의미를 잃는다
Silicon Valley is both a very missionary place, it’s also a very mercenary place. Sometimes people lose their minds With when they, when big money gets involved
when it Stops becoming, to borrow Goodhart’s law, when it stops becoming just a byproduct and more of a measure, it stops having meaning.