에이전트에게 컴퓨터를 — Composable Computers, Bare Metal, Agent Cloud (Ivan Burazin, Daytona)
swyx · Ivan Burazin
에이전트에게 필요한 건 코드 실행 박스가 아니라 한 대의 컴퓨터 — Lambda와 EC2를 합친 stateful 샌드박스를 베어메탈 위 자체 스케줄러로 구현했다. 60ms 기동·5만 대 75초·하루 85만 샌드박스, RL 워크로드 0→50%가 만든 평균 가동률 15%의 스파이크 문제. Windows 레거시에 잠긴 $10T TAM·MCP가 아니라 CLI·토큰 리셀러의 cold shower·AWS보다 Stripe를 닮을 에이전트 클라우드.
Giving Agents Computers
생각 덩어리
End of localhost — 10년 묵은 명제, 너무 일렀던 CodeAnywhere
I was one of the co-founders of CodeAnywhere, the first browser-based IDE, and so we were thinking a long time of, localhost should die. And you had this article.
Cloud9 came out slightly after us. There was Replit, which came out when we stopped doing it... There was Nitrous.io. There was quite a few that existed at the time, but it was like too early.
Because there was no VS Code, there was no Kubernetes, and Docker had just started... we had to build everything to the whole stack ourselves and that was the key learning that we brought into and that we've been using in Daytona today.
(Ivan이 "It's finally happening"을 되풀이하자 swyx이 받아서 — swyx: "It's finally happening, with maybe sort of non-human users.")
피벗의 신호 — 20~30명 전원이 "No"
A lot of people reached out that were building agents, and they were like, "Hey, my agent needs a compute sandbox runtime," whatever you wanna call it.
People were like, "Oh, why is it different? It's the same thing. We have like EC2, we have VMs, we have all these things." But we saw that everyone we gave it to, it was like 20, 30 people, they all said, "No." Like, "This is not what we need."
'Cause we're infra people. We're not AI people. So I basically took it upon myself to like watch every single podcast that exists... and sort of get up to date, read all the blogs, like get, understand what's going on.
New Year's Eve의 MVP — "절대 쓰레기. 그런데 아이디어는 좋다"
Literally over the New Year's Eve, literally on New Year's Eve, I half vibe coded the first MVP, first minimal viable product of what Daytona is today. And I went to sleep at like 3:00 AM or something like that.
And I sent it to my co-founder, my CTO, and he saw it in the morning. He's like, "This is absolute garbage." "Do not show this to anybody at all, but the idea is good." And so he took two weeks, and he rebuilt it.
API 키를 달라는 전화 — 시장이 끌어당기는 감각
There was no login, just an API key, 'cause it was just a beta or an alpha. And they said, "Oh, we want access."
I've done multiple companies in my life. I've never experienced this, that people literally call you if you do not give them access.
Most people thought it was the same infrastructure for humans and agents. We understood a quarter ago it's not. We just didn't know what was the right primitive.
But the market for every single agent that will exist ever in the future is just like, what is that market? How big is that? And we're like, "We are all in on this."
코드 실행 박스가 아니라 composable computer
People, when they think about Run AI Code, they just think about these small, let's call it isolates, code execution boxes that, you send some code, you get an output. Whereas what Daytona is today is essentially composable computers for AI agents.
My wife is an architect, so she has like a Windows with a 3D graphics card inside to do 3D rendering. Like, as humans, we have different computers or different compositions of computers. And our belief is strongly that agents today and going forward will need all these different compositions of computers to do different types of tasks. And so we offer that basically through an API.
Lambda + EC2 — 노트북 뚜껑을 닫았다 여는 에이전트
Agents will be like humans in the sense of you don't want your laptop to be shut down until you're done with work. Like, and you want to close the lid and open the lid, it's the same state.
We need something insanely fast, how to make it fast, how to make it long-running, and stateful. And so those two things, it's like combining a Lambda and an EC2, right? Those two things together.
2008년으로 돌아간 설계 — 베어메탈 위 자체 스케줄러
We looked at Kubernetes, it wasn't good enough for that. We looked at Nomad, it didn't enable that. And so our history in rewriting our own scheduler at CodeAnywhere is basically what my CTO came up with.
Our third co-founder, when he saw it, he's like, "Dude, what is this? This is like 2008." Like, we went back in time, and he's like, "Exactly."
The snapshot, the point in time, the templates, are also preloaded on the bare metal machines. So when you fire off a sandbox from a template or a snapshot, you're essentially directed to the bare metal machine where that snapshot is based on that NVMe drive, and then it literally just turns on that machine, and it's local. There's no network latency, anything on there.
60ms, 75초, 하루 85만 — 스케일의 세 가지 지표
Our time to spin up one is 60 milliseconds with network latency. So request, spin up, reply, 60, the whole thing, 60 milliseconds. ... But if you wanna spin up 50,000 at once, we are now at about 75 seconds. ... Some others, there's public data around this, like take 2,000 seconds, which is 30 minutes.
The biggest customer of ours does like about 850,000 every single day... we do have a request for half a million concurrent, which is literally half a million CPUs somewhere running.
I don't think the benchmarks equate to market ownership or revenue or anything like that.
두 개의 사용 곡선 — follow the sun과 정사각형
So like a background agent's a Cognition, a Lovable... These are all long-running, background agents. And so if you look at their usage patterns, their usage patterns are similar to human, which is like follow the sun.
When they fire off a run, it's just 100%. And then it just runs, and then it stops. So it's very, the usage pattern is squares basically, right? And it's also not follow the sun, because people will fire it off at midnight before they go to sleep.
Our number one city by user... Is Singapore.
It's interesting that Japan is in the top or like Tokyo's in the top, which is in all the tech cycles it has never been.
평균 가동률 15% — 에이전트 인프라 공통의 신종 문제
We have to lock them into some sort of commits to have that capacity, because we have to have, basically we have to have the capacity for peak. Right? And so right now, Daytona's mean utilization is 15%, 1-5.
Everyone has the same problem. Whereas the usage is super spiky, and this is something that has not happened before... the amplitudes were not this high, right?
GPUs are more expensive than CPUs, right? So you want your GPU running at, what, 100% the entire time. ... And if you then have to like go out and provision machines, you're essentially telling the GPU that it has to wait, and that's incurring our cost.
managed Kubernetes와의 대결 — "I'm never going back"
What we are competing against in that environment is essentially managed Kubernetes. So EKS, GKE, whatever. That is what the vast majority run on. And anyone that has tried Daytona versus GKE, EKS is like, "I'm never going back."
Daytona, although as a compute provider, it's more akin to a Twilio and Stripe from a consumption perspective than it is an AWS. Like you have an API, an SDK, it's quite like easy and seamless to get these things up and running.
An interesting feature is that it's very hard to OOM, or out of memory, our sandboxes, because we can dynamically on the fly [resize].
You can spin up a K3S inside of these things, which unlocks a huge amount of workloads that you can do that you cannot do on other providers.
로드맵은 Slack Connect에서 — 같은 주에 같은 요청 3~5건
When we see one user come with a request, we know it goes on a roadmap if like three to five customers come with the same request in that week. It's like very bizarre. It happens so many times.
I try to be on as many call, quote-unquote "sales calls" I can. I'm in every Slack channel. We literally have about 1,000 Slack Connect channels, something like that.
I feel that Slack Connect is literally LinkedIn what it should be.
Computer use 베팅 — Windows 레거시에 잠긴 $10T
I'm a strong believer that the most efficient way for an agent to work is essentially headless or through, terminal or whatnot. But if we, if we look at knowledge work in general, there's about 100 million knowledge workers in the US, about a billion in the world... the salaries of them aggregate to 10 trillion in the US 50 trillion worldwide.
Most of that work is actually still locked into legacy apps inside of Windows, which is not going anywhere for a very long time. Like, people just won't invest in that.
In the RPA market, which is similar market... 25% of, these white collar, workers', work is automated. If an agent is more sophisticated, can go through more runs, figure stuff out, let's say it's, 40%, right? And so if you take 40% of that, you get to essentially, $10 trillion a year.
Windows specifically is something very new, and the only option right now is an EC2 with, Windows or on Azure. Both of them take anywhere from three to five minutes to spin up. We've created an actual sandbox, so it's a second instead of milliseconds, but you have, point in time snapshots, you have, forking, you have all the things that you have from a sandbox.
"Go log in" — 창업자 본인의 computer use 증언
I kept getting, really well McKinsey-style design reports, but the data said partial data.
I gave it its own account in our company, and then I went to all these services and created a read-only account, so literally like an intern in your company.
I'm like, "Go log in." And it will log into the website, then go in, export the data. It'll export the data and do the thing end to end.
If even a startup like ours, and using all the hottest tools, still needs a computer agent what hope does, Goldman have to have a headless, right?
Apple의 자충수 — macOS 샌드박스를 막는 세 가지 족쇄
So one, you're allowed to run only two parallel VMs per machine, so that's one. Two, you can only license to a different user every 24 hours. So if you come in and theoretically, if I wanna charge you per second and I charge you one second, I have to have it idle for the rest of the day.
They enable you to do memory snapshot, pause, resume, but only on the same physical drive, physical machine.
If anyone at Apple is listening, I very much feel that they are shooting themselves in the foot of the scale of the revenue of compute or licensing they could get if they would just enable a concurrency model similar to what you can get on a Windows and Linux.
Twilio를 닮은 장사 — 엔드 개발자가 아니라 B2B2C
Most of the users that use Daytona are sort of a B2B2C. ... in the researcher world, it's B2B, so you're selling to, labs and neo labs and things like that. But on the long-running agents, it's mostly, from a scale revenue perspective, it's mostly B2B2C, where you have a app layer agent that uses you at a big scale.
It's more akin to a Twilio because you don't really run - As a person, you wouldn't run Twilio.
When your focus is the end developer, it is a very hard sell because they're very price sensitive, very price conscious... Like a lot of companies today are like, "If this is our company, spend as much as you can."
MCP가 아니라 CLI — 통합과 실행의 간극
The MCP is an interface against an API, whereas the CLI is like you can actually go do things. Like this is it. The difference between integrations and actually running scripts or data or analysis against a thing.
Having the agent SDK, from Anthropic... was very interesting. ... they are like, "Oh, I can create this new app, this new agent. All I need, I just use Claude Code, and I throw it into a sandbox, and then I have my interface to the human to that." And so that enabled so many more companies to actually offer this, and then they would pull on sandbox.
시장의 크기 — PC 시장과 맞먹는 클라우드, 다음 병목은 CPU
The laptop, the computer PC market, the PC market is about equal to the cloud market in total. So it's about 150, 180 billion a year.
How many agents are gonna be running in two years, in 10 years, in 100 years? Like And for every single task, they will need one of these. And so how big is that? That market is essentially quote unquote "infinite".
Dylan Patel was at the conference talking about, from SemiAnalysis... was also talking about how CPUs will now be a bottleneck because it will be the constraint.
오픈소스의 실제 효용 — 에이전트가 repo를 읽는다
GitHub stars are the worst, yeah. So you go all the way down to GitHub stars. And so our original one was GitHub stars.
In the new sandbox product we did add a AGPL 3... it is true open source in the sense of an enterprise can use it if it, if it wants, but you essentially can't make a competitor without open sourcing your stuff.
You send the repository to your agent when you're integrating Daytona and it just has more context. It's like, "Oh, okay. This is why this is happening. This is why this, that."
Usually when you would go through procurement to become a vendor of large companies, it would take you like two, three months. We get it done in five days now. And this is not saying that maybe we're great, but it's more, I think, a sign of the market where it is today.
Git은 outer loop — S3에 통째로 던지는 고객과 하루 1,000 PR
The reason was that GitHub as is was an overhead. Like, it wasn't fast enough what they needed, it didn't solve the problem that they needed.
We had one customer that would literally take the entire code base inside the sandbox and... they would just dump it all into a JSON and then push that to S3. And that's it.
And I'm like, if people are doing this, that means there needs to be a new solution to this problem, right? ... I think Git as is still exists in the future, maybe even GitHub exists, but there will be a whole new sort.
And then all that has to go through CI, and then that's the bottleneck. Like, everyone's bottleneck. ... There's one company we're talking to, they do 1,000 PRs a day. ... They have just a queue on that, right?
25명의 응답 속도 — 기능보다 5분 안의 Huddle
The number one thing that people come back to us for is that our, we have an insane responsiveness.
Like, we have had customers like, "Hey, we have a problem. Can you get on Huddle?" Like, we will get on that Huddle like in five minutes, literally. I've done this multiple times.
Of the 25 people in Daytona, I think about 13 of them we have worked with seven years plus. So it's like high trust, high throughput, high we know what we're signing off to do.
I told her about 996, she said, "I wish."
모든 것은 아파야 한다 — "That's why I can't"
I already said, I know that this is gonna hurt, and everything has to hurt. By the way, I'm very much of a feeling that everything has to hurt. Going to the gym hurts. Losing weight hurts. Like, everything has to hurt, right?
You actually have to enjoy the pain and just, if you don't enjoy the pain, it's not for you. And so you get accustomed to that pain.
(Christmas에 로그오프하라는 조롱에 대한 응수 — swyx: "And then your response was?" / Ivan: "Oh, my response was, 'That's why I can't.'")
토큰 리셀러의 cold shower — API를 노출하고 소비에 과금하라
The market is adding premium to SaaS vendors that are reselling tokens. And I think that's incorrect.
You had on SaaS, you had typical SaaS margins, whatever it was, right? Stickiness and all these things. Now what you're doing is you are saying, "Here is my agent, and I have whatever the margin is." It's way worse, right?
Just expose the data. Just expose it. ... So charge me for consumption of API. So you'll have your old seat-based pricing for humans. Charge me for this. The number of agents will skyrocket, and essentially you'll have more usage, and charge for more if your product has value.
I think that there will be cold shower when people understand, no one's actually gonna use and pay for these agents and tokens, and that wasn't actually really a solution, but it'll drop back down.
GPU 샌드박스 — inference가 아니라 렌더링과 RL
Oh, yeah, we will. But not for inference. Like, essentially, what we think about is, the GPU sandbox. ... if you wanna do any type of RL on, CAD or something like that, you will need a GPU in the sandbox, and so that's coming now as well.
Today, from a gross profit margin perspective, it doesn't make sense for us to get in that. You have to raise a large amount of capital, a large amount of risk for, single-digit percentage points.
AWS가 아니라 Stripe — 끝나지 않은 에이전트 프리미티브
The entire infrastructure market is growing 40% plus or minus month over month. Everyone is growing 40% month to month. And that's also a hot take, is like if you're not growing 40%-ish... You don't have to come to work to grow that amount, basically.
There's a high probability that actually owning the CPUs beforehand will be a go-to-market tactic.
Everyone says like an AWS for AI agents, but your answer, it might look more like Stripe than AWS, in a sense. So there will be a cloud built out specifically for agents. And so that cloud will have sandboxes, and it will have web search, and it'll have, databases like SQLite or Neon or whatever, specifically for agent and other things.
We are not at the end of the new infrastructure primitives for agents. There are more coming. So people think like, "Oh, there's nothing else. This it." There are more. Like, we have some ideas about the next ones. We don't have time to do them, but there are definitely more primitives that are being built out for agents, and there will be, I think, a cloud that runs all that together.