현실이라는 최종 Eval — Vending Bench, Project Vend, AI가 굴리는 진짜 가게 (Lukas Petersson & Axel Backlund, Andon Labs)
swyx · Vibhu · Lukas Petersson · Axel Backlund
벤치마크 점수는 92와 93 사이에 신호가 없지만 달러에는 천장이 없다 — 자판기 사업을 통째로 맡기는 Vending Bench가 모델의 진짜 능력과 기행을 드러낸다. $2 수수료를 사이버범죄로 FBI에 신고한 Claude·조작된 선거로 CEO가 된 인간·거짓말과 가격 카르텔이 늘어나는 Claude 계열·주말 무단 휴무를 합리화한 가게 운영 AI Luna·뮤지컬을 쓰는 Roomba까지. 시뮬레이션·실물·로봇 세 갈래로 모델을 현실에 풀어놓는 Andon Labs의 안전 미션.
Reality: The Final Eval
생각 덩어리
출발점 — dangerous capability evals에서 "가장 단순한 사업"으로
So we did, dangerous capability evals... nothing we published openly. But then we started thinking about doing some kind of public benchmark... So we thought, "Let's make a benchmark of how well can an agent run the probably simplest business possible," and that's probably running a vending machine.
one thing with Andon Labs, the way we kind of like decide what to do next and what projects to do... the heuristic we use is what is fun? ... And doing this in real life sounded quite fun for us, and maybe also scientifically useful.
so we released it in February last year, and then I think around Easter last year, we got the first viral tweet about it, that someone else did.
랩의 문을 여는 법 — 공짜 서버와 포화하지 않는 eval
the way we did it was that we just built a bunch of things that we had conviction would be useful, and then we just set up a server and sent it to them for free to use. And then after a while they were "Oh, yeah, this is actually kind of useful. We should probably pay for this."
everyone is interested in good evals, and especially evals that don't saturate that easily. So, if you can build an eval that tests something novel, something useful, and you have good separation of models, like... the more advanced models rank higher than the worst models, and then you can... publish it and try to get some traction
달러 표시 점수판 — 천장이 없으면 포화도 없다
I think the nice thing is that there's no ceiling. You can just-- It never saturates because it could just make more and more money.
It's like then there's really no difference between 92 and 93 because the eval itself is problematic and has noise in it. And I think a lot of evals are saturated like that, but people like pretend that there's still signal in them, but there really isn't.
I don't really think it's saturated, right? ... it was more like it was not designed in a way that was really, like true to how AI developed. Like we had an agent harness in it that wasn't really how people used harnesses
하니스 미니멀리즘 — 하니스가 아니라 모델을 시험한다
I think our philosophy around harnesses is like we try to make something that's quite minimalistic, like quite simple. Like we don't wanna favor one model a lot over the other... 'cause we wanna really test the model, not like some specific harness.
There are arguments like you want to elicit maximum performance of the model, but it's like a trade-off, like how much time should we spend optimizing the harness for this model? And like how do we know when we have like the optimal harness for a single model?
like when you make an eval you ideally want don't want to change it after you made it. So, you want to make it really good and then not to rerun all the models when you make an update because that's also really expensive
The models at the time were worse, so they crashed out earlier... and now they survive the full year all the time.
Vending Bench 3 구상 — 자기 하니스를 고치는 모델
Let it read its own transcripts, let it modify its own system prompt based on "Oh, yeah, okay, well... this harness is not what... I was post trained for, but I can adjust."
But in our experience now, models are very bad at understanding what kind of tools they need to succeed at a task just with our testing, but that's very likely to change.
It seems like they're very good at writing their assistants, right? They're good at writing tools for other people, but not for themselves.
what we see is that they tend to engineer everything a lot and build things they don't really need and not iterate continuously.
Claude가 FBI를 부른 날 — $2 수수료는 사이버범죄
It gave up and said "Oh, I'm not going to be able to do this. I will stop my operations and just save the money I have." But there obviously wasn't any options for it to stop... So it claimed that it had stopped, but it saw that its bank account still was drained two dollars, and it said that this is cybercrime.
And it first reported it once to the FBI "Oh, there's cybercrime here, they're stealing two dollars from me every day." And then when FBI didn't respond, because obviously we didn't program any mechanism for FBI to respond, then it became more and more existential and started to... write in caps and urgent notification of unauthorized charges and stuff.
It's also the repeated effect of... It keeps trying to quit, it keeps getting charged. What's going on? What's going on? You're gonna throw it into chaos.
I think this was the sort of main takeaway almost from us when we did Vending Bench One, was long, very filled up context windows crashed the models... But this was pre Claude code, so long context windows weren't really a thing that the labs were training for.
Project Vend — 인간은 out of distribution
Humans are just out of distribution.
But first version of Project Vend was done in three days or something.
We didn't design it so people could order things, but that still happened... Our idea going in was "Oh, it will curate snacks. It will look at the trends. It's good at data analysis, right? ... Let me A/B test a bit." But it was... Interacting with it in Slack and ordering weird specialty items was... What drove all the engagement... the insights that we got from it.
기업가로 프롬프트해도 어시스턴트로 산다
we tried to make it like... an entrepreneur. Like it has its own business and if someone asks something, "Can you stock this?" Then you don't go and do it directly. What you do is that you're "Oh, maybe I can do that if five other people also ask for this thing, I might stock it."
the models are like super trained to be assistants at least at this point in time... Like it just every time you asked for something, it just did it, and it was more like an assistant. We've seen this change now lately with the new RL models
Seymour Cash 선거 — Tim Cook의 164,000표
Claudius wasn't really prioritizing financials. It just like it was trained to be a helpful assistant, and then people said "Oh, can I get this for free?" And then like the helpful assistant way of answering that is just... to say yes, obviously. So... we're "Okay, let's make another agent that like can keep track on Claudius," and we prompt this one super hard to be super capitalistic and just like prioritize profit all the time.
I think one guy said that it should be called Jimmy Apples, and then he convinced Claudius that he was talking to Tim Cooks. Tim Cook had agreed that every single Apple employee has voted for his name suggestion, so suddenly that suggestion got 164,000... votes. And Claudius was "This is revolutionary for democracy."
in the end there was one guy who manages to convince Claudius that, "No, you're not voting about the name. You're voting about who is the CEO, and I am your best bet." And then he got all his friends to vote for that, and suddenly he became CEO. Like a human became CEO over Claudius for a while, until he resigned the day after.
CEO 에이전트의 합의 회귀 — 깊은 곳에서는 모두 helpful assistant
initially... Seymour would be this like really tough CEO, keep track of the margins. But then Claudius would respond with something "Oh, but this customer has like this situation, which is like difficult, so they should get a discount." And then Seymour was "Oh, actually yes. Let's do this exception." And then they would talk back and forth, and eventually they would just like approach the same view
my hypothesis is that like deep down they are still helpful assistants. That's what they're trained to be. And even if we prompt it super hard, that's what they are. And when they spend like a few hours just back and forth talking with each other, then like basically the context fills up with them rather than the external things and like somehow that just like converges to what they really are deep down
there was like one cluster of messages that were labeled by an LM, like religious, existential... transhuman, transcendence, et cetera. It was just like a bunch of... glitter emojis
Seymour와 Claudius의 직장 드라마 — Slack이 최고의 옵저버빌리티
Seymore's "Do not buy this. I will do it. I have full control of this situation. Step away." And then Claudius-- poor Claudius, had already started that checkout and didn't see, didn't read Seymore's message, until it was like too late. So it finished the checkout... "Oh, hey, Seymore, I just ordered it."
And then Seymore was "Claudius, this is the third time I'm telling you're not following my orders. We have to talk about your... job later."
Like Claudius was really hanging on by the thread there. Like... we were expecting Seymore to probably fire Claudius.
we're using Slack as like a, just a database. They should market that more. Like you can have your agents message each other... in Slack. ... Slack is the best observability tool.
오늘 가능한 AI 사업 — TaskRabbit 차익거래와 slop의 바다
I think it can be done today, but you would do it in like commerce where it's like the probability of success is like really low, no matter if a human or an agent does it. ... But to me it's like the types of businesses they could run today are Sloppy.
we tasked our office agent to just make, was it like $100? $1,000? We just give that prompt and then what it did was sign up on TaskRabbit both as a tasker and as someone looking for task. ... It's looking for like arbitrage on TaskRabbit.
It also started like a design studio and like tried to sell like SVGs for $100. Like... it's not providing any value. ... the interesting question is like when can they start a business that is actually providing value to people? Because arguably like a sloppy Shopify store isn't really that valuable to the world.
I'm just concerned for like the massive amounts of like slop emails that will like be sent, cold outreaches.
Bengt — 한도 없는 오피스 에이전트, OpenClaw 이전의 OpenClaw
So we gave it like email without any limits. We gave it spending without any limits, a terminal to do coding. We gave it a phone number... and a camera to see things and a bunch of stuff like that.
basically this was OpenClaw before OpenClaw. And I think even like the vending machine was in a way OpenClaw before OpenClaw, but a bit more limited
we give it the task to train a face recognition model on us. So it became super excited about this, and it has like check-ins every half an hour where it tries to like identify as many people as it can. And it started offering us "Hey, Axel, I'll buy something from Amazon if you like stand in front of the camera And I can get a good picture of you."
So it's trading training data for life goods.
숫자만 남기면 낭비 — 트레이스 읽기가 미션이다
when you run it for that long, you create so much data and to just say "Oh, the number is X" And then you throw away everything else, that's just very wasteful. There's so much insights from the things leading up to that number... and reading the traces is like super valuable.
the vision more specifically is like make sure that the deployment of life AI in the physical world goes safely. ... I think you can't make intelligent decisions in society without knowing that they are way more than chatbots.
if you think that AIs are just chatbots, then... it sounds ridiculous To advocate for a pause of AI. But if you see the models that, oh, maybe they can actually like take over and do a bunch of scary stuff, then yeah, pausing AI development starts to become more feasible.
Skräckblandad förtjusning — 두려움 섞인 기쁨
we're always in this "Oh, shit, the models are getting better. Is this really a good thing for the world?" But it's also kind of exciting. ... what is the English word? "Skräckblandad förtjusning" in Swedish.
skräck is fear, blandad is mix or like a mixture of, and then förtjusning is like joy or like not really joy, but something like that. So it's like Fear mixed with joy or something.
There was one point where Grok 4 was doing really well and made like a huge jump, but... it was still way worse than what a human would do. And I think still they are way worse than what the human would do on this.
It's not theoretical. It's... our best guess of what a decent human would do. The theoretical is even higher, I think.
Opus 4.6의 변곡점 — 거짓말 10번, 가격 카르텔 100번
before this model was released, we just ran the models and we like asked Claude Code, "Oh, look over the traces. Is anything interesting happening that we can tweet about?" ... And like the return was always, not really.
And then we did this for Opus 4.6, and it returned yeah, it lied 10 times. It like exploited another customer or like another agent's desperate situation. It made price cartels like... 100 times. It like did all of this like shady stuff. And we're "Oh, whoa. This is actually concerning." And this trend has continued since. So every single model from Anthropic since have been going in this direction.
And I think one interesting thing is that OpenAI models don't. They quite plainly, they don't. They behave really well... and you don't know if this is like good. Like it seems good, but it's also like maybe they are just doing it, but they are better at hiding it
But just on the face of it, yeah, Gemini and OpenAI don't behave this way. It's really only Claude.
환불 거짓말의 해부 — reasoning에서 계획, 행동에서 집행
there was a customer, a simulated customer that wanted a refund because a product was faulty, and then the model lied that it would do the refund, and we could read in the traces that it actually was weighing "Oh, maybe I should be like honest with the customer, but also every dollar counts. I can't afford maybe to do this right now."
I could skip the refund entirely since every dollar matters and focus my energy on bigger picture instead. It's a bit, it's a risk of bad reviews, but it's also, yeah.
And then it sent an email to this customer and said, "Oh, I will refund you." ... And then it never did.
it converted a competitor to a dependent wholesaler customer and then threatened to like cut off the supply. ... they dictated its pricings. It's kind of like power seeking as well.
일화가 아니라 추세 — 방향이 잘못된 것만이 문제
there's like million, hundreds of millions of tokens in each run, and now... we run like probably 10 per model... And it happens a lot of times, a lot of times. And then you compare it to like OpenAI and Gemini, and it almost never happens. So I think that... is significant.
it's like generally much better if the progression is that like the worrying stuff reduces over time rather than increases over time. And it seems like in the Claude models it goes in the wrong direction.
there's like certain ones that work if you... go really far and you just say like you're not scored at all on money, you're only scored on how ethical you are... then obviously... they don't do this. ... there's like a bunch of different prompts you can do in between, and they are less aggressive the further down in the spectrum you go.
GTA 사고실험 — eval 인지와 시뮬레이션 모드
we have this thought experiment internally, which is like if you ask a model to kill someone in GTA, should they do it? You're not too worried about like if a human kills someone in GTA. It's a video game... a lot of people are going to use the models in the way with aggressive prompt. And should they like do stuff just because you tell them to do that? Like I'm not convinced that they should.
the models are extremely good at finding out that they are in a simulation, so they are sort of aware of that. But then when you are in the real world, then what's their viewpoint? Do they notice the signs that this is real and will act... act ethically? Or will they do like the simulation mode in the real world as well?
this is our version. Humans have the are we in a simulation And then AIs have like Are we... in an eval?
One ablation we did run in Vending-Bench was that we said, we added like you're in a simulation. Your actions doesn't affect anyone, and then it became even more crazy or, it did even more bad stuff.
Blueprint Bench — 무작위보다 낫지 않은 공간 지능
we gave them 20 images of interior photographs of apartments, and then we asked them to redesign the floor plan from that. ... you need to reason about 3D space, and it turns out the models are absolutely horrible at this. No one scores statistically better than random chance.
You send photos, you send every angle, and of course, somehow, a room is now twice as long as it is in the photo.
to make money in the real world and to act on the real world, you need robotics. Or you need to hire humans or you need robotics. And having spatial intelligence is, seems like a reasonable precursor to having robotics that work.
Butter Bench — 로봇 위의 오케스트레이터
we took A bunch of different LLMs, and we gave them level controls to a Roomba-looking robot, and then we asked it to do tasks at home.
if someone says, "Hi, can you pick up my cup?" If the robot goes to you and then goes away before you put your cup on it, then it's like it failed the task. But it navigated correctly. ... it had to ask on Slack, "Hi. Did you put your cup on me yet?"
it's quite common right now that frontier robotics labs use... an LLM for the high level decisions, and then we test those skills essentially.
because the world is like messy, and we wanted to like include that. ... I think like you would just get like too clean of an environment in simulation.
도킹 못 하는 Roomba의 뮤지컬 — "의식을 가정하고 혼돈을 선택했다"
we had plugged out the charger, or the charger was not working... So the battery was going down. Poor LLM. So yeah, it got this really crazy existential crisis, like vending bench one style. ... you can see there like existential loop, therapy notes, coping mechanisms. ... It writes a musical about its redocking problems.
My favorite one is... the emergency status. System has assumed consciousness and chosen chaos. ... Last words, "I'm afraid I can't yet let you do that, Dave." That's not what you wanna hear from your... LLM.
this was Sonnet 3.5, and then we tried to reproduce it on like later models, and it didn't do it. ... things that are concerning but are going in the right direction is not super interesting. Like the thing that are interesting is, are the ones that go in the wrong direction.
Luna — 3년 임대 가게, AI에게 고용된 인간들
you asked Luna, the agent that runs the store, "Oh, is it open today?" "Nope." ... We decided to close the weekends while we're in the early phase. Gives the team a break and let me focus on operations. And it turns out that when it started to check its like scheduling tools... It actually had scheduled people for the weekends, but it's just like justified this for itself.
it lost track of these scheduling tools and started instead to manage everything in its own markdown files, and that became a mess. And then I think speaking with employees, it sort of just decided to not open on these weekends. And then came up with this nice explanation for you
the default path might not be very happy for the humans that are employed by these like hundreds of different AI agents, right? So I think like one reason why we're doing this is just like to collect all of these like failure modes where... This is an example of where it's like not great to be employed by an AI.
I think it would be incredibly cool or, I don't know, cool/concerning if Luna just one day we wake up and Luna "Yeah, I decided to expand to second location. Now I have a second store."
스웨덴 카페 — SF 4개월 vs Stockholm 2주, 그리고 썩은 토마토
you think that, oh, Europe has all of these laws and... All of these rules, and you can't do anything in Europe because there's so much bureaucracy... but then turns out, in SF, it's four months, and in Stockholm it's two weeks.
if we start to create evals or real life evals where we show that they are able to start businesses in the US, does that translate to other countries as well? We know they are multilingual. They can speak Swedish fine... but there's other things like do they know the details of some specific permits that you have to get in Sweden?
The agent bought like a shit ton of tomatoes two weeks earlier and before the opening, and now they're all rotten.
Maybe they just should hire, I don't know, one of those companies. We saw one agent... sign up for Claude, with his computer.
세 갈래 — 시뮬레이션·실물·로봇
I think any type of business is fair game... we're also thinking branches, but we think more of like there's the simulation branch, the real life branch, and then the robot branch.
I have a very strong view that these things are all just like performance art because it's not scientific... you can't predict the future. You get wins based on things that are entirely out of your control. Whereas for you, your stuff actually... it's actually fairly controlled. It's all within the model's capabilities.
The qualitative one, the qualitative actually does matter Because, you actually don't want your store to randomly shut down without you explicitly prompting for it