Dark Factory — diff를 읽는 속도보다 빠른 출하, 공장 관리자가 된 엔지니어
Vincent Koc
OpenClaw 코어 메인테이너 Vincent Koc의 AI Engineer 강연. 피크 하루 800커밋·개인 3,000커밋의 속도는 운이 아니라 엔지니어링이라는 주장 — 산업혁명의 직공→공장 관리자 전환·새벽 2시에 시작한 2,700커밋짜리 great refactor·스윔레인 공장 운영·reasoning token을 '느끼는' 감각·dot skills·60k PR 트리아지. 2025가 token maxing이었다면 2026은 token efficiency, agent in the loop.
Dark Factory — diff를 읽는 속도보다 빠른 출하, 공장 관리자가 된 엔지니어
생각 덩어리
출하 속도가 밈이 된 프로젝트 — 운이 아니라 엔지니어링
I'm going to talk about what I call doc factories and how OpenClaw ships faster than you can read the diff.
I wake up, there's a new technological advancement. ... It's this joke that we're shipping at insane speed and the velocity is just absolutely phenomenal. And some of you might think, "Oh, this is some luck or we're just like Ralph flipping to the max." I think there's actual engineering work here.
엣지는 원래 janky하다 — VR 고글 3시간과 GitHub rate limit
The funny thing with this one here was that I didn't use it for 5 minutes. I used it for 3 hours. And I played Team Fortress 2, had an absolute blast, and then I vomited for 3 hours after that cuz my vision turned into bee visions.
What I'm trying to say here is that like anything on the edge is going to be janky, it's going to be horrible bit horrific. It's going to be uncharted territory.
Working on OpenClaw and being part of the team that ships ... an insane velocity of commits to a point where I get rate limited by GitHub on an hourly basis is an interesting experience.
산업혁명의 재방송 — 직공의 손에서 taste로 옮겨가는 병목
In my day job, I kind of work in the space of Evals, which everything is sort of structured and there's telemetry and it has to be all perfect. And I work on a project where I have this blind faith in the hardness. And it's this kind of two worlds, but they're starting to come together.
We used to have handlooms in cottages, centralized mills everywhere. Craftsmen were the factory workers. But the bottleneck was the weaver's hands.
We're now switching to a world where engineers writing code in editors, not so much. Swarms across repos. Engineers are becoming factory managers ... And the bottleneck becomes taste, you know, that lovely word.
모두가 몰래 쓰고 있다 — ChatGPT 시절의 데자뷰
Very similar to the ChatGPT era, where everyone denied it at scale that they were using ChatGPT. Everyone was in this absolute fear-mongering sort of world. But what the reality was that everyone was using it. ... And the same thing is happening with these autonomous agents at scale.
Some organizations have openly come up with it. So, for example, Anthropic with their recent work they did on building a new C compiler. We had Spotify saying they're no longer writing code by hand, supposedly. Steve Yegge ... saying he pushes about 50 PRs a day total solo. He calls himself a vibe maintainer. I can kind of relate to that.
And OpenClaw where we're pushing at the peak, we were doing 800 commits a day. And realistically, like there's about 10 to 15 core maintainers all with day jobs. It's kind of astronomical in terms of scale.
하루 3,000커밋 — 커밋 히스토리에 찍히는 수면 시간
For me, this was March 15th. ... Where I hit close to 3,000 commits per day. And if you actually took look, my commits actually stop when I go to sleep. So, if you want to see when I go to sleep and when I wake up and how many hours of sleep I have, you can just take a look at my commit history.
This is going to become the norm everywhere else. ... this scale of velocity is going to be normal. And trying to review PRs and go through all this nonsense may not work. But somewhere in the mix is engineering. There is a form of engineering that's going to happen.
Commit maxing에서 bot looping으로 — 루프에 보상 구조 넣기
So, we did commit maxing. You know? Let's just go there. Smash as many commits as we can.
And this reminds me of Ralph looping, right? ... where you're like, "Hey, I'm just going to like give you a task. I'm going to burn tokens for like 8 to 9 hours." And you're waiting, you know? ... You're hoping something happens. Maybe something happens. I don't know.
What if we had a bit more of an opinionated approach to this? ... what if we call it bot looping? ... Do we need more than just tokens? ... what does that reward mechanism look like? ... Yes, let's run loops, but let's be a bit more smart about how we do this.
Nvidia에서의 하루 — 둘이 합쳐 에이전트 60~70개
Right about the time you saw those 3,000 commits, ... this was the day before, I was at Nvidia with Peter ... And they were like, "Hey, we're building Nemo Claw." I'm like, "What?"
He's running about maybe 15 code sessions and he's got his Mac Studio at home he's VPN'd into. I'm running another like 10 or 15. And collectively between us, we're probably running with sub agents included, maybe up to 60, 70 agents. ... on the foreground, maybe 15 swim lanes, if you want to call it that.
Funny thing is, we're working on Nemo Claw on one side, but one maintainer decided, "I'm going to move some stuff around. I'm going to move a couple of folders around." And that was moving entire channels. So, like all our conversations with like MS Teams and Slack ended up moving to another location in the code base.
Great Refactor — 새벽 2시, "코드베이스 전체를 갈아엎자"
We have lots of people raising PRs. And what they actually want is to build features. The thing is, we don't want to give everyone every single feature that they want, in which case it becomes bloat.
The challenge becomes who do I say no to? It's not about saying yes. In a world where tokens are cheap, I can just say yes to absolutely everyone and merge everything in. But that's going to turn this code base into an absolute fire dump.
The vision was actually we need to cut this code base down. We need to rip it into pieces. And a plugin architecture somewhat made sense. Imagine if you're OpenAI or Mistral or Anthropic, what if you own that piece of the provider code and it was handed to you and it was separate from everything else.
It was 2:00 in the morning. We're tired. We thought, "Why not refactor the entire code base?" Sounds like a splendid idea.
과적합 테스트가 구원이 되다 — green이면 원형에 가깝다
So, 2,700 commits later, close to a million lines of code change, touching 82% of the core code base, plugins were launched.
The night before, I think it was like 1:00 in the morning. I'm trying to go to sleep. And the tests are not passing. And I was like, was I Icarus and did I fly too close to the sun? ... did I vibe too hard? I actually generally thought I vibe too hard.
The saving grace was these awful sort of unit tests that AI code loves to generate that actually ended up over-fitting on our code. So, when we completely ripped everything out, we still had these tests that were like extremely over-fitting. And as long as they would go green, we knew we were kind of somewhat close.
공장 관리자의 스윔레인 — CI·기능·버그·릴리즈 레인
In my case, I call it my factory. It's many code sessions. ... Very simple. I have swim lanes. ... it could be five, it could be 10, it could be 20.
Imagine you're a factory manager and you have a production line below. ... you might have a case where you have ... CI to one side, you might have features on one side, you might have bugs on another. ... I just told them, "Take your time. Make sure the test pass. Just commit. Just push them through."
Maybe five is actually looking at new P0s and P1s. ... It might be using GitHub. We have agents that run inside of a Discord channel. So, when we do a release, we might be like, "Hey, what's happened in the last 2 hours that I need to be paying attention to?"
병목은 토큰이 아니다 — 컴퓨트, 뇌 용량, 그리고 worktree 지옥
What ends up becoming quite interesting is tokens are no longer the problem. ... depends who you ask. What really ends up becoming the problem is just raw compute and my brain space in order to sort of keep an eye on all of these sessions.
So, in Harness we trust. ... The one thing I have complicated in my life is adopting Git work trees and I kind of wish I hadn't. ... every PR I touch ends up becoming a new Git work tree. I end up with like something close like 70 or 80 active Git work trees in any given day on my machine and that's kind of hell.
Realistically, I should have adopted what Peter and other people do is just like clone the repo 10 times and point 10 different ... Codex sessions to each one.
The trick here is that like I haven't done any magical sauce. I don't use plan mode or spec mode. I have a conversation with the agent and we work through it and we find a way to make it work.
Reasoning token을 느낀다 — 빨간 드레스의 여인이 보일 때까지
If anyone's watched The Matrix and seen the scene where Neo goes over is like, "How do you know? How do you read the text?" And the guy's like, "Oh, you know, I've been doing this for a while, so I can see like woman in red dress or guy walking dog." And you start to have this like relationship where you can feel the reasoning tokens.
There's times where I'm looking at the swim lane. I'm like, "This sounds off." It doesn't sound off because of what it's doing. It sounds off because of how it's explaining itself to me. It's waffling. It's not making sense. It doesn't seem to know what it's doing.
This feels a lot like how I would manage people. ... if I had someone working for me and they started downright bullshitting, I'd be like, "Wait a minute. What's going on?" So, in these cases, I might just nuke the session.
That experience feels very much like intuitive and building that intuition, I've been able to get to because of the sheer volume of token maxing I've had to go through in the previous year.
Agent Development Environment — dot files처럼 관리하는 dot skills
There is engineering work. I call this the agent development environment. ... the process goes I have skills. I call it dot skills, similar to ... dot files. Both of my dot skills and dot files is available on GitHub. It's all open source. Go for it.
You could just say, "Go Codex." I've been using this skill in my last 2 weeks. Go through the Codex sessions, read the logs, make improvements to the skill.
I would then take that skill and deploy that into my open core or take that into my ... personal environment and I'll use something like vercel.skills.sh as like a mechanism ... there's a process to how I manage and maintain my skills as an engineer.
60k PR과 가짜 Slack — 노이즈를 신호로 읽는 법
There's this kind of running joke that every maintainer that joins the project decides to try and tackle like, "Oh my god, we have 6,000 PRs. How are we going to solve it? I'm going to cluster everything and like figure this out." 60k? How many? 60k. ... This is like a semantic graphing, vector embedding on the entire GitHub stuff. ... everyone else has the same problem, so they decide to send their flavor of the PR issue. Becomes utter noise.
This might be a signal for me to say, "Okay, if there's enough pressure coming on one issue, it must be big enough that all these other clankers decided it's a big problem. Maybe I should go and address it."
There is evals, surprisingly. ... after all this refactoring work, we decided to make a fake Slack of sorts with both synthetic models and real models, so we can run evaluation loops to check that each of the providers and the channels work.
"에이전트 10개를 어떻게 관리하나요" — 모델도 에이전트도 아닌 프로세스
"How do you manage 10 plus agents?" And this is something that you're thinking. I asked them back, "How do you manage 10 plus staff?" And they had no answer for me.
For me it was not like a new paradigm. But, I think for engineers and people working with these coding agents at scale, it's the soft skills that matter. It's how do you ask your agent, "What's going on?" How do you know when they're not bullshitting you? And how do you run that factory? So, it's no longer about the model or the agent. It's about the process.
2025 was about token maxing. 2026 is about not wasting them. It's about token efficiency. It's about agent in the loop.