Twelve hundred and forty-seven people applied to this thing. They took fifty. I still don't totally know how I landed in that fifty, and that feeling ended up mattering more than I expected, so let me start there.
- event
- World's Largest Hermes Buildathon
- host
- GrowthX
- where
- San Francisco
- when
- July 11, 2026 · 8-hour sprint
- track
- AI as an Agency (of 3: Virality · Revenue · Agency)
- scale
- 1,247 applied · 50 built
- built on
- Hermes Agent (Nous Research) + OpenAI
- result
- did not place — and worth every minute
- links
- live build ↗demo ↗event ↗handbook ↗repo ↗
The night before, the organizers dropped a message with the stats on the room. Median 8.3 years of experience. 31% active founders, a lot of them ex-FAANG or YC. Nearly half were engineers, half of those staff or above. Sixty percent already building on LLMs day to day. I read that on my couch after my kid went down and thought, pretty clearly, that I did not belong in that room. I'm a hardware guy. I design electronics. I don't ship software for a living, I've never broken four figures in MRR, and I've never raised a dollar. The next morning I met someone doing over a million in ARR who, when I asked him about it, was so relaxed and humble about the whole thing that it kind of reset my idea of what confident looks like. He wasn't there to win anything. He just wanted to build something fun and meet some people.

I did not have that energy. I had something to prove. Hold that thought, because it's the whole story.

## What I built
The pitch, roughly how I gave it to the judges: builders are too busy shipping to keep up with frontier research, so I built a research agency staffed by AI agents. You point it at your repo, the agents pull the newest papers, pick the ones that actually fit your code, run real experiments in the cloud, and open pull requests with measured performance gains attached.
to keep up with AI, you need to be unemployed
499.7K views · Feb 26, 2026
I was on the “AI as an Agency” track, one of three. The other two were Virality and Revenue. The track is what it sounds like: a team of agents standing in for a whole human function. A manager agent that plans, specialists that execute, work handed between them, memory that persists. My system mapped onto that cleanly, and the part I'm proud of is that none of it was faked. It ran real experiments and opened two real PRs against a real benchmark.

## They literally gave us the answer key
Here's the thing that still gets me. GrowthX published the entire scoring rubric ahead of time. Points equal (level minus one) times a weight, every parameter scored one to five. The root parameter, “a working product shipping real output,” was worth up to 80 and kept paying past the ceiling. Observability was 28 points, and the handbook flat out said most teams skip it, so, free points. Then the partner power-ups: wire up each of the six sponsors for real, 25 apiece, all six worth 150. A hundred and fifty. That's almost the entire technical base of the rubric, sitting there for anyone willing to integrate some tools and screenshot the dashboards.

I knew all of this. I read the handbook twice. And I still walked in and pointed almost all of my energy at the hardest, deepest 80-point parameter, because the imposter thing had me trying to out-engineer the room.
## What actually went well
Not all of it was a lesson in humility. A few evenings that week, after bedtime, I knocked out the two things most likely to kill me on the day. I found an agent benchmark light enough to run over and over in the cloud, cheap and fast, and I proved out that I could run the whole thing inside Cloudflare containers before I ever showed up. When the clock started, those weren't question marks anymore. That prep is the only reason I had anything real to show.

I also finally learned what “observability” means, and it happened in the least glamorous way possible. Deep in the crunch, a Cloudflare mentor said the word to me and it just clicked. It's being able to see what your system is doing while it runs in the cloud, instead of staring at a black box and hoping. I'd never run real infrastructure before, so this was genuinely new to me. Watching each agent, the cost of every step, where things stall, turns out to be its own discipline, and you don't have to build it from scratch. There are tools for exactly this. Sounds obvious written down. It changed how I think about anything that runs on its own.
On top of that I got a lot better at pitching something technical. My last hackathon project fit in a sentence. This one didn't. So I built the explanation like pulling focus: open on the most abstract version, what goes in, what happens, what comes out, then slowly zoom into the stack. The old founder framing still carried it. Here's a specific pain, here's how I take it away.
## Where I got it wrong
I broke my own number one rule, which is get to a working demo first and polish later. I know this rule. I've won on this rule. But intimidation is a terrible project manager, and mine kept telling me to build something more ambitious to justify my seat. So I over-scoped.
It cost me in the most concrete way possible. That 150 in power-ups needed a submission page with screenshots proving I'd used each tool. I was assembling that page at 3:58 for a 4:00 deadline and had to skip it entirely just to get my submission in on time. I left the biggest, easiest bucket of points on the table. Not because I couldn't do it, but because I spent my hours on the hardest thing instead of the smartest thing. My pitch was under-rehearsed for the same reason. Too much scope, no minutes left to practice.
The other miss was the idea. I came in married to the research-agency concept. But two of the three tracks were Virality and Revenue, and the whole room was chasing traction and money in eight hours. There were moments I could've read that and bent toward something simpler and sharper. I didn't. Falling for your solution instead of the problem is the oldest mistake there is, and I made it in front of exactly the crowd you'd least want to make it in front of.
## Reading the room
This is the part I keep coming back to. Three of the five finalist demos were dialed straight into what people want, and they proved it in real time. One was an end-to-end clipper that raced to publish the first clip of a livestream, since the first clip is the one that goes viral, and it pulled 600-plus views on TikTok during the event. Wild.
But the one that really got me: somebody took r/AmIOverreacting, that subreddit where people post a screenshot of a fight and ask strangers to judge who started it, and turned it into an app you drop straight into your group chat. That's such a good read on how people actually behave. Obvious the second you see it, invisible until someone builds it.
For years I'd watch VCs tweet about founders who just get what people want, and I never really understood what they meant by it. I think I do now. It is genuinely hard to unhook yourself from the thing you want to build and look honestly at the thing people would actually use. That gap is the whole skill. None of these demos were the most technically complicated builds in the room. They were the most aware. I optimized for depth. They optimized for whether anyone would care, and the day rewarded theirs.
## Is it even a good product?
Here's the harder question, honestly asked. Is this even a good product? I think the need is real. There are founders, small teams, and solo builders who genuinely can't keep up with how fast AI research moves, and an auto-researcher that reads new papers and turns them into working improvements would help those people a lot. The field matters and the pain is real.
I just don't think what I built is the right shape for it. The tools are right. Cloudflare is a great home for this. Hermes is genuinely impressive. And there's something special in an assistant with persistent memory that dreams between runs, builds its own skills, and gets better at researching over time. Pulling papers nightly, writing and testing code in the cloud, cross-referencing what it read against what it actually shipped. You point it at your repo and let it run forever. Run it nightly against a sales agent you want to push further, or a model you're fine-tuning against some in-house benchmark, and let the improvements stack up.
But the form factor I demoed isn't it. I built a whole GUI, a kanban board with cards sliding across columns, and if I'm honest, most of that was there so I could wire in the Convex sponsor. In real life this thing wants to be a CLI, or an MCP with a thin dashboard you can log into to check on it. Something you point and forget. The catch is that a bare command line doesn't give you great steering over an agent loop like this, and steering matters here, so I don't fully know where it lands. That's the part I'm still chewing on. What I am sure of is that the auto-researcher is a real and important idea, that builders would benefit from it, and that I brought the wrong container for it to a room that was grading the container.

## The small, human stuff
For the record, and because I do actually replay these days: carry mints. I ran on diet energy drinks the whole time and my breath paid for it, which is not the edge you want when you're pitching two feet from a judge. I laughed when I clocked it, then wrote it down, because that's the point of going back through the tape.
So where do I land. I think I built something legitimately good. A real system, real experiments, real output, solo, on a clock, on a stack I'd mostly just learned. The quality wasn't the problem. I brought the right build to the wrong scoreboard, and I let feeling out of place push me toward proving instead of winning. I can do the hard technical thing. The muscle I need to grow is the other one: reading a room, picking the metric that matters, and being honest about what people actually want instead of what I feel like making. I know which of those I'm good at, and I know which one I owe myself. That's worth more to me than the trophy would've been.
See you at the next one.