frank@walsh:~/blog$ cat vibe-hardware.md

Vibe hardware is aimed at the wrong half of the job

Vibe hardware, from someone who designs boards with AI every day · 2026-10-04

[ note ] written in a personal capacity. Work examples are kept generic on purpose: no products, no internal systems by name. Nothing here is my employer's view.

On September 3rd OpenAI launched GPT-6 Astra, and one of the launch demos was fifteen seconds of the model doing PCB layout in KiCad: placing parts, routing copper, “turning an electronic schematic into a manufacturable PCB.” Within a week my feed was wall to wall circuit boards, from flight controllers to a star tracker the model ordered from JLCPCB itself, and the name everyone settled on was vibe hardware.

▸ tail -f vibe-hardware · what it built[1/16]
K
Kai Yang@ChihYang04

GPT-6 Astra doing PCB layout in KiCad 🤯🤯

Image from Kai Yang's post

965K views · Sep 3, 2026 · open on X ↗

// the board from OpenAI's launch video, the clip everyone shared. fifteen seconds, condensed from a 2 min 54 s run.

I design production hardware at a big company, and I use AI to do it every working day, so I have an opinion on all this. It isn't the one the feed suggests. The demos point the model at the one part of the job it's worst at, and the real gains are hiding in the less photogenic work around the board.

[ tl;dr ] the whole argument
  • 1.Harness and context decide almost everything. Everyone gets the same model. What you hand it is yours.
  • 2.Keep it close to text. KiCad over binary formats, CLIs over MCP, a CLAUDE.md that reads like a header file, and never a raw datasheet PDF.
  • 3.It's superb at translation. Flex pinouts, Gerber diffs, DFM triage: anything that moves information between tools that never talked.
  • 4.It's bad at layout. On real industrial boards the best model routes 12.6% of nets cleanly. Humans route 93.6%.
  • 5.Half of “taste” is physics nobody grades. EE needs a layout benchmark with field solvers in the loop.
[ info ] scope
who
a systems EE at a big hardware company, ex Tesla, using an agent every working day
covers
schematic, layout, and the work around them
skips
firmware, which has been covered to death
tools
Claude Code, KiCad at home, my employer's tools at work
shelf life
written October 2026. parts will be wrong by Christmas

## 1. Know what your model is good at

Two terms do most of the work in this post:

  • ▸Harness: everything wrapped around the model. The tools it can call, the files it can read, the instructions it starts with, the loop it runs in.
  • ▸Context: what's in front of it when it answers. Which schematic, which datasheet, which slice of the system. (Tobi Lütke and Andrej Karpathy made “context engineering” the name for this last year.)

Everyone gets the same model, more or less on the same day. The harness and the context are yours, though, and in my experience they account for nearly all of the difference between a result that saves you an afternoon and one that sends you back to the datasheet angry.

The best thing I've read on this is Claude-shaped science, by Matthew Schwartz, a Harvard physicist who spent months trying to get Claude to do physics the way he does physics. It didn't work, and the interesting part of the essay is what he did next:

> Instead of treating Claude like the collaborator I wanted it to be, I started to treat it like the collaborator it actually is.

He went looking for the problems that suited it, built a harness called BootLoops to steer it toward them, and came out the other side with thirty-six manuscripts. That is the whole skill in electrical engineering too. A model in 2026 is superhuman at some parts of board design and worse than an intern at others, and the job is knowing where that line sits and keeping the model on the right side of it. The catch is that the line moves every few months, so whatever you learn about it has a short shelf life.

## 2. Harness: get the model close to the files

Models are best at text. Every other format is a translation layer, and every translation costs you some accuracy and a lot of tokens. Three rules fall out of that, and they're the ones I'd give anyone setting up a harness for hardware work:

▸ three harness rules
rulewhyevidence
Store designs as textThe agent reads the schematic, netlist, and board directly, and can patch the tool itself.KiCad; i2cjak's Backplane shows every agent edit to a board live
Prefer CLIs to MCPEvery MCP server loads its tool definitions into context on every turn.Playwright MCP: 13.7k tokens before you've done anything. MCP costs 4–32× more tokens
Write CLAUDE.md like a header fileDeclarations stay in context; definitions stay on disk until something calls them.Anthropic's progressive disclosure; HumanLayer's “prefer pointers to copies”

If I were starting a hardware company tomorrow, I wouldn't pick an ECAD tool that stores designs as binaries, because every binary file is a wall between the model and your design. The honest counterexample is Eli Hughes, who wrote open-source parsers to crack Altium files open and feed the netlists to Claude and Codex for design reviews. So binary isn't a dead end. Someone just has to build the bridge before the model can cross it, and with KiCad the bridge is already there.

At the other extreme is atopile, which describes the whole circuit as code. When I was at Tesla there was a real push to use it, because we wanted to move as fast as humanly possible, and I understood the appeal. My problem with it was, and still is, that there's no schematic: you read code to find out how a buffer is hooked up. Models love that. Electronics engineers want a schematic, and I don't think that changes because the machine would prefer otherwise.

On CLIs: at work, a lot of my harness is small command-line tools the agent wrote for our internal web tools, things like the parts database, the issue tracker, and the place suppliers post their DFM comments. I describe what I want out of the tool, the agent builds the command, and I've stopped caring much how they work inside. Anthropic's own docs back the approach: “CLI tools are the most context-efficient way to interact with external services.”

And here's the header file. In embedded C, a header declares what exists and where it lives, and the definitions only get pulled in when something calls them. A CLAUDE.md should work the same way, and mine looks roughly like this:

▸ cat CLAUDE.md · declarations, not definitions
// always loaded. about forty lines. each one points somewhere.
parts:kb/parts/index.md // one line per part; open the one you need
datasheets:kb/datasheets/<mpn>.md // OCR'd markdown, never the PDF
system:kb/system/interconnect.md // what this board plugs into
tools:parts · tracker · gerber-diff // run --help to learn each one
rule:cite the datasheet page for every pin claim // the CSA lesson, below
// the files on the right stay on disk until a question needs them. a pointer costs one line.

One gotcha worth knowing: Claude Code's @path import loads the whole file at launch, which makes it an #include, not a pointer. A plain sentence saying where the file lives, and when to open it, does the job better and costs one line.

## 3. Context: the model is an expert in a box

The mistake I see most often, and I see it from good engineers, goes like this: paste a circuit question into a chat window, get a wrong answer, and decide the model is dumb.

Try this instead. You're a very good EE, and someone locks you in a box, slides a schematic you've never seen under the door, and asks whether the current-sense amp is hooked up right. There's no datasheet and no system diagram, and you have no idea what the board plugs into. You'd guess, you'd mostly be right because you're good, and some fraction of the time you'd be confidently wrong. That's the model, every time you ask it something cold. Anthropic's prompting guide has a politer version, “a brilliant but new employee who lacks context,” but the box is closer to how it looks from the model's side of the door.

I did exactly this to myself this week. I was designing a current-sense amplifier circuit and hadn't put the part's datasheet in my knowledge base yet, and the model told me, with total confidence, that the ground pin could go to a negative rail. It can't. The model didn't get any smarter between that answer and the right one; the only thing that changed was the datasheet.

So what does “context” actually mean for a board? In my experience it's four things, roughly in order of how often people forget them:

  • ▸Every datasheet on the sheet, as text. More on that below.
  • ▸The system around the board. Every connector, harness, and module on the other end. At Tesla I could export that from an in-house tool. Most places, it lives in your systems engineer's head, so go get it.
  • ▸History. Past issues, design rules, the last three revisions.
  • ▸An index on top, so it only opens what it needs. I use a version of Karpathy's LLM wiki. It works; it's also a chore to keep current.

My biggest pet peeve is the raw datasheet PDF dropped into a chat with “design me a circuit around this.” It feels like you've handed the model everything it needs, which is exactly what makes it a trap. Three things go wrong:

  • ▸It isn't really text. A PDF stores instructions for drawing glyphs. Reading order and tables have to be reconstructed (LlamaIndex explains why).
  • ▸It's expensive. Claude sees each page as an image plus text, 1,500 to 3,000 tokens a page. One datasheet can eat your context before you've asked anything.
  • ▸Agents grep it. The ones that shell out to a text extractor are searching a document that was never meant to be searched.

So I pre-digest every datasheet before the model ever sees it. Baidu's Unlimited-OCR runs locally on my laptop through MLX and handles the bulk of the text for free, and anything it isn't confident about goes to Claude, which is slower and costs money but can actually read a table. Ari Mahpour at Altium built something similar, which made me feel a little less crazy for doing it.

▸ ./datasheet-to-tokens.sh · top to bottom
1
datasheet.pdf
↓
2
Unlimited-OCR on my laptop (MLX)
reads every page; flags what it can't handle
↓
body text
written straight out as markdown
tables + figures it flagged
cropped to an image, sent to a headless claude -p, written back as markdown
↓
3
kb/datasheets/<mpn>.md
merged, greppable, a fraction of the tokens
↓
4
one new line in the index
so the model knows it exists
// the free local model does the bulk. the expensive one only sees the parts that need eyes.

None of this should be necessary forever. Datasheets are written for people, their main reader is quickly becoming a model, and somebody is going to fix the format. Here's where the vendors stand as of this fall:

▸ machine-readable datasheets, october 2026
whowhat exists
Microchipa free, public MCP server for its catalog
TIa JSON product API, approved customers only
ST, Infineon, Renesasno official server as of August
datasheets.mda startup re-extracting the PDFs
a standardnone for a whole datasheet, graphs and all

I'd bet heavily on that last row changing within a couple of years. Machine-readable datasheets were the backup idea on my YC application, and I still think every agent that touches hardware is going to need them.

## 4. What it's actually good at: translation

With the harness and the context in place, the thing that surprised me most isn't any single task. It's that the model is a universal adapter. Hardware engineering is full of files that were never designed to talk to each other (the schematic, the board, the mechanical model, the flex outline, the supplier's DFM report), and a big part of my day used to be carrying information between them by hand. The model can carry it, and it doesn't get bored.

The best example is a test flex I designed recently. One of our boards has a 36-signal board-to-board connector, and we wanted every signal broken out to 2.54 mm headers so we could probe them on the bench. The complication was that the reference flex has a bend and this one had to come out straight, so the pinout couldn't just be copied across. I gave the agent the reference pinout and the four-layer FCCL stackup and let it go.

It assigned every net, planned the fanout, drew the outline in Matplotlib, and exported it as IDX for the mechanical side. (Drawing a board outline turns out to be the same problem as drawing an SVG, which these models are extremely good at.) I checked every pin by hand afterwards and changed nothing. It isn't fabbed yet, so the bench gets the last word, but an afternoon of careful, boring work became one prompt and a review.

▸ translations I run every week
fromtowhat it catches
reference flex pinoutnew flex breakout + IDX outlinean afternoon of manual pin mapping
board → flex → connector → boardinterconnect table, checked in 3Dmirrored footprints, flips through a bend, pin 1 landing on pin 36
Gerber rev A + rev Ban overlay diffwhat actually moved, without flicking between windows
the layouta picture for the mechanical engineereverything grayed out except the one thing they need
supplier DFM commentsa triage list against our guidelineswhich comments are real. most aren't
schematican Excel quick-start calculatorthe design math, in a format every engineer can poke at

The honest gap is simulation. Even with text netlists, I haven't had good results getting it to build and run LTspice models, though I also haven't put real harness work into that, so some of the blame is mine. Everything else I use it for is below, sorted by where it sits in the job. Steal whatever's useful.

▸Everything I use it for23 uses across office work, parts, schematic, layout, and manufacturing, each ratedexpand
## office and search
issue-tracker archaeologysomeone hit this three years ago on another program. the agent finds it with a query instead of me with a memorygreat
chat historysummarize the thread I just got pulled into; have we seen this before in the tool support channelsgreat
email and chat draftsboring, worksgreat
status decks and design reviewspoint it at a project's chat history and files. prompt: fewer words, more visuals, keep the detailgood
questions over the KBwhat's the SWD pinout; what's the input leakage; table these op amps by offset voltagegreat
## parts and datasheets
parts database searchfind an op amp that meets the spec, has an active lifecycle, and is already used in another product. subagents crawl, I pickgreat
datasheets to markdownthe OCR pipeline abovegood
standardsmany are already in the weights. the rest, once out of PDF, are easy to querygood
## schematic
quick-start calculatorsscreenshot or netlist of the skeleton, then an Excel model of the circuit to play with. Excel because every other engineer already knows how to use onegreat
format translationschematic to SPICE netlist, netlist to spreadsheet, and backgood
symbols and footprintsabout 80% of the way from the datasheet. you check the restgood
ERC setup and runsnatural language in, rules out, then it runs themgood
DFMEA, FMEA, HARAall text. with system and board context, close to automatic. this is what sev10 is forgreat
SPICE simulationhasn't gone well for me in LTspice. I also haven't put real harness work into itmeh
## layout and mechanical
flex fanout planningpinout copied from a reference flex, breakout planned, outline drawn in Matplotlib, exported as IDXgreat
pin-1 interconnect chainsboard to flex to connector to board, checked in 3D through the bendsgreat
Gerber diffsoverlay two revisions and report what moved. it has every coordinategreat
pictures for mechanical engineersthe layout, grayed out except the one thing they need to seegreat
design rulesfrom a standard or the fab's capability sheetgood
STEP fit checksboard in enclosure, at home. cuts out a round trip to an MEgood
placement and routingsee section fivemeh
## manufacturing and debug
DFM triagepull the supplier's comments from the tracker, check each against the design files and our guidelines, flag the few that need a humangreat
log analysisupdate and flashing failures on a large embedded product. paste the logs, give it the contextgreat
// steal anything. the verdicts are mine, as of this month.

## 5. It can rotate a shape. It can't route a board.

Which brings us back to the feed. The thing everyone is excited about is layout, and layout is exactly where the models are weakest.

I used to say LLMs aren't shape rotators, and it turns out that's too blunt. They've gotten very good at flat, discrete spatial problems, the kind of puzzle you can write down as a grid. What they still can't do is continuous geometry in three dimensions, or hundreds of interacting constraints spread over a large area at once, and that happens to be a pretty good description of a circuit board.

▸ spatial.log · best model on each test vs the human baseline
[can] abstract grid puzzlesARC-AGI-3 ↗
solvable
100%
model
99.9%
GPT-6 Astra · Sep 2026 · 100% means every game beaten as efficiently as a human
[can] 2D mental rotationSpatialViz ↗
human
90%
model
91.3%
GPT-5 · Dec 2025
[can't] 3D mental rotationSpatialViz ↗
human
79.2%
model
33.8%
GPT-5-mini · Dec 2025 · chance is 25%
[can't] multi-view spatial reasoningMMSI-Bench ↗
human
97.2%
model
45.2%
Gemini 3 Pro · 2026, third-party eval
[can't] routing real boards, nets DRC-cleanOmniRouting ↗
human
93.6%
model
12.6%
best model, no tools · Aug 2026
// each row is a different test, model, and date; there are no 2026 numbers for most of them yet. the pattern is what matters: flat and discrete is solved, 3D and routing are not.

The board-specific numbers are worse. OmniRouting took 1,681 real industrial boards, each with a placement that engineers had already proven routable, and asked models to finish the job:

▸ OmniRouting, Aug 2026: share of nets connected and DRC-clean
who routed itnets clean
human engineers93.6%
a classic algorithmic router (PcbRouter)56.2%
best model, with every tool28.0%
best model, on its own12.6%
// the models also routed ground as ordinary traces instead of pours.

PCBWorld found the same shape: a GPT-5.4 agent cleanly routed 65% of small real boards and none of the medium ones. A tiny RL policy trained only against a DRC checker beat it on both. That last part matters, and I'll come back to it: the small model won because something was grading it.

As far as I can tell there are two reasons, one about how the model sees the board and one about how it remembers it:

  • ▸It can't see the copper. Images arrive as 28-pixel patches. The arithmetic is below.
  • ▸It has no picture to update. The board exists to the model as a list of coordinates, and nothing redraws a mental map when one moves. The Spatial Competence Benchmark calls the result “locally plausible geometry that breaks global constraints.” Every segment looks fine. The board is shorted.
▸ what one visual token sees
stepvalue
the imagea 100 mm board, shown edge to edge at 2576 px
one patch28 × 28 px (Claude's vision docs)
so one token coversabout 1.1 × 1.1 mm of board
inside that squarea 0.1 mm trace, its clearance, and a whole 0.4 mm-pitch BGA ball

The board everyone piled on was Microduck, a four-layer robot board that Astra “one-shotted.” Here's the layout from the original post, so you can play spot-the-problem before reading the replies:

▸ view x-microduck.jpg · open the post on X ↗KiCad layout of the Microduck robot board, routed by GPT-6 Astra
// the Microduck board, routed by GPT-6 Astra in KiCad. look at the screw terminals along both edges and the connectors along the top.
▸ what experienced EEs said, with links
whowhat they called outabout
i2cjakconnectors you can't physically reach; watch the model "like a hawk" so it doesn't paint itself into a cornerMicroduck
BlindViascrew terminals blocked so you can't get a wire inMicroduck
Luke Westonthe USB-C connector, the FFC connector, and signal integrity on the camera diff pairsMicroduck
DeepPCB9 of 32 vias in pads, five track widths under one empty net class, a split power planeMicroduck
Michael W.USB D+/D- routed terribly, a GPIO expander added to dodge routing, "the most awful buck converter layout I have ever seen"his own AI-routed board
kelina Shenzhen PM whose new clients are software engineers with vibe-designed boards that don't workthe fab's view
// all of these are also in the carousel at the top, in amber.

That's only what shows up in a screenshot, too. Most of these boards would probably work on a bench, which is all a demo needs. Start thinking about ESD, EMC, SI and PI, or about building ten thousand of them, and they fall apart. The buck converters are the worst offenders: sprawling hot loops, giant switch nodes, and inductors on the far side of the board from the IC. I'm surprised some of them turn on.

What I'd do instead is let the model drive the tool that was actually built for geometry, and keep it on the parts of the job it's good at:

  1. 1.The model reads the stackup, the datasheets, and the fab's capability sheet.
  2. 2.It writes the net classes, widths, clearances, and layer rules. Nobody likes doing this, which is why autorouters get a bad name.
  3. 3.The autorouter does the geometry. Given good constraints it will escape a BGA with dogbones on the layers you pick (there's an old EEVblog video on exactly this).
  4. 4.The model checks DRC and diffs the result; a human routes or reviews the critical nets.

JLCPCB's review of an Astra board suggests the model already does a version of this: it hand-routed the power and switching nets and gave the rest to Freerouting. I haven't run this loop end to end myself yet, and it's the next thing I want to try.

## 6. Taste is scar tissue

The usual answer to all of this is that the models lack taste. I think that's half right, so let me start with the half that is.

Taste comes from doing it wrong once. A board fails EMC, and you go to the chamber, and you spend days learning about component orientation, loop area, and the geometry of fields until you find the thing. After that you look for that extreme on every board you see. A fab explains what an unbalanced stackup does in reflow, and from then on you can glance at a layout and tell whether the designer was thinking about how much copper the acid would take off each layer. None of it came from school.

Here's the same board laid out both ways, with the six places I look in the first thirty seconds. Flip the toggle and watch the same six spots change.

▸ view thirty-seconds.kicad_pcb
GND pourJ1 USB-CD1U1 MCUU2CinL1CoutJ2stackup, side viewL1 52% copperL2 94% copperL3 90% copperL4 48% coppermirrors match:etches evenly,stays flatbuck hot loopreturn currenttop copperbottom copper123456
1ESD at the connectorThe TVS diode (D1) sits on the USB lines right at the connector, so a zap is clamped before it goes anywhere.
2Buck input loopInput cap tight against U2. The hot loop (amber) is tiny.
3Switch nodeInductor right next to the switch pin. The switch node is a small copper island.
4PlacementDecoupling caps hug every side of U1, and the passives sit in rows with one orientation. Someone thought about pick-and-place and rework.
5Ground returnSolid ground under the USB pair. The return current (cyan) runs straight back underneath it.
6Copper balanceEach layer roughly matches its mirror in the stackup (L1↔L4, L2↔L3). It etches evenly and stays flat in reflow.
// an illustration, not a real design. same netlist both ways; both connect and both would pass a basic DRC. flip it and look at the same six spots.

A newbie can look at a board from a top-tier company and see that it's good, but can't tell you why. The model is in the opposite position. It has the words: the loop-area rules, the EMC textbooks, and every app note on buck layout ever written are almost certainly in the weights. What it doesn't have is the analogies, the reflex that says this looks like the board that failed in the chamber last spring. It designs from first principles every time, because it's still in the box.

## 7. Half of taste is physics nobody grades

Now the half I don't buy. Look at those six callouts again, and notice that most of them aren't really taste at all. They're physics with a number attached:

▸ what an EE glances at vs what it actually is
the glancethe physicswhat measures it
① ESD placementinductance of the clamp pathpartly rules, partly extraction
② buck input loophot-loop inductanceparasitic extraction
③ switch node sizeradiating copper area at the switching edgeEM solver
④ decoupling placementPDN impedance vs frequencyPDN analysis
⑤ slot in the groundreturn-path discontinuitySI field solve
⑥ copper balancecopper area ratio per layera twenty-line script
// we call it taste because nobody can run a field solver in their head, so we compress years of results into a glance.

Models get good at whatever can be checked automatically, which is why they got good at code first: you can run the tests, and the tests don't care how confident the model sounded. So it's worth asking what actually gets checked in EE today:

▸ what AI benchmarks grade today, easiest first
checkgraded byopen, headless tool that could grade it
every net connectedOmniRouting, PCBWorld, every demokicad-cli
DRC / ERC cleanOmniRouting, PCBWorld, every demokicad-cli
circuit works in SPICE at tolerance cornersEEBench (schematic only)ngspice
copper balance per layernobodykicad-cli exports + a script
PDN impedance, DC IR dropnobodyElmer, ngspice
impedance, crosstalk, return pathsnobodyopenEMS, gerber2ems
radiated emissionsnobodyopenEMS (slow)
works at ten thousand unitsnobodynone. only reality grades this
// fab DRC and commercial SI/PI tools check several of these for human designers. nobody scores a model on them.

EEBench, from the atopile team, is the best EE benchmark there is: real parts, ngspice at tolerance corners, BOM cost, no LLM judge. Claude Opus 5.5 leads at 75%. But its methodology puts layout out of scope, and the routing benchmarks only check connectivity and DRC.

I couldn't find a single published example of a language model laying out a board against a PDN or field-solver reward. Which means the Astra demo was optimized for exactly what it showed, a board that connects, because that's the only thing anyone is grading. Remember the PCBWorld result: the small model beat the big one because something was checking its work.

I can't see the labs building this on their own. Field solvers are slow, the good commercial ones are license-gated, and EE is a small market next to code. The models would also need to get much better at physics before the scores meant anything: HWE-Bench, which asks for board-level schematics from scratch and checks them in simulation, tops out at 8%.

So here's my ask. Look at the right-hand column of that table: every “nobody” row already has an open-source tool that runs headless. Someone should wire them into one benchmark that scores a layout on copper balance, PDN impedance, IR drop, and return paths, on top of DRC. atopile already showed that if you build the environment, labs will train on it.

Build the one that measures the part we actually ship. If you're building it, I'd like to help.

## Two boxes

Here's where I land, for now. Vibe hardware is real, and it's going to be huge, mostly for people next to electrical engineering: firmware, software, and mechanical engineers who need a bench tool and no longer have to ask someone like me for one. That's great. The serious end will look more like software did. Anyone can build an app now, but very few people can run a platform for millions of users, and the ones who do use AI to go faster, not to go away.

Maybe I'm biased; I've spent my whole adult life getting good at this. But the difference from software is the verify loop. Code fails in seconds. A board fails in the chamber, six weeks after you sent it out. Until that loop is something a model can run, it will keep learning the parts of the job that can be checked, and those aren't the hard ones.

There are two boxes in this post. The first is the one we put the model in every time we ask it something cold, and that one's easy to open: hand it the datasheet, the system, the pointers. The second is the one it was trained in, where the only thing anyone grades is whether the board connects. That box has to be opened from the outside, with a field solver. Until someone does, I'll keep doing the layout, and let it do everything else.

▸ cat sources.md · every link in this post · click to expand