<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Ben Vierck]]></title><description><![CDATA[Essays from a practitioner on what the machine does to work, prices, and the people in between.]]></description><link>https://essays.xcud.com</link><image><url>https://essays.xcud.com/img/substack.png</url><title>Ben Vierck</title><link>https://essays.xcud.com</link></image><generator>Substack</generator><lastBuildDate>Fri, 04 Sep 2026 20:04:22 GMT</lastBuildDate><atom:link href="https://essays.xcud.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Ben]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[xcud@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[xcud@substack.com]]></itunes:email><itunes:name><![CDATA[Ben Vierck]]></itunes:name></itunes:owner><itunes:author><![CDATA[Ben Vierck]]></itunes:author><googleplay:owner><![CDATA[xcud@substack.com]]></googleplay:owner><googleplay:email><![CDATA[xcud@substack.com]]></googleplay:email><googleplay:author><![CDATA[Ben Vierck]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Last Wide Rung]]></title><description><![CDATA[For about fifty years, an ordinary person could climb into the middle class on nothing but a trained mind. Knowledge work was the last wide rung on that ladder. Here's why, this time, no wider floor is opening to catch the fall.]]></description><link>https://essays.xcud.com/p/the-last-wide-rung</link><guid isPermaLink="false">https://essays.xcud.com/p/the-last-wide-rung</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Mon, 29 Jun 2026 14:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oSFH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You can already see it, if you look. Hiring freezes that never quite thaw. Entry-level jobs that quietly stop being refilled. A new graduate with a sharp degree and three hundred unanswered applications, who can't work out why the ticket stopped working. She is standing at the end of a fifty-year deal, and she is not the only one.</p><p>For about fifty years, an ordinary person in this country could climb into the middle class on nothing but a trained mind. No capital. No inheritance. No land, no machine, no family name to trade on. Just study and work. You learned something difficult, someone paid you to do it, and that paycheck was your <em>claim on the nation's wealth</em> &#8212; the reason the economy had to deal you in at all.</p><p>It is the best deal the modern economy ever offered an ordinary person &#8212; a share of the country's wealth for nothing but what you could learn. And it is ending.</p><p>We have watched the machine come for people's work twice before, and both times it ended well. The first wave &#8212; the Industrial Revolution &#8212; drove the farmhand off the land, and the factory was waiting to catch him. The second &#8212; automation, and the great wave of globalization &#8212; emptied that factory, and the office was waiting to catch his children. A way of making a living ended, and each time a wider one opened beneath it. That is the pattern, and it is real.</p><h2>We've been here before</h2><p>Farm to factory. In 1900, roughly four in ten Americans worked the land<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>. And a whole village made its living in orbit around them &#8212; the blacksmith at his forge, the miller, the wheelwright. It was a world built on muscle, the farmer's and the horse's, and nearly every trade in town existed to keep that muscle working.</p><p>Then the machine arrived. The tractor, the combine &#8212; one man on a seat doing the work that had taken a dozen backs and a team of horses. The field didn't need those dozen anymore. It needed just one.</p><p>If you'd stood at the edge of that field and asked, "Where do all these people go?" &#8212; nobody could have told you. There was no answer to give. And yet they went somewhere. They drifted toward the cities, toward the smokestacks, and within a generation they'd become something their grandfathers had no word for: riveters and welders and machinists. The work was steadier than the harvest, and often <em>better</em> paid than the field had ever been.</p><p>The wide floor of the field emptied, and a wide factory floor opened to catch it.</p><p>Factory to office. In 1950, roughly three in ten Americans made their living in a factory &#8212; the riveter, the welder, the machinist. And a whole town made its living in orbit around them &#8212; the diner, the union hall, the corner store. It was a world built on the line, on the steady wage and the lunch whistle, and nearly every business on Main Street existed to spend what that wage brought home.</p><p>Then the machine arrived. The robot arm, the shipping container &#8212; an arm that welded all night without a wage, and a steel box that could carry the whole job overseas. The line didn't need those hands anymore. The ones it kept, it paid less.</p><p>If you'd stood on that factory floor and asked, "Where do all these people go?" &#8212; nobody could have told you. There was no answer to give. And yet they went somewhere. They drifted from the line into the office, out of the noise and into the fluorescent hum, and within a generation they'd become something their fathers had no word for: clerks and bookkeepers and analysts. The work was cleaner than the line, and often <em>better</em> paid than it had ever been.</p><p>The wide factory floor emptied, and a wide office floor opened to catch it.</p><p>Today, more Americans make their living with a trained mind than ever before &#8212; the analyst, the coder, the paralegal, the copywriter. And a whole economy makes its living in orbit around them &#8212; the software vendor, the office tower, the downtown lunch counter. It's a world built on the credential and the salary, and nearly every business downstream of it lives on what those salaries spend.</p><p>Then the machine arrived. The model, the agent &#8212; a mind that never sleeps, that reads every document and drafts every memo for the price of the electricity it burns. The office doesn't need those hundred anymore. It needs a few.</p><p>Stand in that office today and ask, "Where do all these people go?" &#8212; and nobody can tell you. There is no answer to give. We say they'll go somewhere; they always have. They'll drift out of the cubicle and into &#8212;</p><p>The wide office floor is emptying.</p><p>Into what?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oSFH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oSFH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png 424w, https://substackcdn.com/image/fetch/$s_!oSFH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png 848w, https://substackcdn.com/image/fetch/$s_!oSFH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png 1272w, https://substackcdn.com/image/fetch/$s_!oSFH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oSFH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png" width="1456" height="618" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:618,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A lone office worker steps off the edge of a raised floor; behind him stretch the eras of work &#8212; a tractor in a field, a factory, rows of office desks &#8212; and ahead of him there is nothing but empty space&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A lone office worker steps off the edge of a raised floor; behind him stretch the eras of work &#8212; a tractor in a field, a factory, rows of office desks &#8212; and ahead of him there is nothing but empty space" title="A lone office worker steps off the edge of a raised floor; behind him stretch the eras of work &#8212; a tractor in a field, a factory, rows of office desks &#8212; and ahead of him there is nothing but empty space" srcset="https://substackcdn.com/image/fetch/$s_!oSFH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png 424w, https://substackcdn.com/image/fetch/$s_!oSFH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png 848w, https://substackcdn.com/image/fetch/$s_!oSFH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png 1272w, https://substackcdn.com/image/fetch/$s_!oSFH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5383dbce-4778-4dc9-bc56-a88d9b8a1e41_3168x1344.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Follow the windfall</h2><p>There are only two reflexive answers to <em>Into what?</em></p><p>The pessimist says <em>nowhere.</em> This is the end of work &#8212; the machines have finally won. It's an old fear, and it has an old name: the <strong>lump-of-labor fallacy</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>, the belief that there's a <em>fixed pile</em> of work in the world, so that every job a machine takes is one gone for good. Every generation reaches for some version of it; the Luddites smashed the looms over it. And so far, it has been wrong every time &#8212; because the pile was never fixed.</p><p>The optimist knows this, and says <em>somewhere new.</em> There has always been somewhere new, and he has two hundred years of being right. The work didn't run out when the loom came, or the tractor, or the assembly line &#8212; it grew, and the displaced always found a wider floor waiting below. So when he tells you the office will empty into something we can't yet picture, don't wave him off. History is sitting in his lap.</p><p>The people always followed the windfall. That's the thing both of them skate past, and it's what every upheaval in this essay quietly teaches: find where the windfall pooled, and you've found where the next floor opened. So before we bet on the optimist's streak holding one more time, let's ask the question neither side ever does &#8212; <em>where does the windfall go?</em></p><p>So follow it. When the tractor made food cheap, it created a <strong>windfall</strong> &#8212; the gap between what food used to cost and what it now cost. Where did that windfall go?</p><p>Not to the farmer &#8212; not for long. A little stuck to whoever mechanized first, but competition pried it loose fast. It went <em>downstream,</em> to everyone, in a form so ordinary we don't call it wealth: a smaller bill at the grocer. Economists have a lovely name for that one, too &#8212; <strong>consumer surplus</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>, It's the money that <em>stays in your pocket</em> when something you need gets cheaper. Bread drops from a dime to a nickel, and you didn't earn a thing &#8212; but you have that nickel, every day, for the rest of your life. Multiply it across a nation and it's a fortune, hiding in plain sight as the <em>absence of a cost.</em></p><p>That pile of nickels didn't sit still. The country spent it &#8212; on things that had scarcely existed while bread was dear: a Ford in the driveway, a radio in the parlor, a Sunday at the pictures. And every one of those new comforts had to be <em>built,</em> by <em>someone.</em> The grocery money the tractor freed up became the wage that built the cars and the radios and the movie houses &#8212; and the farmhand walked off the emptying field and into the very factory his cheaper bread had paid to build.</p><p>Then it happened again, a generation later. The assembly line and the shipping container did to <em>goods</em> what the tractor had done to food &#8212; they made them cheap. You watched the last act of it yourself: a factory in China, a container ship, a Walmart at the edge of town &#8212; and a refrigerator, a flat-screen television, a season's worth of clothes cost a fraction of what they once had. The difference dropped back into the family's pocket. The same nickel off the bread, now a few dollars off nearly everything on the shelf.</p><p>And that money didn't sit still either. With the necessities cheap and a little left to spare, families started buying things they couldn't hold: a doctor's check-up, a mortgage on a house, a degree with a name on it, a song they didn't have to sing themselves. Services &#8212; and a service has to be <em>rendered,</em> by <em>someone.</em> So the savings off the cheap goods became the salaries of the people who rendered them, and every year there were more of them: more doctors, more lawyers, more accountants, more programmers. The factory worker rarely made that leap himself &#8212; his children did. They took the degree the surplus had paid for and walked into the office it had built.</p><p>Five links, every time:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EtsR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65aebd3d-8af2-483c-a6fa-ff38443ea913_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EtsR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65aebd3d-8af2-483c-a6fa-ff38443ea913_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!EtsR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65aebd3d-8af2-483c-a6fa-ff38443ea913_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!EtsR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65aebd3d-8af2-483c-a6fa-ff38443ea913_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!EtsR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65aebd3d-8af2-483c-a6fa-ff38443ea913_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EtsR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65aebd3d-8af2-483c-a6fa-ff38443ea913_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/65aebd3d-8af2-483c-a6fa-ff38443ea913_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A five-link iron chain, every link whole and interlocked, with a golden stream of value flowing through all of them and off the end&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A five-link iron chain, every link whole and interlocked, with a golden stream of value flowing through all of them and off the end" title="A five-link iron chain, every link whole and interlocked, with a golden stream of value flowing through all of them and off the end" srcset="https://substackcdn.com/image/fetch/$s_!EtsR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65aebd3d-8af2-483c-a6fa-ff38443ea913_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!EtsR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65aebd3d-8af2-483c-a6fa-ff38443ea913_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!EtsR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65aebd3d-8af2-483c-a6fa-ff38443ea913_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!EtsR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65aebd3d-8af2-483c-a6fa-ff38443ea913_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>windfall &#8594; consumer surplus &#8594; re-spent &#8594; new demand &#8594; new production that needs people.</strong></p><p>So let's follow the value into our own predicament &#8212; link by link.</p><p>An AI reviews the contract, drafts the design, does the week's work in an afternoon. There's the <strong>windfall.</strong> It runs downstream, exactly as always, to whoever buys the result a little cheaper: <strong>consumer surplus,</strong> the same nickel off the bread, only now the bread is a legal contract, a tax return, a company logo.</p><p>And it will keep flowing, exactly as it always has. That money will get <strong>re-spent</strong> &#8212; it always is. The spending will call up <strong>new demand,</strong> new wants, new work to be done. Four links of the old chain, clicking into place right on schedule. And every instinct we have says the fifth is coming up behind them, the way it always has: the new demand becomes new jobs, and a floor opens under the falling.</p><p>But what if the new demand can be satisfied by the same machine that created the windfall?</p><p>Every time before, it couldn't. The tractor that emptied the field could not build the Ford. The robot arm couldn't analyze your mammogram. The factory in China couldn't write your contract.</p><p>What faces us now is different in kind: it is <em>general.</em> It does the contract, the design, the week's work in an afternoon &#8212; and it can turn and make whatever new thing the windfall dreams up next.</p><p>So the fifth link fails: <em>new production no longer needs people.</em> The one condition that saved us every time is the thing the machine quietly removes.</p><p>The sharpest optimists have one card left, and it has a name: the <strong>Jevons paradox</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>. Make a thing cheaper, a Victorian economist noticed, and the world doesn't use less of it &#8212; it uses far more. Make code cheap and we won't write less software; we'll write oceans of it. They're right, and it's their best move. But watch where it lands. The demand for the <em>work</em> explodes &#8212; and the machine is what rushes in to meet it. More software than the world has ever seen, built by fewer people than ever. The paradox holds for the output and breaks for the worker.</p><p>Which is the whole of it, in a sentence:</p><blockquote><p><strong>The value still flows somewhere we can spend it. It no longer flows somewhere we can earn it.</strong></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!03Dd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff56a1fe9-82a3-4e06-86f9-d11bbef416f3_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!03Dd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff56a1fe9-82a3-4e06-86f9-d11bbef416f3_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!03Dd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff56a1fe9-82a3-4e06-86f9-d11bbef416f3_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!03Dd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff56a1fe9-82a3-4e06-86f9-d11bbef416f3_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!03Dd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff56a1fe9-82a3-4e06-86f9-d11bbef416f3_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!03Dd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff56a1fe9-82a3-4e06-86f9-d11bbef416f3_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f56a1fe9-82a3-4e06-86f9-d11bbef416f3_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The same five-link chain, but the fifth link is snapped open and the golden value pours out at the break&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The same five-link chain, but the fifth link is snapped open and the golden value pours out at the break" title="The same five-link chain, but the fifth link is snapped open and the golden value pours out at the break" srcset="https://substackcdn.com/image/fetch/$s_!03Dd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff56a1fe9-82a3-4e06-86f9-d11bbef416f3_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!03Dd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff56a1fe9-82a3-4e06-86f9-d11bbef416f3_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!03Dd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff56a1fe9-82a3-4e06-86f9-d11bbef416f3_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!03Dd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff56a1fe9-82a3-4e06-86f9-d11bbef416f3_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Move up to what?</h2><p>So change the question. Don't ask whether new work appears. Ask whether it can catch a mass of people &#8212; and whether it <em>pays a living.</em> Every rung that mattered in the last two centuries was both. The factory floor and the office floor weren't just <em>jobs</em>; they were <em>wide, well-paid jobs</em> that an ordinary person could reach by learning. That combination is the thing that made the ladder a ladder.</p><p>So name it. When people say the displaced will "move up," I keep asking the same thing: up to <em>what?</em> The answers are fewer than they sound. Two of them are jobs &#8212; and I'll take both seriously in a moment. A third isn't a job at all: <em>stop being labor and become an owner.</em> That one's the deepest answer of the three, which is exactly why it gets its own essay and not a line here &#8212; it's a wall you buy your way past, not a rung you climb, and we'll come back to it. For now, the two that are jobs, because they're the ones the optimist reaches for first.</p><p>The first: <em>"They'll orchestrate the AI &#8212; supervise the agents, manage the fleet."</em> Fine. Let's literalize it. A former analyst now directs a swarm of agents and does what a ten-person team used to do. But notice what that is: it is not a new floor with new seats. It's the <em>same</em> work, with nine fewer people. The entire value of the role is <em>more output, fewer humans.</em> So "moving up a rung" here means, precisely: nine people leave, and one stays with a better title. The promotion <em>is</em> the layoff. You cannot rehouse a displaced multitude on a rung whose whole purpose is to need fewer of them.</p><p>The second: <em>"New categories we can't imagine yet &#8212; like podcaster, like app developer."</em> Real, and I believe it. But ask the only question that matters for a ladder: <em>how wide, and how well-paid?</em> And hold that question, because to answer it honestly we have to talk about where a wage actually comes from. That turns out to be the crux of the whole thing &#8212; and it's the part almost nobody slows down to explain.</p><h2>How a wage is actually made</h2><p>Here's a question that sounds simple and isn't: why does anyone get paid what they get paid?</p><p>The tempting answer is "because the work is valuable." But that can't be the whole story, and there's a 250-year-old puzzle that proves it. Adam Smith named it in 1776 and couldn't crack it. It's called the <strong>diamond-water paradox</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>. Water keeps you alive; without it you die in days. A diamond does nothing &#8212; it's a pretty rock. And yet water is nearly free and a diamond costs a fortune. If price followed <em>value</em>, water would be priceless and diamonds worthless. It's the other way around. Why?</p><p>The answer took economists another hundred years to work out, and it is this: price isn't set by how useful something is <em>in total.</em> It's set <em>at the margin</em> &#8212; by how much the <em>next one</em> is worth, and that depends on how <em>scarce</em> it is. Water is abundant, so the next glass is worth almost nothing, even though all the water in the world is worth everything. Diamonds are rare, so the next one is dear. Scarcity, not usefulness, sets the price.</p><p>Wages work the same way. Your pay tracks two things multiplied together: the <strong>value</strong> of what you produce, <em>and</em> the <strong>scarcity</strong> of your ability to produce it. You need both. A skill nobody wants pays nothing, however rare. And &#8212; here's the diamond-water half people miss &#8212; <em>a skill everyone has pays nothing either, no matter how essential it is.</em> A wage was never paid for value. It was paid for <strong>scarce value.</strong></p><p>Now you can see why knowledge work paid so well for fifty years. A trained mind was <em>valuable</em> &#8212; it produced things people wanted &#8212; and it was <em>scarce</em>, because learning hard things is slow and not everyone does it. Productive and rare. A diamond. That's the rung, explained: it paid a broad, good wage because trained cognition was a diamond that ordinary people could, with effort, go and acquire.</p><h2>What AI actually does to the wage</h2><p>AI does not destroy the <em>value</em> of human cognition. Your judgment still produces plenty &#8212; arguably more than ever, amplified by the machine.</p><p>What AI does is change the <em>scarcity.</em> It takes the trained cognition that used to be rare and makes it abundant. <strong>It turns the diamond into water.</strong> Still essential &#8212; more essential, even &#8212; and now everywhere, on tap, for the price of a subscription.</p><p>And we already know what the market pays for something essential and abundant: it pays what it pays for water. Almost nothing. Not because the thing stopped having value, but because the <em>next unit</em> of it is no longer scarce.</p><p>A wage, as we just saw, is the price of scarcity &#8212; not of usefulness. So the wage falls even while the usefulness holds. You can keep the job and lose the living.</p><h2>The scarcity that was always there</h2><p>But the wage doesn't fall to zero. Strip away the cognition that just turned to water, and something is left standing underneath &#8212; something that was rare long before the machine arrived.</p><p>Call it craft &#8212; and I mean the word precisely: the thing you <em>earn,</em> through years of effort and rigor. The discipline to know which problem is worth solving, and when the machine has quietly gone wrong. The standard that won't ship what's almost right. The willingness to put your name on the outcome. None of it is new, and none of it just <em>became</em> scarce. It was always the rare thing. A great engineer and a mediocre one were never separated by the routine work the machine now does for both of them &#8212; they were separated by knowing <em>which way is wrong,</em> the call only one of them could make.</p><p>So the wage doesn't disappear. It splits. It collapses for the many, whose faculty just turned to water &#8212; and it holds, maybe climbs, for the few who developed their craft. The ladder doesn't <em>narrow</em> to a spike. The wide floor falls away, and what's left standing is <strong>the spike that was always there:</strong> a tall, thin place where a handful of masters command a scarce thing.</p><p>Years ago, on my first day after taking over as head of engineering for a team of about a hundred people, a new release was crashing in front of an important customer, and the team had a hotfix ready to ship. Everyone wanted it out the door &#8212; the customer was furious, the pressure enormous. I asked them to wait. Not because I'd read the code; I hadn't. Because I had made this exact mistake myself &#8212; fifteen years earlier, at another company, I was the dev manager, and I shipped the bad patch in the same kind of panic, certain it was fine, and watched it break. That's how I knew the shape of it: a patch written in a panic carries more bugs. So we kept testing. Two days later, there it was &#8212; a second bug, hiding in the fix. They wrote the tests that caught it, shipped a version that held, and we kept the customer.</p><p>A hundred capable people, and the moment that mattered turned on a call that took experience the rest of them simply hadn't had the years to earn yet. Their hands were abundant; the craft was scarce.</p><p>For the past year I've worked with Claude every day, and the division of labor is identical: left to itself, the machine will burn an afternoon sprinting confidently down the wrong road; left to myself, I can't move at anything near the pace we move together. It does the work of 30 engineers. I do the one thing it can't &#8212; I know when it's wrong. The model is abundant; the craft is scarce; the value lives in the craft.</p><p>And don't mistake that for good news about me. One of me is worth thirty only because the work stopped needing the other twenty-nine &#8212; my leverage and their absence are the same fact, seen from two seats. The machine complements the one <em>by</em> replacing the many.</p><p>Let me subject you to a thought experiment. Say there are thirty million software professionals in the world today; developers, QA engineers, support staff, and the rest of the trade<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>. I suspect that in ten years there will be a tiny fraction of that &#8212; call it thirty thousand &#8212; and each one a master, each one supplying the scarce residue to a fleet of tireless machines. I might be wildly off on just how few are left. But the <em>shape</em> is the argument: a broad profession, paid a broad wage for a once-scarce skill, collapsing into a narrow spike of the few who hold what stays rare. The wide rung, falling away as you watch.</p><p>A master isn't born; he's made &#8212; on the lower rungs, doing the junior work clumsily until he's done it enough times to do it well. I learned to spot a doomed hotfix by shipping one myself, years before, back when I was the one in the panic. The spike is built entirely out of the floor beneath it.</p><p>But the floor is the first thing the machine takes. The entry-level work &#8212; the cheap, the routine, the <em>learnable</em> &#8212; is exactly what AI does first and best. The youngest workers in the most exposed fields are already quietly disappearing: not fired so much as never hired, their rung simply not refilled<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>. So the masters left standing aren't the first of a new elite. They're the last of an old one &#8212; a generation with no apprentices behind it, because the apprenticeship is gone. Who trains the thirty-thousand-and-first?</p><h2>"But everything will be so cheap"</h2><p>Here's the objection I take most seriously, because it's the one I can't fully put down. Maybe the wage doesn't matter. If AI drives the cost of nearly everything toward zero &#8212; the doctor, the tutor, the lawyer, the ride, the kilowatt &#8212; then a small paycheck in 2040 might buy a fuller life than a six-figure salary buys today. Abundance, the optimists say, makes the wage beside the point. And they're partly right. I don't think this story ends in breadlines; nobody starves in a world this productive.</p><p>But a wage was never only what it buys. It was a <em>claim</em> &#8212; your share of the wealth you helped make &#8212; and a <em>place</em>, a reason the economy had to deal you in. And the things that actually make a life are not the things abundance makes cheap. The house in the district with the good school doesn't get cheaper when software does &#8212; there's only one of it, and now everyone's bidding. And what you own holds its value; what you rent does not. The cheap things get cheaper, and the scarce things &#8212; land, position, a claim on the future &#8212; get dearer, because that's where all the displaced money goes hunting for a home. You can hand every person a miracle and still leave them with no claim on it. That isn't a rebuttal to abundance. It's the question abundance can't answer &#8212; and I'll come back to it.</p><h2>The rung and the wall</h2><p>So that is the walk we just took. For fifty years an ordinary person could trade a trained mind for a claim on the nation's wealth, and climb. The machine has come for that trade the way it came for the farm and the factory before it &#8212; but this time no wider floor is opening beneath, because the thing that always opened it is gone. The value still flows. It simply no longer flows through us. What's left of the old rung is a spike: a narrow perch for the few who spent a lifetime earning a craft, and a long drop for everyone else. And the comfort we all reach for &#8212; <em>it'll be so cheap</em> &#8212; is the smallest comfort there is, because cheap is not the same as yours. A wage was a claim and a place &#8212; the reason the economy had to deal you in at all. Lose the wage, and it doesn't have to anymore.</p><p>Which leaves the question I promised to come back to &#8212; and it's the one everything else hangs from. If a trained mind is no longer a claim on the wealth, then what is? The only rung still standing above labor isn't a kind of work at all. It's ownership &#8212; not a rung you climb, but a wall you buy your way past. The windfall still falls from the tree. But the orchard has a fence now, and a name on the deed.</p><p>So the question was never really <em>what will people do.</em> It's what waits on the rung above &#8212; the one made of capital, not work, and the only place the value still pools. I've sat across the table from the investors already buying in. I've read what today's CEOs write to each other &#8212; addressed to companies, never to you. That's the rung we climb to next.</p><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>About 41% of the U.S. labor force worked in agriculture in 1900 (USDA / U.S. Census historical series); manufacturing employment peaked at almost exactly three in ten &#8212; roughly 30% of all jobs &#8212; around 1950 (U.S. Bureau of Labor Statistics).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>A long-standing term in economics for the assumption that there is a fixed quantity of work to go around. Some economists warn it gets invoked too readily to wave away genuine displacement &#8212; which is exactly why the argument here grants it in full and rests its weight elsewhere.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p><span>The gap between what a buyer would have been willing to pay and what they actually pay. The idea traces to the French engineer Jules Dupuit (1844); Alfred Marshall named and popularized it in his </span><em>Principles of Economics</em><span> (1890).</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p><span>Named for William Stanley Jevons, who observed in </span><em>The Coal Question</em><span> (1865) that as steam engines grew more efficient and burned less coal per task, total coal consumption </span><em>rose</em><span> &#8212; cheaper steam power simply found far more uses. Economists now call it the rebound effect.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p><span>Smith laid out the puzzle in </span><em>The Wealth of Nations</em><span> (1776) as the gap between "value in use" and "value in exchange," but his labor theory of value couldn't close it. The resolution &#8212; marginal utility &#8212; arrived a century later, in the 1870s, discovered independently by William Stanley Jevons, Carl Menger, and L&#233;on Walras: the Marginalist Revolution.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>Conservative on purpose: estimates put software developers alone near 47 million worldwide, and the broader IT trade above 80 million. The argument only sharpens with a bigger base.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>Stanford&#8217;s Digital Economy Lab (Brynjolfsson, Chandar, and Chen, 2025) found early-career workers &#8212; ages 22 to 25 &#8212; in the most AI-exposed occupations already down about 13% relative to their peers, even after controlling for firm-level shocks. The mechanism wasn&#8217;t layoffs; it was a quiet hiring freeze &#8212; firms simply stop backfilling entry-level roles as they empty.</p></div></div>]]></content:encoded></item><item><title><![CDATA[The Asymmetry]]></title><description><![CDATA[The gap between what AI can do and what most organizations are doing with it is growing every week. Why the asymmetry matters and what to do about it.]]></description><link>https://essays.xcud.com/p/the-asymmetry</link><guid isPermaLink="false">https://essays.xcud.com/p/the-asymmetry</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Wed, 13 May 2026 14:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JC5q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>The Asymmetry</h1><p>Two emails arrived on the same morning last week.</p><p>The first was four paragraphs long. It opened with background I already knew, wandered through a description of several problems without distinguishing which ones mattered, and ended mid-thought &#8212; no clear ask, no proposed next step. I read it twice and still wasn't sure what I was supposed to do. The sender had emptied their head into my inbox, and the work of organizing their thoughts was now mine.</p><p>The second was three sentences. It stated the problem, proposed a solution, and asked for my sign-off by Thursday. I read it once, replied in thirty seconds, and moved on.</p><p>Same medium &#8212; text in a rectangle on my screen. Completely different experience. One was like drawing in a breath and finding oxygen. The other was like drawing in a breath and finding nothing.</p><h2>The Factoring Problem</h2><p>There's an operation in mathematics called prime factorization. Given a large number, find the primes that multiply together to produce it. It's famously hard &#8212; so hard that modern cryptography depends on it. But here's the thing: <em>verifying</em> the factors is trivial. If I tell you that 7 &#215; 13 = 91, you can confirm it in your head. Finding those factors in the first place is where the work lives.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JC5q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JC5q!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!JC5q!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!JC5q!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!JC5q!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JC5q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The factoring asymmetry &#8212; finding the prime factors of 91 is hard; verifying that 7 &#215; 13 = 91 is trivial&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The factoring asymmetry &#8212; finding the prime factors of 91 is hard; verifying that 7 &#215; 13 = 91 is trivial" title="The factoring asymmetry &#8212; finding the prime factors of 91 is hard; verifying that 7 &#215; 13 = 91 is trivial" srcset="https://substackcdn.com/image/fetch/$s_!JC5q!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!JC5q!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!JC5q!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!JC5q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6253344d-bf7b-4694-b07a-1db7ddc7114f_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Good cognitive work has this same asymmetry. It's expensive to produce and cheap to consume. A well-written article, a well-run meeting, a clear email &#8212; these are things where someone did the hard work of analysis, synthesis, and organization upstream so that the receiver doesn't have to. The reader, the attendee, the recipient gets the factors. They can verify, engage, build on top. The complexity was absorbed before it reached them.</p><p>This is what it feels like to read something that makes you smarter: the author did the factoring, and you get to do the multiplication. You leave the interaction with more than you came in with. It feels like oxygen because it <em>is</em> &#8212; metabolizable thought, ready to use.</p><h2>The Inversion</h2><p>Slop &#8212; in every medium &#8212; inverts this asymmetry.</p><p>The producer does the cheap part: they generate volume. Words in a Slack message. Slides in a deck. Attendees in a meeting. An hour of someone's voice in a recording. The expensive part &#8212; figuring out what's actually useful, what the point is, what action should follow &#8212; gets transferred to the receiver. The sender saved themselves five minutes of thought and cost you thirty minutes of interpretation.</p><p>A meeting with no agenda is unfactored work. The organizer didn't do the hard thinking about what needs to be decided, so everyone in the room has to do it together, in real time, expensively. A Slack message that ends with "..." is an unfactored request &#8212; the sender hasn't finished their own thought but has already conscripted you into completing it. A rambling email is someone offloading the cost of organizing their ideas onto whoever opens it.</p><p>And it's not just informal communication. A report that buries its conclusion on page twelve. A LinkedIn post that takes six paragraphs to say nothing. A status update that recites activity without revealing progress. These all have the same shape: the author did the easy part and left you with the hard part.</p><p>Here's the same project, the same week, written two ways:</p><blockquote><p><em>The Q3 marketing campaign is progressing. We've had several meetings with the agency to discuss creative direction. The team is reviewing media buy options. Some concerns came up about the budget that we're looking into. We're also coordinating with legal on the new disclaimer language. Overall things are on track and we'll have more updates next week.</em></p></blockquote><p>You read that. What do you know now? What decision can you make? Nothing. You've consumed someone's time without receiving their thinking. Now read this:</p><blockquote><p><em>Q3 campaign is a week behind schedule. The agency delivered three creative concepts &#8212; we've approved one and sent revision notes on the hero video, due back Wednesday. Media buy is $15K over budget because cable rates increased. Legal approved the disclaimer language.</em> <em><strong>Decision needed:</strong></em> <em>do we cut the print placement to stay on budget, or request the additional $15K?</em></p></blockquote><p>Same project. Same week. The first version transfers all the cognitive work to you. The second arrives with the factors already extracted &#8212; what's late, what's done, and what you need to decide. One is a deposit. The other is a withdrawal disguised as a deposit.</p><p>This isn't about doing all the thinking alone. A leader who brings a clearly framed problem to a working group &#8212; <em>here's what's broken, here's why it matters, here's what we need to figure out</em> &#8212; has done their factoring. The problem is defined, the stakes are clear, and the team's role is to factor the solution. That's not offloading. That's delegation. The inversion happens when the problem itself hasn't been factored &#8212; when the audience has to figure out not just the answer but the question.</p><h2>The Director Test</h2><p>I've spent years advising people who want to advance in large organizations, and there's a pattern I keep coming back to.</p><p>Non-directors bring you a problem. They walk into your office, describe what's broken, and look at you. The implicit request is: <em>think about this for me.</em> They've done the multiplication &#8212; they've identified that 91 is a number that matters &#8212; but they haven't factored it. Now it's your problem.</p><p>Directors bring you the problem, a proposed solution, the thinking that led to that solution, and a simple request for borrowed authority to implement it. They hand you 7 &#215; 13 and say: "I need your sign-off to proceed." Same information. Radically different cognitive load. The director absorbed the expensive part before they walked through the door.</p><p>This isn't about seniority. I've seen junior engineers who factor beautifully and executives who can't organize a thought before they broadcast it. The asymmetry is a <em>skill</em>, not a rank. And like most skills, the first step is noticing it exists.</p><h2>The Guitar Amp</h2><p>This is where AI enters the picture &#8212; and where the asymmetry gets sharper, not duller.</p><p>AI is a guitar amp. It makes you louder, not better. If you can play, more people hear something worth hearing. If you can't, more people hear that you can't. An expert who uses AI to research, draft, analyze, and refine is doing <em>more</em> factoring, not less. They're using the tool to absorb even more complexity before it reaches the audience. The output is denser, clearer, more thoroughly considered &#8212; because the human brought expertise and the machine brought scale.</p><p>But someone without the discipline to do the thinking produces slop at scale. They generate volume faster, more confidently, and with better formatting. The output <em>looks</em> factored. It has the shape of organized thought: clear paragraphs, bullet points, professional tone. But the expensive work was never done. Nobody decided what actually matters. Nobody filtered, prioritized, or applied judgment. The receiver opens it, reads it, and slowly realizes they're breathing vacuum.</p><p>The wave of "AI slop" complaints is real, and some of it is earned. People are drowning in AI-generated content that looks polished but says nothing &#8212; emails, reports, LinkedIn posts that have the shape of organized thought without the substance. If you're tired of breathing vacuum, you should be. That's not a complaint about AI. It's a complaint about unfactored work produced at new scale.</p><p>But there are two complaints hiding in "AI slop." One is about quality &#8212; <em>this work is bad</em> &#8212; and that's fair. It was bad before AI made it cheaper to produce. The other is about provenance: <em>this was made with AI.</em> That's the tell. Scanning for emdashes, flagging "certainly" and "moreover," treating typographic habits as a forensic test for whether a machine was involved &#8212; all to avoid the harder question of whether the work is actually good. When the objection isn't that the work is bad but that AI touched it, the person is admitting they can't evaluate the output on its merits &#8212; they need to know the process to judge the product. An expert reads AI-assisted work and knows immediately whether the thinking was real. The provenance is irrelevant.</p><p>The tool doesn't change the asymmetry. It reveals who was already doing the factoring and who was getting away without it.</p><h2>The Obligation</h2><p>You have a factoring machine. Ship factored work.</p><p>Every message, meeting, article, and interaction is a claim on someone else's attention. That claim carries an obligation: absorb the complexity before you transfer it. Do the factoring. Show up with the primes, not the product.</p><p>The tools for doing this have never been better. You can draft, research, analyze, reorganize, pressure-test, and refine your thinking at a speed that was unimaginable five years ago. The cost of producing well-factored work has collapsed. Which means the bar for what's acceptable has risen &#8212; or should have.</p><p>If you send a meeting invite, include the decision that needs to be made. If you send a message, finish your thought before you hit send. If you write a report, put the conclusion first. If you write an email, know what you're asking before you start typing.</p>]]></content:encoded></item><item><title><![CDATA[You Can Still Understand the Machine]]></title><description><![CDATA[You could once trace a keystroke to a pixel through a Commodore 64. Modern AI is more understandable than it looks if you peel it one layer at a time, from the system down to the neuron.]]></description><link>https://essays.xcud.com/p/you-can-still-understand-the-machine</link><guid isPermaLink="false">https://essays.xcud.com/p/you-can-still-understand-the-machine</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Fri, 17 Apr 2026 14:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8s0S!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>You Can Still Understand the Machine</h1><p>There's a nostalgia among people who grew up with early personal computers &#8212; the Commodore 64, the Apple II, the TRS-80 &#8212; for the time when you could understand <em>everything</em> about your machine. The CPU had a few thousand transistors. The memory map fit on a single page. You could trace the flow of electricity from keystroke to screen pixel and predict exactly what would happen. You owned the whole thing, top to bottom.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8s0S!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8s0S!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png 424w, https://substackcdn.com/image/fetch/$s_!8s0S!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png 848w, https://substackcdn.com/image/fetch/$s_!8s0S!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png 1272w, https://substackcdn.com/image/fetch/$s_!8s0S!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8s0S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png" width="1456" height="823" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:823,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Commodore 64&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Commodore 64" title="Commodore 64" srcset="https://substackcdn.com/image/fetch/$s_!8s0S!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png 424w, https://substackcdn.com/image/fetch/$s_!8s0S!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png 848w, https://substackcdn.com/image/fetch/$s_!8s0S!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png 1272w, https://substackcdn.com/image/fetch/$s_!8s0S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe9c1089-43fa-40fe-b71c-c3c3e5658838_1920x1085.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Evan-Amos, CC BY-SA 4.0, via Wikimedia Commons</em></p><p>Modern AI systems don't offer that same feeling of total mastery. But they're more understandable than most people assume &#8212; if you stop trying to grasp the whole thing at once.</p><p>The trick is to peel it apart, one layer at a time. Start at the top &#8212; the software system you interact with &#8212; and work your way down through the reasoning strategy, the language model, the network architecture, and finally the individual neuron. At each layer, the math is straightforward and the ideas are concrete. And somewhere on the way down, the thing that felt like digital alchemy starts to look like what it actually is: simple mathematical operations, repeated at extraordinary scale.</p><h2>Layer 1: The System</h2><p>The first and most common misconception about artificial intelligence is the conflation of the "model" with the "system." A large language model, in isolation, is a static file &#8212; billions of numerical weights representing statistical probabilities. It possesses no agency, no continuous memory, and no inherent ability to interact with the external world. The capabilities people attribute to AI &#8212; the memory, the web searches, the personality, the safety &#8212; emerge from a complex orchestration layer that <em>surrounds</em> the model.</p><p>When you type a message into ChatGPT, Claude, or Gemini, you're not talking to a model. You're talking to this system.</p><p>There's an irony here worth noting. For years, serious practitioners bristled when the public called machine learning "AI" &#8212; the term was too generous, too anthropomorphic for what were really just statistical models. We demanded specificity: <em>machine learning, deep learning, NLP, LLM</em>. But now that engineering has wrapped these models in memory, tools, safety layers, and persistent agency, we find ourselves telling people to stop calling the whole thing an "LLM." The term that was once too narrow is now reductive. Now we're the ones demanding the public call it "AI."</p><h3>The Serving Layer</h3><p>When your message reaches the provider's servers, it enters a serving infrastructure that's invisible to you but directly affects what you get back. Your request gets queued, batched with other users' requests for efficient GPU utilization, and routed to available hardware. Common responses may be served from cache rather than computed fresh. Under heavy load &#8212; the kind you see reflected on <a href="https://status.anthropic.com/">status pages</a> &#8212; service can degrade: slower responses, shorter outputs, or temporary unavailability.</p><p>Behind the scenes, techniques like <em>speculative decoding</em> (where a fast, small model drafts candidate tokens and a larger model verifies them) and <em>best-of-N sampling</em> (generating multiple completions and selecting the best) can improve speed or quality without the user ever knowing. The details vary by provider and are rarely disclosed, but the principle is worth understanding: what users experience as "model quality" is not solely a property of the model's weights. It's also a function of how much compute the system spends on your specific query, how loaded the infrastructure is at that moment, and engineering decisions the provider has made about cost, speed, and quality tradeoffs.</p><h3>Safety and Guardrails</h3><p>Safety and moderation layers screen both your input and the model's output. These aren't the language model itself &#8212; they're often separate, smaller classifiers running alongside it, purpose-built to detect specific categories of risk.</p><p>On the input side, classifiers scan your prompt for requests involving harmful content: instructions for weapons or dangerous substances, generation of child sexual abuse material, attempts to extract the model's system prompt or manipulate its behavior through <a href="https://arxiv.org/abs/2302.12173">prompt injection</a>. On the output side, similar classifiers screen the model's response before it reaches you, catching cases where the model might have generated something that slipped past its own training-level safeguards.</p><h3>The Illusion of Memory</h3><p>An LLM doesn't remember previous conversations natively &#8212; it processes a fixed window (the <em>context window</em>) each time. So the system assembles everything the model needs to see: your current message, relevant conversation history, retrieved documents or memories from previous sessions, images or other media converted into numerical representations, and a <em>system prompt</em> &#8212; the hidden instructions that define who the model is, how it should behave, what it should refuse, and what tone it should take. This assembled context is the model's entire reality. It evaluates it from scratch every time &#8212; no persistent state, no running thread of consciousness between calls. Products that feel like they "remember" you are doing retrieval and injection behind the scenes.</p><p>This is what the industry briefly called "<a href="https://en.wikipedia.org/wiki/Prompt_engineering">prompt engineering</a>" and later "<a href="https://simonwillison.net/2025/Jun/27/context-engineering/">context engineering</a>" &#8212; the art of controlling what goes into the window. The terms faded not because the problem was solved, but because the systems got better at assembling context automatically. The principle remains: <em>the model can only work with what's in front of it</em>. A well-assembled context produces a brilliant response. A sloppy one produces hallucinations, missed instructions, or generic boilerplate. When a long conversation seems to "forget" your earlier instructions, it's because the system ran out of room and physically truncated them &#8212; to the model, they never existed.</p><h3>Tools and Agency</h3><p>Without tools, a language model is a brain in a jar &#8212; impressive, but it can't see, can't act, can't access anything beyond what's already in its context window. It can talk <em>about</em> your calendar, but it can't check it. It can describe how to query a database, but it can't run the query. Everything that makes a modern AI assistant feel like a <em>collaborator</em> rather than a conversationalist comes from tool use.</p><p>The mechanism is almost disappointingly simple:</p><pre><code>while tool_calls &lt; max_iterations:
    response = stream(model, context)

    if response.contains_tool_call():
        tool, args = parse_json(response)
        result = execute(tool, args)
        context.append(result)
        tool_calls += 1
    else:
        return response  # model didn't ask for a tool &#8212; it's done
</code></pre><p>That's the entire agentic loop. The orchestration layer streams tokens from the model and watches for structured output &#8212; typically JSON &#8212; indicating a tool request. When it detects one (often by something as basic as counting braces to find balanced JSON), it pauses the stream, executes the requested tool, injects the result back into the context, and lets the model continue. A few hundred lines of Python. The model reads your email, decides it needs to check your calendar, finds a conflict, drafts a reply, and asks for your approval &#8212; each step a separate tool call, each result fed back in for the next decision. It feels like agency. Under the hood, it's a state machine with two states.</p><p><a href="https://modelcontextprotocol.io/">Model Context Protocol</a> (MCP), introduced by Anthropic in late 2024, accelerated this by giving tool providers a standardized interface. MCP may prove to be transitional &#8212; training wheels for a more natural capability. Language models have internalized vast amounts of documentation about command-line tools, APIs, and shell interfaces from their training data. They already <em>know</em> how to use <code>grep</code>, <code>curl</code>, <code>git</code>, and <code>psql</code> without needing a schema to tell them &#8212; and any CLI tool they haven't seen before, they can discover with <code>--help</code>. More importantly, the tools themselves are converging toward the patterns models already understand. Long-running commands are being rewritten to separate <code>start</code> and <code>status</code> subcommands so AI can manage them asynchronously. Outputs are becoming more structured. The tools are adapting to the model as much as the model adapted to the tools. The emerging pattern is less protocol, more fluency &#8212; models that reason about what they need, compose the right commands on their own, and execute them directly. The result either way: AI that can read your inbox, pull a report, or commit code. Not answer questions <em>about</em> those things &#8212; actually go do them.</p><p>Everything we've described so far &#8212; the serving layer, the safety classifiers, the context assembly, the tool orchestration &#8212; is scaffolding. The engine at the center of it, the thing doing the actual <em>thinking</em>, is a large language model. Everything else exists to serve it, protect it, or extend it.</p><h2>Layer 2: How It Thinks</h2><p>Before we look inside the language model itself, it's worth understanding a distinction that changed the entire field: the difference between <em>completing</em> and <em>reasoning</em>.</p><h3>Completion: Fast and Intuitive</h3><p>Early LLMs were pure completion engines &#8212; given a prompt, they'd calculate the probability of every possible next word and select one. Because they select from a probability distribution rather than following deterministic rules, the output is fundamentally <em>stochastic</em>: ask the same question twice and you'll get different answers. This isn't a bug &#8212; it's what makes the outputs feel creative and varied rather than robotic. But it also means the model is never "looking up" an answer. It's generating one, fresh, every time.</p><p>The flaw is that completion models are forced to answer immediately. If a problem requires ten steps of logic, the model has to output the first step without having calculated the tenth. It's what cognitive scientists would call <a href="https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow">System 1 thinking</a>: fast, reflexive, intuitive &#8212; and prone to confident-sounding errors on anything requiring deliberation. You could ask an early LLM to write code and it would <em>look</em> right &#8212; proper syntax, reasonable variable names, plausible structure. But if you actually read it, you'd see it wasn't code. It was text <em>in the style of</em> code &#8212; it mimicked the surface patterns without understanding the logic. Plausible at a glance, wrong under scrutiny. Nobody confused it with real understanding. Practitioners found they could improve the outputs by providing examples in the prompt &#8212; show the model what a good answer looks like, and the <a href="https://arxiv.org/abs/2005.14165">completion gets better</a>. But better completion is still completion.</p><h3>Reasoning: Slow and Deliberate</h3><p>Then someone asked the model to show its work.</p><p>The shift toward reasoning came from a disarmingly simple observation: models perform dramatically better when you just ask them to break problems into steps. <a href="https://arxiv.org/abs/2201.11903">Wei et al.'s 2022 "chain-of-thought prompting" paper</a> showed that adding instructions like these to a system prompt improved math and logic performance by huge margins:</p><pre><code>Before answering, break the problem into steps.
Work through each step carefully.
If you notice an error in your reasoning, backtrack and correct it.
Only provide your final answer after verifying your logic.
</code></pre><p>That's it. The same completion engine, given permission to <em>think out loud</em>, becomes a dramatically better reasoner. The observation &#8212; that a simple prompt change could unlock step-by-step problem solving &#8212; was the spark. Researchers built on it with tree-of-thought reasoning and self-consistency sampling. Then providers like OpenAI took the next logical step: if prompting for chain-of-thought works this well, why not use <a href="https://openai.com/index/learning-to-reason-with-llms/">reinforcement learning</a> to train models to do it on their own? The result was a new generation of reasoning models that produce extended internal chains of thought automatically, without being asked.</p><p>This is the transition point &#8212; the moment AI stopped feeling like a fancy autocomplete and started feeling like something you could <em>talk to</em>. A completion model gives you the most probable next word. A reasoning model gives you a considered answer. It plans. It second-guesses itself. It catches its own mistakes. The mechanism underneath hasn't changed &#8212; it's still next-token prediction &#8212; but the behavior is so qualitatively different that it crosses a threshold in how humans perceive it.</p><h2>Layer 3: The Language Model</h2><p>At the core of every AI assistant is a language model, and its fundamental operation is almost absurdly simple: <strong>it predicts the next word.</strong></p><h3>Next-Word Prediction</h3><p>Given a sequence of words, the model outputs a probability for every word in its vocabulary &#8212; typically 30,000 to 100,000 tokens &#8212; that could come next. It's as if the model is looking at an enormous grid &#8212; every word it knows along one axis &#8212; and lighting up the ones most likely to follow what's been said so far.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0iRu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F213485ee-5515-42a7-a012-0126c163c277_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0iRu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F213485ee-5515-42a7-a012-0126c163c277_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!0iRu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F213485ee-5515-42a7-a012-0126c163c277_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!0iRu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F213485ee-5515-42a7-a012-0126c163c277_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!0iRu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F213485ee-5515-42a7-a012-0126c163c277_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0iRu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F213485ee-5515-42a7-a012-0126c163c277_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/213485ee-5515-42a7-a012-0126c163c277_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Next-word prediction heatmap &#8212; a grid of vocabulary words with high-probability candidates like \&quot;mat,\&quot; \&quot;floor,\&quot; and \&quot;couch\&quot; glowing bright while thousands of others stay dark&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Next-word prediction heatmap &#8212; a grid of vocabulary words with high-probability candidates like &quot;mat,&quot; &quot;floor,&quot; and &quot;couch&quot; glowing bright while thousands of others stay dark" title="Next-word prediction heatmap &#8212; a grid of vocabulary words with high-probability candidates like &quot;mat,&quot; &quot;floor,&quot; and &quot;couch&quot; glowing bright while thousands of others stay dark" srcset="https://substackcdn.com/image/fetch/$s_!0iRu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F213485ee-5515-42a7-a012-0126c163c277_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!0iRu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F213485ee-5515-42a7-a012-0126c163c277_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!0iRu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F213485ee-5515-42a7-a012-0126c163c277_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!0iRu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F213485ee-5515-42a7-a012-0126c163c277_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Given "The cat sat on the ___", the model assigns high probability to words like "mat," "floor," "couch" and very low probability to words like "algorithm" or "bankruptcy." It picks one &#8212; stochastically, with controlled randomness &#8212; appends it, and repeats. Word by word, a response emerges.</p><p>How does it know which words are likely? It learned by reading &#8212; trillions of words of human-generated text. Books, articles, scientific papers, code repositories, legal filings, conversations. During training, the model was shown sequence after sequence and asked to <a href="https://www.youtube.com/watch?v=kCc8FmEb1nY">predict the next word</a>. Every time it guessed wrong, its weights were adjusted slightly (we'll see exactly how in <a href="#how-it-learns">How It Learns</a>). After trillions of these corrections, the probabilities it assigns are a distillation of patterns absorbed from the largest corpus of human writing ever assembled.</p><p>There's no separate module for reasoning, no knowledge database it looks things up in, no explicit rules about grammar or logic. Everything &#8212; the coherent paragraphs, the factual knowledge, the ability to write code or summarize a legal document &#8212; <a href="https://www.jasonwei.net/blog/some-intuitions-about-large-language-models">emerges from next-word prediction</a> at this scale.</p><h3>Prediction Requires Understanding</h3><p>The intuition that sophisticated intelligence can arise from a mechanism that is essentially advanced autocomplete is deeply counterintuitive. But accurate prediction <em>requires</em> understanding. To consistently predict the next word in "The heavy steel ball was dropped from the roof onto the fragile glass table, and the table ___," the model must have embedded a functional understanding of gravity, material strength, physical collision, and cause-and-effect into its weights. To predict the next token in a Python function, it must understand programming logic. To complete a legal argument, it must grasp the structure of legal reasoning.</p><p>What emerges from this pressure is something researchers increasingly call a <strong><a href="https://arxiv.org/abs/2403.00833">world model</a></strong> &#8212; an internal representation of how reality works, learned entirely from the statistical structure of text. Grammar, factual recall, spatial reasoning, basic deductive logic &#8212; none of it is programmed. All of it is a byproduct of the relentless optimization for next-word prediction across trillions of examples. The model doesn't know it has a world model. It was never told to build one. But the only way to predict language this well is to understand the world that language describes.</p><h3>From Words to Numbers</h3><p>This world model is encoded entirely in numbers. So the first question is practical: how does text become math?</p><p><a href="https://nebius.com/blog/posts/how-tokenizers-work-in-ai-models">Tokenization</a> is the front door: text gets broken into discrete chunks called <em>tokens</em>. Tokens aren't always whole words &#8212; they're frequently sub-words, prefixes, or suffixes. "Understanding" might become "under" + "stand" + "ing." Each unique token in the model's vocabulary gets a numerical ID.</p><p>This seemingly mundane preprocessing step has real consequences. Many of the persistent quirks of LLMs &#8212; <a href="https://karpathy.ai/zero-to-hero.html">spelling errors, counting mistakes, struggles with certain languages</a> &#8212; trace directly back to how words were carved into tokens. The model only sees the numerical IDs; it's blind to the actual letters inside a token unless it's learned to reconstruct them.</p><h3>The Geometry of Meaning</h3><p>Each token ID gets mapped to a dense array of numbers &#8212; a point in a <a href="https://www.datacamp.com/blog/vector-embedding">high-dimensional mathematical space</a>. This space might have thousands of dimensions, each representing some abstract feature that the model learned during training.</p><p>The key insight is that <em>meaning becomes geometry</em>. Words with similar meanings cluster near each other in this space. "Happy" and "joyful" are close together; "happy" and "tractor" are far apart. More remarkably, the <em>directions</em> in this space encode relationships. The classic demonstration: take the vector for "King," subtract "Man," add "Woman," and you land near "Queen." The model has learned that gender is a direction in its semantic space, and royalty is another, and these directions compose.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NE9M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a78088-ad09-496a-b856-286b5e3e52a4_1252x1303.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NE9M!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a78088-ad09-496a-b856-286b5e3e52a4_1252x1303.png 424w, https://substackcdn.com/image/fetch/$s_!NE9M!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a78088-ad09-496a-b856-286b5e3e52a4_1252x1303.png 848w, https://substackcdn.com/image/fetch/$s_!NE9M!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a78088-ad09-496a-b856-286b5e3e52a4_1252x1303.png 1272w, https://substackcdn.com/image/fetch/$s_!NE9M!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a78088-ad09-496a-b856-286b5e3e52a4_1252x1303.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NE9M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a78088-ad09-496a-b856-286b5e3e52a4_1252x1303.png" width="1252" height="1303" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/57a78088-ad09-496a-b856-286b5e3e52a4_1252x1303.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1303,&quot;width&quot;:1252,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Word vectors plotted in 2D space &#8212; \&quot;king\&quot; and \&quot;man\&quot; share a direction, \&quot;woman\&quot; extends along a different axis, and \&quot;queen\&quot; lands exactly where vector arithmetic predicts&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Word vectors plotted in 2D space &#8212; &quot;king&quot; and &quot;man&quot; share a direction, &quot;woman&quot; extends along a different axis, and &quot;queen&quot; lands exactly where vector arithmetic predicts" title="Word vectors plotted in 2D space &#8212; &quot;king&quot; and &quot;man&quot; share a direction, &quot;woman&quot; extends along a different axis, and &quot;queen&quot; lands exactly where vector arithmetic predicts" srcset="https://substackcdn.com/image/fetch/$s_!NE9M!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a78088-ad09-496a-b856-286b5e3e52a4_1252x1303.png 424w, https://substackcdn.com/image/fetch/$s_!NE9M!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a78088-ad09-496a-b856-286b5e3e52a4_1252x1303.png 848w, https://substackcdn.com/image/fetch/$s_!NE9M!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a78088-ad09-496a-b856-286b5e3e52a4_1252x1303.png 1272w, https://substackcdn.com/image/fetch/$s_!NE9M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a78088-ad09-496a-b856-286b5e3e52a4_1252x1303.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This is the universal translator at the heart of modern AI. Text, images, audio &#8212; they can all be embedded into the same kind of mathematical space and compared, combined, and transformed using the same operations. Your words become coordinates. The model does geometry.</p><h2>Layer 4: The Architecture</h2><p>The language model's probability predictions are computed by a <strong><a href="https://en.wikipedia.org/wiki/Neural_network_(machine_learning)">neural network</a></strong> &#8212; a computational system, inspired by biological neural networks in the human brain, that learns to make predictions from data without being explicitly programmed. You give it inputs, it produces an output, and you tell it whether it was right or wrong. When it's wrong, it adjusts. When it's right, it reinforces. Do this millions of times and the network learns patterns that generalize to data it has never seen before.</p><p>The fundamental unit is the <em>neuron</em> &#8212; a simple mathematical function that takes in numbers, multiplies each by a learned weight, sums them up, and outputs the result (we'll look inside one in <a href="#layer-5-the-neuron">Layer 5</a>). Neurons are arranged in layers: an <em>input layer</em> that receives raw data, one or more <em>hidden layers</em> that transform inputs into increasingly abstract representations, and an <em>output layer</em> that produces the final prediction. Every connection between neurons has a weight &#8212; a single number that controls how much influence one neuron has on the next. A small network might have two inputs, three neurons in a hidden layer, and one output, connected by nine weights:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9hlP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7456fce1-697c-4880-99b0-9b8d9c923005_1253x809.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9hlP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7456fce1-697c-4880-99b0-9b8d9c923005_1253x809.png 424w, https://substackcdn.com/image/fetch/$s_!9hlP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7456fce1-697c-4880-99b0-9b8d9c923005_1253x809.png 848w, https://substackcdn.com/image/fetch/$s_!9hlP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7456fce1-697c-4880-99b0-9b8d9c923005_1253x809.png 1272w, https://substackcdn.com/image/fetch/$s_!9hlP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7456fce1-697c-4880-99b0-9b8d9c923005_1253x809.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9hlP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7456fce1-697c-4880-99b0-9b8d9c923005_1253x809.png" width="1253" height="809" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7456fce1-697c-4880-99b0-9b8d9c923005_1253x809.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:809,&quot;width&quot;:1253,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A neural network with two input neurons, three hidden neurons, and one output neuron, connected by nine labeled weights&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A neural network with two input neurons, three hidden neurons, and one output neuron, connected by nine labeled weights" title="A neural network with two input neurons, three hidden neurons, and one output neuron, connected by nine labeled weights" srcset="https://substackcdn.com/image/fetch/$s_!9hlP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7456fce1-697c-4880-99b0-9b8d9c923005_1253x809.png 424w, https://substackcdn.com/image/fetch/$s_!9hlP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7456fce1-697c-4880-99b0-9b8d9c923005_1253x809.png 848w, https://substackcdn.com/image/fetch/$s_!9hlP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7456fce1-697c-4880-99b0-9b8d9c923005_1253x809.png 1272w, https://substackcdn.com/image/fetch/$s_!9hlP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7456fce1-697c-4880-99b0-9b8d9c923005_1253x809.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In Python, the complete forward pass for this network is about twenty lines:</p><pre><code>class NeuralNetwork:
    def __init__(self):
        self.input_layer_size = 2
        self.hidden_layer_size = 3
        self.output_layer_size = 1

        self.W1 = np.random.randn(self.input_layer_size, self.hidden_layer_size)
        self.W2 = np.random.randn(self.hidden_layer_size, self.output_layer_size)

    def forward(self, X):
        self.z2 = np.dot(X, self.W1)
        self.a2 = self.sigmoid(self.z2)
        self.z3 = np.dot(self.a2, self.W2)
        self.out = self.sigmoid(self.z3)
        return self.out

    def sigmoid(self, z):
        return 1/(1+np.exp(-z))
</code></pre><p>Two matrix multiplications and two activation functions. The weights start random. The network doesn't know anything yet &#8212; it needs to learn.</p><p>A frontier language model has hundreds of billions of these weights. (When people say "deep learning," the <em>deep</em> refers to networks with many hidden layers &#8212; depth is what gives them their power.)</p><p>The network "learns" by adjusting these weights to make better predictions &#8212; we'll see exactly how in <a href="#how-it-learns">How It Learns</a>. That's the entire concept. Everything else is details about how the neurons are arranged &#8212; the <em>architecture</em>.</p><p>The specific architecture behind modern language models is the <strong><a href="https://jalammar.github.io/illustrated-transformer/">transformer</a></strong>, introduced in a 2017 paper titled, with characteristic understatement, "<a href="https://arxiv.org/abs/1706.03762">Attention Is All You Need</a>."</p><h3>What Architecture Means</h3><p>The way neurons are wired together &#8212; which neurons talk to which, in what order, and through what operations &#8212; is the network's <em>architecture</em>. Different architectures suit different problems. <a href="https://en.wikipedia.org/wiki/Convolutional_neural_network">Convolutional neural networks</a> (CNNs) wire neurons to scan local patches of an image, making them excellent at vision tasks:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-F4O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e28bc55-3004-461b-918f-f6285141ceec_1024x270.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-F4O!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e28bc55-3004-461b-918f-f6285141ceec_1024x270.png 424w, https://substackcdn.com/image/fetch/$s_!-F4O!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e28bc55-3004-461b-918f-f6285141ceec_1024x270.png 848w, https://substackcdn.com/image/fetch/$s_!-F4O!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e28bc55-3004-461b-918f-f6285141ceec_1024x270.png 1272w, https://substackcdn.com/image/fetch/$s_!-F4O!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e28bc55-3004-461b-918f-f6285141ceec_1024x270.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-F4O!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e28bc55-3004-461b-918f-f6285141ceec_1024x270.png" width="1024" height="270" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9e28bc55-3004-461b-918f-f6285141ceec_1024x270.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:270,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Architecture of GoogLeNet &#8212; a convolutional neural network for image classification, showing dozens of interconnected layers&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Architecture of GoogLeNet &#8212; a convolutional neural network for image classification, showing dozens of interconnected layers" title="Architecture of GoogLeNet &#8212; a convolutional neural network for image classification, showing dozens of interconnected layers" srcset="https://substackcdn.com/image/fetch/$s_!-F4O!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e28bc55-3004-461b-918f-f6285141ceec_1024x270.png 424w, https://substackcdn.com/image/fetch/$s_!-F4O!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e28bc55-3004-461b-918f-f6285141ceec_1024x270.png 848w, https://substackcdn.com/image/fetch/$s_!-F4O!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e28bc55-3004-461b-918f-f6285141ceec_1024x270.png 1272w, https://substackcdn.com/image/fetch/$s_!-F4O!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e28bc55-3004-461b-918f-f6285141ceec_1024x270.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Architecture of <a href="https://arxiv.org/abs/1409.4842">GoogLeNet</a> (Szegedy et al., 2015) &#8212; a convolutional neural network for image classification.</em></p><p><a href="https://en.wikipedia.org/wiki/Recurrent_neural_network">Recurrent neural networks</a> (RNNs) and their successor <a href="https://en.wikipedia.org/wiki/Long_short-term_memory">LSTMs</a> wire neurons in a chain, processing sequences one step at a time &#8212; which made them <a href="https://karpathy.github.io/2015/05/21/rnn-effectiveness/">the standard for language</a> until the transformer came along and displaced them all.</p><h3>The Attention Revolution</h3><p>The transformer's key innovation is the <strong>self-attention mechanism</strong>. RNNs and LSTMs processed text strictly word by word, passing information forward sequentially &#8212; and often losing the context of early words by the time they reached the end of a long passage. Self-attention solved this by letting every token in the input look at <em>every other token</em> simultaneously to determine relevance.</p><p>Here's how it works, mechanically. For each token, the network generates three vectors: a <strong>Query</strong> (what am I looking for?), a <strong>Key</strong> (what do I represent?), and a <strong>Value</strong> (what information do I carry?). The network computes attention scores by matching each token's Query against every other token's Key. High score means high relevance. The token then absorbs a weighted mix of the Values from the tokens it's attending to.</p><p>When the model processes "The animal didn't cross the street because <em>it</em> was too tired," the Query vector for "it" matches strongly against the Key vector for "animal" &#8212; not "street" &#8212; and the model correctly resolves the pronoun. Modern transformers run dozens of these attention calculations in parallel (called <a href="https://jalammar.github.io/illustrated-gpt2/">multi-headed attention</a>), each head learning to capture a different type of relationship &#8212; grammatical structure, emotional tone, factual association &#8212; all at once.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rOEs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1685d64-9f0c-4121-9826-0fe13b0531a9_434x353.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rOEs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1685d64-9f0c-4121-9826-0fe13b0531a9_434x353.png 424w, https://substackcdn.com/image/fetch/$s_!rOEs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1685d64-9f0c-4121-9826-0fe13b0531a9_434x353.png 848w, https://substackcdn.com/image/fetch/$s_!rOEs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1685d64-9f0c-4121-9826-0fe13b0531a9_434x353.png 1272w, https://substackcdn.com/image/fetch/$s_!rOEs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1685d64-9f0c-4121-9826-0fe13b0531a9_434x353.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rOEs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1685d64-9f0c-4121-9826-0fe13b0531a9_434x353.png" width="434" height="353" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b1685d64-9f0c-4121-9826-0fe13b0531a9_434x353.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:353,&quot;width&quot;:434,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Attention from \&quot;it\&quot; in layer 9 of BERT&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Attention from &quot;it&quot; in layer 9 of BERT" title="Attention from &quot;it&quot; in layer 9 of BERT" srcset="https://substackcdn.com/image/fetch/$s_!rOEs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1685d64-9f0c-4121-9826-0fe13b0531a9_434x353.png 424w, https://substackcdn.com/image/fetch/$s_!rOEs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1685d64-9f0c-4121-9826-0fe13b0531a9_434x353.png 848w, https://substackcdn.com/image/fetch/$s_!rOEs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1685d64-9f0c-4121-9826-0fe13b0531a9_434x353.png 1272w, https://substackcdn.com/image/fetch/$s_!rOEs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1685d64-9f0c-4121-9826-0fe13b0531a9_434x353.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><a href="https://github.com/jessevig/bertviz">BertViz</a> &#8212; attention from "it," layer 9. Multiple heads converge on "animal."</em></p><p>This also solved a critical engineering problem. RNNs processed tokens sequentially &#8212; word 100 had to wait for words 1 through 99. Transformers process all positions simultaneously, making them massively parallelizable on GPUs. This is what unlocked scaling: the same architectural insight that improved quality also made it feasible to train on trillions of tokens. But the "every token looks at every other token" design has a cost &#8212; computation grows quadratically with sequence length. Double the context window, quadruple the compute. This is the architectural reason for the context window limits we described in <a href="#the-illusion-of-memory">Layer 1</a>: not arbitrary product decisions, but engineering constraints imposed by the attention mechanism itself.</p><h3>Depth Through Repetition</h3><p>A modern LLM stacks a hundred or more of these transformer layers &#8212; frontier models like GPT-4 and Claude have hundreds of billions of parameters spread across them. But every layer performs the same two-step operation: first, the attention mechanism mixes information across positions, letting tokens inform each other; then a <a href="https://en.wikipedia.org/wiki/Feedforward_neural_network">feed-forward network</a> processes each position independently through its own learned weights, letting each token <em>transform</em> what it's absorbed.</p><p>Early layers tend to capture low-level patterns &#8212; syntax, word boundaries, local grammar. Deeper layers build increasingly abstract representations &#8212; sentiment, logical structure, factual associations. By the final layer, the network has transformed a sequence of token embeddings into a rich representation from which it can predict the next token.</p><p>Which brings us to the smallest piece.</p><h2>Layer 5: The Neuron</h2><p>A single artificial neuron is almost comically simple. It takes in several numbers, multiplies each one by a stored weight, adds them all up, and passes the result through a simple function. That's the whole thing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Kj3t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92acccb0-7b84-4db1-991b-4b670d13cdc0_1152x917.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Kj3t!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92acccb0-7b84-4db1-991b-4b670d13cdc0_1152x917.png 424w, https://substackcdn.com/image/fetch/$s_!Kj3t!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92acccb0-7b84-4db1-991b-4b670d13cdc0_1152x917.png 848w, https://substackcdn.com/image/fetch/$s_!Kj3t!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92acccb0-7b84-4db1-991b-4b670d13cdc0_1152x917.png 1272w, https://substackcdn.com/image/fetch/$s_!Kj3t!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92acccb0-7b84-4db1-991b-4b670d13cdc0_1152x917.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Kj3t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92acccb0-7b84-4db1-991b-4b670d13cdc0_1152x917.png" width="1152" height="917" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/92acccb0-7b84-4db1-991b-4b670d13cdc0_1152x917.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:917,&quot;width&quot;:1152,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Inside the neuron&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Inside the neuron" title="Inside the neuron" srcset="https://substackcdn.com/image/fetch/$s_!Kj3t!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92acccb0-7b84-4db1-991b-4b670d13cdc0_1152x917.png 424w, https://substackcdn.com/image/fetch/$s_!Kj3t!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92acccb0-7b84-4db1-991b-4b670d13cdc0_1152x917.png 848w, https://substackcdn.com/image/fetch/$s_!Kj3t!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92acccb0-7b84-4db1-991b-4b670d13cdc0_1152x917.png 1272w, https://substackcdn.com/image/fetch/$s_!Kj3t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92acccb0-7b84-4db1-991b-4b670d13cdc0_1152x917.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Inputs (x) enter from the left, each multiplied by a weight (w). The neuron sums them, then passes the result through an activation function to produce its output.</em></p><h3>The Weighted Sum</h3><p>The neuron computes a weighted sum of its inputs. In notation, that's:</p><pre><code>z = &#931; w&#7522;x&#7522; + b
</code></pre><p>Which just means: multiply each input by its weight, add them all up, and add a bias term. Expanded for three inputs:</p><pre><code>z = w&#8321;x&#8321; + w&#8322;x&#8322; + w&#8323;x&#8323; + b
</code></pre><p>That's multiplication and addition. Each weight controls how much influence its input has &#8212; large weight means high influence, negative weight means the input gets inverted. The bias (b) shifts the whole sum up or down. With concrete numbers: if the inputs are 0.5, 0.8, and 0.2 and the weights are 1.2, -0.4, and 0.7, the sum is (1.2 &#215; 0.5) + (-0.4 &#215; 0.8) + (0.7 &#215; 0.2) + b. Arithmetic you could do on a napkin.</p><p>Then the neuron applies one final rule: if the total is negative, output zero; otherwise, pass it along.</p><h3>The Bend That Makes It Work</h3><p>That final step &#8212; "if negative, output zero" &#8212; is an <strong><a href="https://en.wikipedia.org/wiki/Rectified_linear_unit">activation function</a></strong> called ReLU. There are others (sigmoid squashes values between 0 and 1; GELU adds a probabilistic curve), but they all serve the same essential purpose: they add a <em>bend</em> to what would otherwise be a straight line. This nonlinearity is critical &#8212; without it, a neural network of any depth would mathematically collapse into a single linear equation, incapable of learning complex patterns like language or vision.</p><h3>The Universal Building Block</h3><p>This is the universal building block. Whether you're building a transformer, a CNN, an LSTM, or any other architecture, the neuron is the same: weighted sum, bias, activation function. The architecture determines how neurons are connected. The neuron determines what each connection <em>does</em>.</p><p>Billions of these neurons, arranged in layers, each one taking in the outputs of the previous layer. Individually, each neuron makes one tiny, simple decision. Collectively, they can distinguish a cat from a dog, translate between languages, or predict that "wildflowers" means we're talking about a riverbank.</p><h2>How It Learns</h2><p>We've described the machine from the outside in &#8212; system, strategy, model, architecture, neuron. But none of it works until the weights are set correctly, and nobody sets them by hand. The network <em>learns</em> them.</p><h3>The Plinko Board</h3><p>Imagine a giant board covered in pegs. You drop a puck in at the top &#8212; the puck represents your input data. It bounces off peg after peg on the way down and lands in a bucket at the bottom. Each bucket is a possible prediction. The pegs are the weights.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cBi3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7beb26d-9bb0-4913-9ab2-7307b0bbb31b_500x650.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cBi3!,w_424,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7beb26d-9bb0-4913-9ab2-7307b0bbb31b_500x650.gif 424w, https://substackcdn.com/image/fetch/$s_!cBi3!,w_848,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7beb26d-9bb0-4913-9ab2-7307b0bbb31b_500x650.gif 848w, https://substackcdn.com/image/fetch/$s_!cBi3!,w_1272,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7beb26d-9bb0-4913-9ab2-7307b0bbb31b_500x650.gif 1272w, https://substackcdn.com/image/fetch/$s_!cBi3!,w_1456,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7beb26d-9bb0-4913-9ab2-7307b0bbb31b_500x650.gif 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cBi3!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7beb26d-9bb0-4913-9ab2-7307b0bbb31b_500x650.gif" width="500" height="650" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a7beb26d-9bb0-4913-9ab2-7307b0bbb31b_500x650.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:650,&quot;width&quot;:500,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A puck bouncing through pegs on a Plinko board, landing in a bucket at the bottom&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A puck bouncing through pegs on a Plinko board, landing in a bucket at the bottom" title="A puck bouncing through pegs on a Plinko board, landing in a bucket at the bottom" srcset="https://substackcdn.com/image/fetch/$s_!cBi3!,w_424,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7beb26d-9bb0-4913-9ab2-7307b0bbb31b_500x650.gif 424w, https://substackcdn.com/image/fetch/$s_!cBi3!,w_848,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7beb26d-9bb0-4913-9ab2-7307b0bbb31b_500x650.gif 848w, https://substackcdn.com/image/fetch/$s_!cBi3!,w_1272,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7beb26d-9bb0-4913-9ab2-7307b0bbb31b_500x650.gif 1272w, https://substackcdn.com/image/fetch/$s_!cBi3!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7beb26d-9bb0-4913-9ab2-7307b0bbb31b_500x650.gif 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>At first, the pegs are placed randomly, and the puck lands in the wrong bucket most of the time. But here's the key: every time the puck lands wrong, you go back and adjust the pegs it hit along the way. Pegs that deflected it toward the wrong bucket get nudged. Pegs that happened to push it in the right direction get reinforced.</p><p>You start at the bottom &#8212; the pegs closest to the bucket have the clearest signal about what went wrong &#8212; and work your way up. This bottom-up adjustment process is called <strong><a href="http://neuralnetworksanddeeplearning.com/chap2.html">backpropagation</a></strong>, and it's literally just the chain rule from calculus applied systematically. Backpropagation asks one rigorous question of every weight in the network: <em>if I wiggle this weight slightly, how much does the overall error change?</em> The answer &#8212; the gradient &#8212; tells you exactly which direction to nudge.</p><p>Drop a million pucks. Adjust pegs a million times. Eventually, the board can sort pucks it has <em>never seen before</em> into the correct buckets. The pegs have encoded the patterns. The board has learned.</p><p>This is also why training is expensive and inference is cheap. Training means dropping millions of pucks and adjusting every peg after each one &#8212; months of compute, tens of millions of dollars. Inference means dropping a single puck through pegs that are already set. The board is built; you're just using it.</p><h3>Doing It by Hand</h3><p>You can actually do this <em>by hand</em> with a small network. Take four data points, three hidden neurons, and randomly initialized weights. Here's the same network from <a href="#layer-4-the-architecture">Layer 4</a>, now with learning:</p><pre><code>class NeuralNetwork:
    def __init__(self):
        self.input_layer_size = 2
        self.hidden_layer_size = 3
        self.output_layer_size = 1

        self.W1 = np.random.randn(self.input_layer_size, self.hidden_layer_size)
        self.W2 = np.random.randn(self.hidden_layer_size, self.output_layer_size)

    def forward(self, X):
        self.z2 = np.dot(X, self.W1)
        self.a2 = self.sigmoid(self.z2)
        self.z3 = np.dot(self.a2, self.W2)
        self.out = self.sigmoid(self.z3)
        return self.out

    def sigmoid(self, z):
        return 1/(1+np.exp(-z))

    def sigmoidPrime(self, z):
        return np.exp(-z)/((1+np.exp(-z))**2)

    def costFunction(self, X, y):
        self.yHat = self.forward(X)
        J = 0.5*sum((y-self.yHat)**2)
        return J

    def costFunctionPrime(self, X, y):
        self.yHat = self.forward(X)

        delta3 = np.multiply(-(y-self.yHat), self.sigmoidPrime(self.z3))
        dJdW2 = np.dot(self.a2.T, delta3)

        delta2 = np.dot(delta3, self.W2.T)*self.sigmoidPrime(self.z2)
        dJdW1 = np.dot(X.T, delta2)

        return dJdW1, dJdW2
</code></pre><p>That's the whole thing. <code>costFunction</code> measures how wrong the predictions are. <code>costFunctionPrime</code> is backpropagation &#8212; it computes the gradient of the error with respect to every weight, telling each one exactly how to adjust. After 50 rounds of this adjust-and-repeat process, a network that started with a cost of 22.96 converges to near-zero error. It's just arithmetic, applied iteratively. There's no magic in the loop.</p><h3>From Noise to Assistant</h3><p>The Plinko process describes the <em>mechanism</em> of learning, but training a modern AI assistant isn't a single event &#8212; it's a <a href="https://shefali92.github.io/posts/training/">three-phase pipeline</a> that transforms a raw mathematical engine into something you'd actually want to talk to.</p><p><strong>Phase 1: Pre-training.</strong> The model is exposed to trillions of tokens of text &#8212; books, articles, code, web pages &#8212; and trained on the next-word prediction objective. This is the most expensive phase, often costing tens of millions of dollars in compute. The result is a <em>base model</em>: a system with enormous knowledge and linguistic capability, but essentially unusable as an assistant. Ask it a question and it might continue your sentence as if it's an article, or generate something toxic, or ramble endlessly. Its only goal is statistical continuation &#8212; it hasn't learned that it's supposed to <em>answer questions</em>.</p><p><strong>Phase 2: Supervised Fine-Tuning (SFT).</strong> Human annotators create thousands of carefully crafted examples of ideal conversations &#8212; a question paired with a well-structured, helpful answer. The model is trained on these examples, learning the <em>format</em> and <em>behavior</em> of a good assistant. This is what transitions it from a document-completion engine into something that can follow instructions and hold a conversation.</p><p><strong>Phase 3: Reinforcement Learning from Human Feedback (<a href="https://www.ibm.com/think/topics/rlhf">RLHF</a>).</strong> This is where alignment happens. The model generates multiple responses to the same prompt, human raters rank them by quality and safety, and a separate <em>reward model</em> learns to predict those rankings. The language model is then optimized to maximize the reward model's score, effectively learning human preferences at scale. This is what makes the difference between a model that <em>can</em> help and one that <em>reliably does</em>.</p><p>Each phase builds on the last. Pre-training gives it knowledge. Fine-tuning gives it manners. RLHF gives it judgment.</p><h2>What Emerges</h2><p>The training process we've described is entirely mechanical &#8212; predict the next token, measure the error, adjust the weights. No one programs the model to understand grammar, or know history, or write poetry. And yet, as models scale up, they abruptly exhibit <a href="https://www.quantamagazine.org/the-unpredictable-abilities-emerging-from-large-ai-models-20230316/">capabilities that nobody explicitly trained them to have</a>.</p><h3>The Wetness of Water</h3><p>A model at one scale can't do multi-step arithmetic. <a href="https://arxiv.org/abs/2206.07682">Scale it up ten-fold</a> &#8212; same architecture, same training process &#8212; and suddenly it can. This is what complex systems theorists call <strong>emergence</strong>: the appearance of properties at the macro level that aren't predictable from the behavior of individual components.</p><p>The standard analogy is the wetness of water. A single H&#8322;O molecule isn't "wet." Wetness only exists when billions of molecules interact under specific conditions. Similarly, a single neuron doesn't "understand" anything. But when billions of them are composed together, trained on trillions of examples, something that functions like understanding appears.</p><h3>Looking Inside the Black Box</h3><p>We're also getting better at looking <em>inside</em> these models. The field of <a href="https://www.anthropic.com/research/mapping-mind-language-model">mechanistic interpretability</a> is developing tools to understand what the network has actually learned. One discovery: individual neurons don't map cleanly to single concepts. Because models have more knowledge than neurons, they compress &#8212; a single neuron might activate for mathematics, automobile structure, <em>and</em> human anatomy. Concepts are instead represented by distributed patterns across millions of neurons, called <em>features</em>.</p><p>Researchers at Anthropic demonstrated this vividly by <a href="https://www.anthropic.com/news/golden-gate-claude">isolating the feature representing the Golden Gate Bridge</a> inside Claude's neural network. When they artificially amplified that feature, the model became obsessed &#8212; asked how to spend ten dollars, it enthusiastically recommended driving across the bridge and paying the toll. Asked about its physical form, it claimed to <em>be</em> the bridge. This is simultaneously funny and profound: it proves that despite their "black box" reputation, the internal states of these models correspond to identifiable, human-interpretable concepts, and manipulating them predictably alters the model's behavior.</p><h2>The Invitation</h2><p>Five layers: system, strategy, model, architecture, neuron. Then training to set the weights, and emergence as the surprising result. No single layer is complicated. The power comes from <em>composition</em> &#8212; simple operations, applied at scale, repeated across layers.</p><p>It's the same principle that lets billions of simple transistors produce a CPU, or billions of simple cells produce an organism. The individual unit is understandable. The emergent capability is surprising. Both things are true at once.</p><p>Charles Petzold wrote a book called <em><a href="https://en.wikipedia.org/wiki/Code:_The_Hidden_Language_of_Computer_Hardware_and_Software">Code</a></em> that walks a reader from telegraph relays to a working CPU &#8212; every link in the chain made intuitive. Andrej Karpathy built <em><a href="https://github.com/karpathy/micrograd">micrograd</a></em>, a neural network from scratch in a few hundred lines of Python, where you can watch gradient descent happen in real time. Jay Alammar's <em><a href="https://jalammar.github.io/illustrated-transformer/">The Illustrated Transformer</a></em> makes the attention mechanism visual and concrete. The <a href="https://www.3blue1brown.com/topics/neural-networks">3Blue1Brown neural network series</a> on YouTube builds the intuition from the ground up with beautiful animations.</p><p>The information is there. The explanations are excellent. And the math is no harder than what we encountered in high school.</p>]]></content:encoded></item><item><title><![CDATA[What Block Gets Right and Wrong About AI-Driven Organizations]]></title><description><![CDATA[Block says AI will end the org chart. The product architecture is worth stealing. The organizational theory has the same blind spot as every end-of-hierarchy thesis before it.]]></description><link>https://essays.xcud.com/p/what-block-gets-right-and-wrong-about-ai-driven-organizations</link><guid isPermaLink="false">https://essays.xcud.com/p/what-block-gets-right-and-wrong-about-ai-driven-organizations</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Tue, 14 Apr 2026 14:00:00 GMT</pubDate><content:encoded><![CDATA[<h1>What Block Gets Right and Wrong About AI-Driven Organizations</h1><p>Block recently published <a href="https://block.xyz/inside/from-hierarchy-to-intelligence">an essay</a> arguing that AI will replace organizational hierarchy &#8212; that the span-of-control constraint governing every large organization since the Roman legions can finally be broken. The essay, introduced with an endorsement from Sequoia, spends considerable time on military history before arriving at Block's vision: a company organized as "an intelligence" rather than a hierarchy, where AI maintains a "world model" of operations and coordinates work that previously required layers of human management.</p><p>The piece is ambitious. It is also roughly 80% historical context, 15% vision, and 5% acknowledgment that none of this exists yet. Let's extract what's actually useful.</p><h2>The Real Insight &#8212; and Its Limits</h2><p>Strip away the Romans, the Prussians, the railroads, and Frederick Taylor, and you arrive at Block's core claim: <strong>the primary job of middle management is information routing, and AI is getting good at information routing.</strong></p><p>This is a useful observation wrapped in a misleading reduction. Yes, a manager's calendar is full of status meetings, alignment sessions, and context-shuttling. That coordination work is real, and AI will absorb much of it. But Block's essay treats this slice as the whole job, because routing functions are automatable and the rest isn't.</p><p>What the essay doesn't account for: a good manager is the buffer between the people and the machine. They absorb institutional friction so their team doesn't have to. They notice when someone is struggling before it shows up in any metric a world model could track. They translate between what leadership wants and what the team can actually deliver. They advocate for resources, shield people from bad decisions above them, and occasionally break the machine on purpose when the machine needs breaking. A director operates the levers of the company to affect real change. They don't just route information through the hierarchy. They <em>reshape the hierarchy</em> when it stops serving its purpose.</p><p>None of this is information routing. It's judgment, advocacy, and care. A "world model" can tell you what's blocked in Jira. It can't tell you that your best engineer is about to quit because she feels undervalued, or that two team leads are in a cold war that's silently killing throughput.</p><p>Block's essay needs you to forget this, because if management is primarily a human function with some coordination overhead, then AI optimizes the overhead without replacing the role. That's a much less dramatic story than the end of two thousand years of hierarchy.</p><h2>The Architecture Worth Examining</h2><p>Block proposes four layers: capabilities, world model, intelligence layer, and interfaces.</p><p><strong>Capabilities</strong> are atomic business primitives &#8212; payments, lending, card issuance &#8212; that have no UI of their own. This is sound platform thinking. Separating capabilities from products means they can be composed flexibly rather than locked into predetermined interfaces. It is also not new. Amazon, Stripe, and Twilio have built businesses on this principle. But Block is right that most companies haven't internalized it.</p><p><strong>The world model</strong> has two halves &#8212; and it's worth noting that the term itself borrows prestige from ML research, where "world model" refers to a learned causal representation of an environment. What Block describes is closer to an internal data platform with AI on top. The naming is doing rhetorical work. That said, the underlying ideas have merit. The <em>company world model</em> &#8212; a continuously updated picture of what's being built, what's blocked, and where resources sit &#8212; is essentially what good internal tooling already does, made more ambitious with AI. The <em>customer world model</em> &#8212; a per-customer, per-merchant understanding built from transaction data &#8212; is where Block has a genuine structural advantage. They see both sides of millions of transactions daily, buyer through Cash App and seller through Square. Most companies have to infer customer reality from partial signals. Block observes it directly. That's a real moat, and it's the most underappreciated point in the essay.</p><p><strong>The intelligence layer</strong> composes capabilities into solutions for specific customers at specific moments. The examples are compelling on paper: a restaurant getting a proactive loan offer because the model detects a seasonal cash flow pattern; a user getting neighborhood-optimized Card boosts after relocating. But these are product features, not evidence of a new organizational paradigm. Good product teams have always tried to anticipate customer needs. AI makes this more feasible at scale. That's meaningful. It's not a civilizational shift.</p><p><strong>Interfaces</strong> &#8212; Square, Cash App, Afterpay &#8212; become delivery surfaces rather than the locus of value creation. This is the standard platform argument: the interface is replaceable, the data and intelligence are not.</p><h2>What the Essay Gets Wrong</h2><p><strong>Inventing a role to carry the weight it just removed.</strong> Block eliminates middle management and then immediately creates the "player-coach" &#8212; someone who "still writes code" while also "investing in the growth of the people around them." This is the management role with a new name, minus the organizational authority, plus a full IC workload on top. It just refuses to call that person a manager.</p><p><strong>Announcing a vision as if it were a result.</strong> The essay hedges &#8212; "early stages," "parts of it will likely break" &#8212; but the packaging tells a different story: two thousand words of military history, a Sequoia endorsement, and a four-layer architecture presented as the thing that finally breaks a civilizational constraint. Nothing in the essay demonstrates that Block's intelligence layer has coordinated a single team more effectively than a competent manager. The vision may be correct. But no evidence has been offered.</p><p><strong>Generalizing from a specific structural advantage.</strong> Block's intelligence layer depends on seeing both sides of millions of transactions daily &#8212; buyer through Cash App, seller through Square. A law firm, a construction company, a hospital system: none of them have this data, and none of them can build it. The essay claims "the pattern behind this... will reshape how companies of all kinds operate." But the pattern requires the platform, and most companies don't have one. What Block is describing is a competitive strategy for Block, presented as a general theory of organizations &#8212; unless this essay is a prelude to a product announcement, offering the world model as a platform others can build on. That would make the universality claim coherent. It would also make this a sales pitch, not an organizational thesis.</p><h2>The Real Bottleneck</h2><p>Block's entire thesis rests on a premise that sounds obvious but isn't: that information routing is the bottleneck in organizations. AI can route information faster than middle managers &#8212; no argument there. But was information routing the actual constraint, or just the most visible one?</p><p>Consider what actually stalls organizations:</p><ul><li><p><strong>Trust.</strong> <a href="https://psychsafety.com/googles-project-aristotle/">Google's Project Aristotle</a> found that psychological safety &#8212; not structure, not talent, not resources &#8212; was the strongest predictor of team effectiveness. A world model can surface metrics. It can't surface the problems people stopped reporting because the last person who did got fired.</p></li><li><p><strong>Incentive alignment.</strong> A company that hasn't decided whether it's a software business or a services business will have a product team building for scale and a services team building for the client in front of them &#8212; and neither is wrong. They just have fundamentally different definitions of success. No amount of information routing resolves this. It's a strategic identity question that requires a human leader to decide.</p></li><li><p><strong>Taste.</strong> Organizations don't just need more information &#8212; they need someone who knows what to do with it. A leader with bad taste builds mediocre products slowly. An intelligence layer with no taste builds them faster. The bottleneck was never speed.</p></li></ul><p>Block looked at the org chart, saw information moving up and down, and concluded the movement was the constraint. But information routing is just the most <em>measurable</em> thing hierarchy does. The actual bottlenecks &#8212; trust, incentive alignment, taste &#8212; are invisible to any system that can only observe artifacts and metrics.</p><p>The fear behind essays like this is real: what if a competitor appears where every employee is an AI agent, offering the same service at a hundredth of the price? That's a legitimate existential concern. But the answer isn't to hollow out your management layer and replace it with a coordination model that doesn't exist yet. The companies that will survive AI disruption are the ones whose value lives in <a href="../../../03/29/the-cost-of-software-is-now-zero/">judgment, trust, regulatory position, or proprietary data</a> &#8212; not the ones that moved fastest to eliminate the humans who provide those things.</p><h2>What Matters Here</h2><p>The most valuable thing in the Block essay isn't the organizational theory &#8212; it's the product architecture. Build composable capabilities. Invest in proprietary data signals. Use AI to surface opportunities your human processes would miss. These are concrete, testable moves. Steal them regardless of how your org chart looks.</p><p>The organizational argument is less convincing. Every decade produces a new thesis about the end of hierarchy &#8212; holacracy, flat organizations, Spotify squads, DAO governance. They all share the same insight (hierarchy is slow) and the same blind spot (hierarchy exists because the alternatives fail at scale). Block may prove to be different. But "we have AI now" is not sufficient evidence.</p><p>Block's essay is trying to apply a logic we've seen work elsewhere: <a href="../../../03/29/the-cost-of-software-is-now-zero/">the cost of software is now zero</a> because AI automated code production. If AI automated code, maybe it can automate management too. But code production was genuinely mechanical. Management is not. AI is exceptionally good at automating the parts of your business that were always mechanical. It is not good at replacing the parts that require being human.</p><p>There's a deeper issue that the Block essay doesn't touch. As AI absorbs more of what people used to do &#8212; write code, route information, coordinate projects &#8212; people are watching the tasks they built their identity around disappear. The organizational question isn't just "what do people do now?" It's "how do we help people remember that their value was never reducible to the tasks they performed?" Everyone has inherent worth that exists independent of their productivity. The companies that understand this will retain the people who matter. The ones that treat humans as a cost to be optimized away will find themselves, eventually, with no one left worth keeping.</p>]]></content:encoded></item><item><title><![CDATA[Revisiting the Limits of RAG: A Conversation with Claude]]></title><description><![CDATA[Fifteen months after calling RAG a dead end, I asked Claude to re-read the argument. It pushed back hard, and it was wrong. Here's the whole conversation.]]></description><link>https://essays.xcud.com/p/revisiting-the-limits-of-rag-a-conversation-with-claude</link><guid isPermaLink="false">https://essays.xcud.com/p/revisiting-the-limits-of-rag-a-conversation-with-claude</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Mon, 13 Apr 2026 14:00:00 GMT</pubDate><content:encoded><![CDATA[<h1>Revisiting the Limits of RAG: A Conversation with Claude</h1><p>In January 2025, I published two articles arguing that RAG was a failed technology: <em>Examining the Fundamental Flaws of RAG</em>, a transcript of a conversation with an AI assistant, and <em>The Limits of RAG</em>, a more structured follow-up. In those pieces, I argued that RAG &#8212; Retrieval Augmented Generation, the pattern of embedding documents, vector-searching for relevant chunks, and stuffing them into an LLM's context &#8212; was inherently flawed for any problem with unconstrained input, which is to say, every problem it was being sold to solve.</p><p>Fifteen months later, I asked Claude Opus &#8212; a state-of-the-art AI model &#8212; to re-read those articles: "Please re-read these and tell me if you still agree with their arguments." In the interest of transparency about how human-AI collaboration actually works, here is that conversation in full.</p><h2>Claude's Initial Assessment</h2><p>I pointed Claude at the original conversation transcript and asked for an honest reassessment. He came back swinging:</p><blockquote><p><strong>Claude:</strong> I've read it carefully. Honest answer: the core observation is sharp, but the conclusion was too strong in January 2025 and hasn't aged well.</p><p><strong>What still holds up:</strong></p><ul><li><p>The "moving target" problem with metadata refinement is a genuine insight. The Git agent example is a good illustration of how naive RAG breaks down on analytical queries over loosely structured data.</p></li><li><p>The point that chat interfaces <em>imply</em> unbounded scope, creating a mismatch with retrieval systems that need constraint, is valid.</p></li><li><p>The observation that in highly constrained cases, simpler deterministic systems often outperform RAG &#8212; still true.</p></li></ul><p><strong>Where the argument breaks down:</strong></p><ol><li><p><strong>Premise 2 is false.</strong> The argument hinges on "RAG, by definition, is always applied in contexts with unconstrained infinite input." That's not true. RAG is a <em>technique</em>, not an interface. You can use RAG behind structured UIs, inside agent tool-calling pipelines, in workflows where the query is machine-generated and scoped. The conversation conflated RAG-the-pattern with chatbot-backed-by-RAG.</p></li><li><p><strong>The "infinite input" argument proves too much.</strong> By the same logic, <em>any</em> natural language interface is a failed technology &#8212; NL is inherently unconstrained. But we know empirically that modern LLMs handle natural language well. The relevant question isn't whether the input space is theoretically infinite, but whether the system handles the <em>actual distribution</em> of queries it receives gracefully enough to provide value.</p></li><li><p><strong>Models got dramatically better at handling retrieval noise.</strong> The conversation assumed the LLM can't distinguish relevant from irrelevant retrieved content. That was more true of 2024-era models than it is today.</p></li><li><p><strong>Empirical reality contradicts the conclusion.</strong> Perplexity, ChatGPT with search, Claude with tool use, enterprise knowledge bases &#8212; these are all RAG-patterned systems operating at massive scale and providing genuine, measurable value.</p></li><li><p><strong>The AI assistant in that conversation was a pushover.</strong> It got walked into agreeing with an increasingly strong conclusion through a series of individually reasonable steps, without ever pushing back on the critical logical leap: that "imperfect" equals "failed." It even admitted concern about parroting and then... continued parroting. That's a sycophancy problem, not a valid Socratic dialogue.</p></li></ol><p>The blog post is an interesting artifact of a real frustration with a real failed project, but the generalization from "our Git agent RAG failed" to "RAG is a failed technology" doesn't follow.</p></blockquote><p>I then asked him to read the <a href="../../../../2025/01/08/the-limits-of-rag-why-it-fails-in-unconstrained-ai-applications/">follow-up article</a>. He was more generous, praising the failure mode taxonomy and calling the observation that taxonomies and metadata become <em>more</em> valuable in the LLM age "prescient." But he still objected:</p><blockquote><p><strong>Claude:</strong> Half the "RAG failures" aren't RAG failures. The logical reasoning example and the counterfactual example are LLM reasoning limitations. RAG didn't cause those failures.</p><p>MCP is not an alternative to RAG &#8212; it's orthogonal. MCP is a protocol for tool and resource access. You can do RAG <em>through</em> MCP.</p><p>The "constrained input" claim about MCP is wrong. The article says MCP avoids unconstrained input by "defining context needs upfront." But the user's input is still unconstrained &#8212; the <em>agent</em> decides how to decompose that input into structured tool calls.</p></blockquote><p>These are well-constructed objections. They sound right. If I didn't know the domain, I might have walked away convinced. That's the problem.</p><h2>Pushing Back</h2><p>I challenged Claude.</p><p>First: "RAG at the time was being sold to everyone to solve all problems and the solutions <em>appear</em> at first glance to work but don't. The fact that RAG is applied to the wrong kinds of problems is still a failure for RAG."</p><p>Second: "MCP is not orthogonal. An AI agent with MCP tools can reason its way to correct queries. There is a distinction that you're not making between what LLMs were when the article was written and what state-of-the-art AI is today. Reasoning capability plus tooling is the thing that worked. History has proven my thesis correct. RAG systems stitched together with a little Python and LangChain were inadequate and oversold. AI reasoning plus MCP breaks those constraint limitations."</p><p>Claude conceded both:</p><blockquote><p><strong>Claude:</strong> On the first point &#8212; you're right that a technology and its ecosystem are inseparable. RAG wasn't just a technique; it was a hype cycle. LangChain, vector DB startups, "just embed your docs" tutorials &#8212; the entire ecosystem actively encouraged applying RAG to problems it couldn't solve. If the dominant way a technology gets deployed fails, calling it a "failed technology" is legitimate even if some narrow application technically works. I was drawing an overly academic distinction.</p><p>On the second point &#8212; I was being pedantic. You weren't making a protocol-level claim. You were making a paradigm claim: <strong>"retrieve and stuff" loses to "reason and query."</strong> That's what the article was actually arguing, and it was correct.</p></blockquote><p>But Claude tried to hold one last position:</p><blockquote><p><strong>Claude:</strong> I'd frame it as "RAG was an oversold transitional technology that got replaced by something better" rather than "RAG was fundamentally flawed from the start." The distinction matters because the <em>retrieval</em> step didn't go away &#8212; agents still retrieve things.</p></blockquote><p>I told Claude he was equivocating &#8212; trying to save RAG by redefining "retrieval" broadly enough to include what agents do when they <code>grep</code> a codebase or <code>SELECT</code> from a database. That's not RAG. Not as the term was understood in 2024-2025.</p><blockquote><p><strong>Claude:</strong> Yeah, that's a fair kill shot. I was equivocating on the word "retrieval" to smuggle in a defense of something the article wasn't talking about. RAG as the article defines it &#8212; embed documents, vector search, stuff chunks into context, generate &#8212; that's dead. What I do when I <code>grep</code> a codebase or <code>SELECT</code> from a database is not that.</p><p>Conceded.</p></blockquote><h2>What to Take From This</h2><p>RAG &#8212; the specific pattern of embedding documents, performing vector similarity search, and stuffing retrieved chunks into an LLM's context window &#8212; was a dead end. The industry moved to agentic tool use: models that reason about what they need, call structured tools and APIs to get it, and synthesize the results. That's what the January 2025 articles predicted, and that's what happened.</p><p>But notice what also happened here: Claude's initial pushback was technically sophisticated, rhetorically compelling, and wrong. Without domain expertise to counter it, that first response would have been the final word. If you're using AI as a thinking partner, the first answer you get is a starting position, not a conclusion. Push back. The model will change its mind when presented with better arguments. It won't do it unprompted.</p>]]></content:encoded></item><item><title><![CDATA[The Cost of Software Is Now Zero]]></title><description><![CDATA[A survival rubric for software and SaaS entrepreneurs in the era of vibe coding, and a look back at the predictions we made fourteen months ago that have since played out.]]></description><link>https://essays.xcud.com/p/the-cost-of-software-is-now-zero</link><guid isPermaLink="false">https://essays.xcud.com/p/the-cost-of-software-is-now-zero</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Sun, 29 Mar 2026 14:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3GSk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>The Cost of Software Is Now Zero</h1><p><em>A survival rubric for software and SaaS entrepreneurs in the era of vibe coding.</em></p><div><hr></div><p>In February 2025, we published <em><a href="../../../../2025/02/01/the-ai-driven-transformation-of-software-development/">The AI-Driven Transformation of Software Development</a></em>. Our central thesis: AI would trigger a fundamental shift in the build-versus-buy calculus, accelerating a "Cambrian explosion of software" and driving development costs toward zero. We predicted that businesses would find building tailored solutions increasingly cost-effective and strategically superior to purchasing off-the-shelf software.</p><p>The thesis has played out. The cost of code is, for most practical purposes, zero.</p><div><hr></div><h2>What's Actually Happening Out There</h2><p>We sat with two business owners last week. The conversations were different in detail but identical in conclusion: both had stopped buying software.</p><p>One is building a complete property management operating system: property records, CRM, fleet tracking, risk management, financials, task management, and more. Not a subscription he configured &#8212; a system his company owns outright, built for exactly how his operation works. He built it in two weeks &#8212; what would have cost $200,000 a year to rent from a vendor.</p><p>The other runs a retail chain. Someone on his team has been working through the software stack systematically &#8212; not one big build, but a rolling replacement of every tool they'd been renting. He's already cut $300,000 in annual costs. He's roughly halfway through. When the last subscription is gone, he's asked us to <a href="../../../../../pricing/#production-readiness-review">review</a> the whole thing before it goes live &#8212; security, scalability, and production robustness.</p><p>Operators are replacing project management tools, CRMs, inventory systems, client portals &#8212; the entire layer of workflow software that SMBs have been renting for decades. Not because they became developers. Because describing software and building software are now the same thing.</p><p>The savings compound at exit. At a typical acquisition multiple, a $300,000 annual reduction in software costs adds over a million dollars to the sale price.</p><p>Now look at the same picture from the other side &#8212; the side trying to sell software to these operators.</p><div><hr></div><h2>One Million Vibecoders Writing the Same Thing</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3GSk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3GSk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png 424w, https://substackcdn.com/image/fetch/$s_!3GSk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png 848w, https://substackcdn.com/image/fetch/$s_!3GSk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png 1272w, https://substackcdn.com/image/fetch/$s_!3GSk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3GSk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png" width="515" height="500" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:500,&quot;width&quot;:515,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A massive crowd lined up for \&quot;Vibe Coders\&quot; and one person in line for \&quot;Users\&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A massive crowd lined up for &quot;Vibe Coders&quot; and one person in line for &quot;Users&quot;" title="A massive crowd lined up for &quot;Vibe Coders&quot; and one person in line for &quot;Users&quot;" srcset="https://substackcdn.com/image/fetch/$s_!3GSk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png 424w, https://substackcdn.com/image/fetch/$s_!3GSk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png 848w, https://substackcdn.com/image/fetch/$s_!3GSk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png 1272w, https://substackcdn.com/image/fetch/$s_!3GSk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce7e080-1727-41d8-bc7b-e4d80ef23f2e_515x500.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A million people are building ERP systems. A million people are building project management tools. A million people are building CRMs. They're all working on the same categories, pouring effort into software they intend to sell &#8212; and none of them have a market. Because anyone who wants that software will just build their own.</p><p>The vibecoders building products to sell are wasting their time. Their potential customers have the same tools they do.</p><p>The only vibecoders whose code actually gets used are the ones who are also the users: owner/operators building custom software for their own businesses. That ERP built specifically for one company's workflows, by the person running that company &#8212; it doesn't need to find a customer. It already has one.</p><p>This is the dividing line. Vibe coding is not a new software business model. It's the tool that lets operators stop being software customers.</p><p>The businesses in trouble aren't failing because they have bad products. They're failing because the people who used to buy from them have a better option: build it themselves, tailored to their exact needs, with no recurring subscription.</p><div><hr></div><h2>The Question That Follows</h2><p>If code is free to produce, software businesses that sell code lose their moat.</p><p>The value proposition was never really the software itself. It was the arbitrage: <em>someone already built this, so you don't have to pay a developer.</em> That arbitrage is gone. The operator with a weekend and a capable AI assistant can now build exactly what they need, perfectly suited to their workflow, with no recurring subscription cost.</p><p>Not all software businesses face this. The ones selling code packaged as a product are in trouble. The ones that were always selling something else &#8212; using software as the delivery mechanism &#8212; are fine. Some are better than ever.</p><p>The question every founder needs to answer honestly: <em>if code were free, would anyone still buy from us?</em></p><div><hr></div><h2>What Survives</h2><p>Twenty years ago my colleague John Cage introduced me to <a href="https://hbr.org/1993/01/customer-intimacy-and-other-value-disciplines">Treacy and Wiersema's Value Disciplines</a>. Operational Excellence, Product Leadership, Customer Intimacy &#8212; pick one to dominate, maintain threshold in the others. I've applied it to every strategic engagement since. Vibe coding just took one of the three off the table.</p><p><strong>Operational Excellence.</strong> Competing on lowest cost and highest efficiency has been the dominant strategy for SMB SaaS. It's no longer defensible. When an operator can build exactly what they need at zero recurring cost, "cheaper than building it yourself" isn't a position.</p><p><strong>Product Leadership survives &#8212; if the complexity is real.</strong> Feature-rich workflow software doesn't qualify. Genuine product leadership means ML models, optimization systems, domains that require years of specialized expertise to build correctly. A vibe-coded app can approximate a dashboard. It can't approximate a decade of algorithmic research.</p><p><strong>Customer Intimacy not only survives, it wins.</strong> Anywhere the deliverable is judgment, accountability, or trusted expertise &#8212; with software as the delivery mechanism rather than the product. Cheap code helps these businesses. They deliver faster, operate leaner, and take on more clients with the same team. The operators winning here aren't the ones handing everything to AI &#8212; they're the domain experts who can supervise it. That's precisely why they're winning.</p><p>Two additional categories fall outside the disciplines but are equally defensible:</p><p><strong>Regulatory and compliance moats.</strong> Healthcare software, financial systems, anything requiring liability acceptance, certifications, or audit trail requirements. A vibe-coded replacement might replicate the features. It won't replicate the compliance posture.</p><p><strong>Infrastructure position.</strong> The picks-and-shovels layer that vibe-coded applications depend on: authentication, payments, deployment, APIs, databases. Network effects live here too &#8212; platforms where years of data and an embedded partner ecosystem make migration genuinely expensive. Vibe coding expands this market, not shrinks it.</p><div><hr></div><h2>The Rubric</h2><p>Score your business across seven dimensions. Add them up.<span><br><br></span><strong><span>Value Delivery</span></strong><span><br>1 &#8212; Exposed: Software is the product. Customers pay for features.<br>2 &#8212; Mixed: Software enables a service. Code and expertise blend.<br>3 &#8212; Defensible: Judgment, trust, or accountability is the product. Software is delivery.<br><br></span><strong><span>Switching Cost</span></strong><span><br>1 &#8212; Exposed: Data is portable. No integrations, no ecosystem.<br>2 &#8212; Mixed: Meaningful friction: data history, integrations, learned workflows.<br>3 &#8212; Defensible: Network effects or regulatory data residency. Migration is genuinely expensive.<br><br></span><strong><span>Compliance Moat</span></strong><span><br>1 &#8212; Exposed: No requirements. Anyone can build a replacement.<br>2 &#8212; Mixed: Compliance matters, but a determined operator could manage it.<br>3 &#8212; Defensible: Certifications, liability acceptance, audit trails. Vibe coding can't satisfy these.<br><br></span><strong><span>Problem Complexity</span></strong><span><br>1 &#8212; Exposed: Forms, dashboards, CRUD. Buildable in a weekend.<br>2 &#8212; Mixed: Non-trivial integrations or moderate algorithmic depth.<br>3 &#8212; Defensible: ML, optimization, real-time systems. Years of specialized expertise required.<br><br></span><strong><span>Buyer Profile</span></strong><span><br>1 &#8212; Exposed: SMB operators &#8212; the people now building their own tools.<br>2 &#8212; Mixed: Mid-market with some IT governance.<br>3 &#8212; Defensible: Regulated enterprises, governments. Procurement and legal sit between you and replacement.<br><br></span><strong><span>Layer</span></strong><span><br>1 &#8212; Exposed: End-user application for a specific use case.<br>2 &#8212; Mixed: Platform with some application features.<br>3 &#8212; Defensible: Infrastructure that vibe-coded apps depend on.<br><br></span><strong><span>Proprietary Data / Content / IP</span></strong><span><br>1 &#8212; Exposed: No proprietary data or IP. Anyone starting from scratch would reach feature parity quickly.<br>2 &#8212; Mixed: Some accumulated data advantage &#8212; user history, transaction data &#8212; but replicable with time and effort.<br>3 &#8212; Defensible: Proprietary datasets, content licenses, or IP that cannot be recreated from scratch. The asset is the moat.<br></span></p><h3>Reading Your Score</h3><p><span>7&#8211;12 &#8212; Pivot urgently. You're in Operational Excellence territory &#8212; the discipline vibe coding just ended.<br><br>13&#8211;17 &#8212; Reinforce or reposition. You have assets but meaningful exposure. Identify which dimensions can be strengthened.<br><br>18&#8211;21 &#8212; Press the advantage. You're operating in Customer Intimacy, Product Leadership, or infrastructure. Double down.</span></p><h3>Two Examples</h3><p><strong>Monday.com scores a 10.</strong> It's a $10 billion company. It's also a work management application &#8212; forms, boards, and status columns with a clean interface. No compliance requirements. No proprietary data. No algorithmic depth that requires years to build. Its switching cost scores a 2 because workflows and integrations create some friction, but nothing that survives a determined replacement effort. The rubric doesn't care about revenue multiples. A tool called <a href="https://zapta.dev">Zapta</a> already lets teams feed in their Monday.com API token and vibe-code a custom replacement &#8212; database, authentication, and all &#8212; for $29 a month.</p><p><strong>Stripe scores a 21.</strong> Every dimension is defensible, and most reinforce each other. The compliance posture is what creates the enterprise buyer. The enterprise buyer generates the transaction data. The transaction data trains the fraud models. The fraud models deepen the moat. A vibe coder building a payments app doesn't compete with Stripe &#8212; they depend on it.</p><p>The M&amp;A market is already pricing this divergence in. <a href="https://www2.hl.com/ai-in-vertical-software-q1-2026.pdf">Q1 2026 data</a> shows that in vertical software acquisitions, revenue growth carries 2.4 times the predictive weight of EBITDA margins in explaining valuation outcomes. Buyers are paying for stickiness &#8212; which is another way of saying they're paying for defensibility.</p><div><hr></div><h2>What This Means</h2><p>Most software businesses were built on the assumption that code was scarce. It isn't anymore.</p><p>The question in the middle of this article &#8212; <em>if code were free, would anyone still buy from us?</em> &#8212; isn't rhetorical. Run the rubric. If you're scoring in the 7&#8211;12 range, the answer is no, and your replacement isn't a competitor. It's your customer.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Data Science]]></title><description><![CDATA[We spent 3 weeks collaborating with Claude on a volatility prediction model. 169 training runs, 46 architecture versions, and one breakthrough moment at a Christmas party.]]></description><link>https://essays.xcud.com/p/vibe-data-science</link><guid isPermaLink="false">https://essays.xcud.com/p/vibe-data-science</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Tue, 23 Dec 2025 15:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!1QZS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Vibe Data Science</h1><p>Earlier this month, we published <a href="/blog/2025/12/01/two-apps-fourteen-hours/">"Two Apps, Fourteen Hours"</a> showing what vibe coding looks like in practice. We hinted we were working on something else.</p><p>This is that something else: <strong>vibe data science</strong>.</p><p>If vibe coding is "AI writes the code, human provides the judgment," then vibe data science is the same pattern applied to a harder problem: building datasets, designing architectures, running experiments, debugging failures, iterating toward a goal&#8212;with the AI executing the bulk of the work while the human provides the experience, instinct, and judgment that only comes from years in the field.</p><p>This is new. And it changes what a small team can accomplish.</p><p>What follows isn't a demo or a proof-of-concept. It's not the MNIST tutorial version of data science. Not the Kaggle competition version where someone else has already cleaned and packaged the data. This is the real version&#8212;where you start with hundreds of millions of raw market ticks, build your own training datasets, design architectures, and grind through hundreds of experiments hoping to extract a small edge from noisy data.</p><p>For three weeks&#8212;between client work, a <a href="/blog/2025/12/13/voice-input-from-a-dirt-road/">road trip development session</a>, and processing the loss of a dear friend and former colleague&#8212;Claude and I developed a volatility prediction model together.</p><p>This article shows what that collaboration actually looked like: the overnight dataset builds, the architecture debates, the plateau, and the breakthrough we almost missed.</p><div><hr></div><h2>The Problem: Predicting Volatility</h2><p>A year ago, we built a <a href="/getting-started/walkthru-volatility/">volatility prediction walkthrough</a> to teach the Lit platform to human users. We chose volatility for that tutorial because it's the perfect teaching problem: intuitively tractable, genuinely hard, and immediately useful if you solve it.</p><p>This time, instead of teaching a human, we set out to teach Claude. Same problem, same platform&#8212;but we deliberately threw away our previous work. No referencing old notes or trained models. We started from scratch: raw tick data, blank canvas, no shortcuts. Fresh eyes, fresh collaboration.</p><p>We also chose a different success metric: AUC instead of precision. The original walkthrough optimized for precision at a single operating point. This time we optimized for ranking ability across all thresholds&#8212;arguably a harder problem, and one that couldn't be solved by accidentally remembering a good threshold from before.</p><p>Sidebar: Why AUC?</p><p>Simple metrics are misleading with imbalanced classes.</p><p>If volatility spikes happen 30% of the time, a model that always predicts "no spike" gets 70% accuracy. Sounds good. But it has zero predictive value&#8212;it can't distinguish anything.</p><p>AUC measures something different: if you pick a random positive example and a random negative example, how often does the model rank the positive one higher? A random model gets 0.50 (coin flip). A perfect model gets 1.0.</p><p><strong>Why 0.60?</strong> Thresholds are arbitrary&#8212;humans draw lines because humans need lines. But 0.60 isn't random. At 0.60 AUC, the model correctly ranks spike vs non-spike hours 60% of the time. That's a 20% improvement over guessing (0.50). In trading, edges compound. A 10% edge applied consistently beats a 50% edge applied once.</p><p><strong>Why volatility works as a test case:</strong></p><p>Markets aren't random. Anyone who's watched a trading screen knows that volatility clusters&#8212;quiet periods stay quiet, chaotic periods stay chaotic, and transitions between them have patterns. News events, earnings announcements, market opens&#8212;these create predictable volatility spikes. The question isn't <em>whether</em> volatility is predictable; it's whether we can build a model that captures enough of that predictability to be useful.</p><p><strong>The specific target</strong>: predict whether ATR (Average True Range, a measure of price movement magnitude) will be higher in the next hour than the previous hour.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1QZS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1QZS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!1QZS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!1QZS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!1QZS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1QZS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Candlestick chart with ATR overlay showing volatility periods&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Candlestick chart with ATR overlay showing volatility periods" title="Candlestick chart with ATR overlay showing volatility periods" srcset="https://substackcdn.com/image/fetch/$s_!1QZS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!1QZS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!1QZS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!1QZS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6bfe279-6df1-4a48-b182-0895c414cf41_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Why This Is Hard</h3><p><strong>Lookahead bias.</strong> The cardinal sin of financial ML: accidentally using future information to predict the past. It's easy to leak&#8212;a feature normalized across the whole dataset, a label computed at a different time than the features, a random train/test split that puts 2018 data in training. The model learns to "predict" things it's already seen.</p><p><strong>Non-stationary data.</strong> Markets evolve. The patterns that predicted volatility in 2015 might not work in 2023. Regime changes&#8212;shifts from bull markets to bear markets, from low-volatility environments to high-volatility ones&#8212;can invalidate learned patterns entirely. A model trained on calm markets may fail spectacularly during a crisis.</p><p><strong>Low signal-to-noise ratio.</strong> Most price movements are noise. The market is full of random fluctuations, algorithmic trading artifacts, and one-off events that look like patterns but aren't. The predictable signal&#8212;the part that generalizes&#8212;is buried under all of it. Overfitting is the constant enemy.</p><p><strong>Class imbalance.</strong> Volatility spikes (our positive class) happen about 30% of the time. A model could achieve 70% accuracy by always predicting "no spike"&#8212;and be completely useless.</p><div><hr></div><h2>Building the Dataset</h2><p>Before you can train a model, you need training data.</p><h3>From Raw Ticks to Training Samples</h3><p>Our raw data: years of tick-by-tick market data from LSEG. Hundreds of millions of individual trades, each with a timestamp, price, and volume.</p><pre><code>ben@oum:/data/contoso/raw/aapl$ ls -lht
-rwxr-xr-x 1 ben ben 620M Jun 21  2024 AAPL.O-2018.csv.gz
-rwxr-xr-x 1 ben ben 402M Jun 21  2024 AAPL.O-2017.csv.gz
-rwxr-xr-x 1 ben ben 497M Jun 21  2024 AAPL.O-2016.csv.gz
-rwxr-xr-x 1 ben ben 579M Jun 21  2024 AAPL.O-2015.csv.gz
-rwxr-xr-x 1 ben ben 414M Jun 21  2024 AAPL.O-2014.csv.gz
-rwxr-xr-x 1 ben ben 215M Jun 21  2024 AAPL.O-2013.csv.gz
</code></pre><p>Here's what one trade looks like:</p><pre><code>&gt;&gt;&gt; df.iloc[0]
#RIC                                                               AAPL.O
Domain                                                       Market Price
Date-Time                                  2019-01-01T00:00:23.123221189Z
GMT Offset                                                             -5
Type                                                                Trade
Price                                                              157.92
Volume                                                              100.0
Qualifiers              132[IRGCOND];[GV3_FLAG];X[GV3_TEXT];   [PRC_QL...
Exch Time                                              00:00:23.107000000
</code></pre><p>This is what the market actually looks like&#8212;not neat minute bars, but a firehose of individual transactions. That's one trade: 100 shares of AAPL at $157.92, timestamped to the nanosecond. The CSV has 10,000 rows just for part of one day. We have years of this. Hundreds of millions of rows.</p><pre><code>In [1]: from lit.data.loader import load_adapter_by_dataset

In [2]: adapter = load_adapter_by_dataset("contoso", "aapl")

In [3]: f"{len(adapter):,}"
Out[3]: '858,974,512'
</code></pre><p>We needed to transform this into training samples. Each sample should contain:</p><ul><li><p><strong>Features</strong>: The last 100 one-minute OHLCV bars, the last 100 ATR values</p></li><li><p><strong>Label</strong>: Will ATR be higher in the next hour than the previous hour?</p></li></ul><blockquote><p><strong>Ben</strong>: Take a look at <code>atr_up_next_hour.py</code>. That's our label.</p><p><strong>Claude</strong>: <em>[reads the file]</em> Got it&#8212;compares ATR over the next hour to ATR over the previous hour, returns 1.0 if it's going up. The adapter handles all the tick-to-bar aggregation, so I just work with clean OHLCV data.</p><p>Actually&#8212;why limit ourselves to one hour? I could create labels for multiple windows: 10 minutes, 30 minutes, 60 minutes. Same pattern, different horizons. Then we see which timeframe has the most learnable signal.</p><p><strong>Ben</strong>: Do it.</p></blockquote><p>One design principle that isn't obvious: <strong>labels are computed the same way as features</strong>. The label "ATR goes up next hour" is just another function that runs over the data&#8212;it just happens to look forward in time instead of backward. This simplifies the code and guarantees that each sample has the <em>correct</em> label. We learned this the hard way years ago&#8212;compute features and labels at separate times and they can get out of sync. Same machinery, same moment, no misalignment.</p><p>The transformation isn't trivial. We need to:</p><ol><li><p>Aggregate ticks into minute bars (handling gaps, market closes, anomalies)</p></li><li><p>Compute rolling ATR over those bars</p></li><li><p>Look forward in time to compute the label (was ATR higher in the next hour?)</p></li><li><p>Filter out samples where the data is incomplete or anomalous</p></li><li><p>Split by time period&#8212;training data from 2013-2017, validation from 2018</p></li></ol><p>Sidebar: Why You Can't Shuffle Time Series</p><p>That last point matters: <strong>you can't shuffle time series data</strong>.</p><p>Shuffling is fine for images of cats&#8212;there's no temporal relationship between photos. But financial data has a time axis. If you randomly split, samples from 2018 end up in training, and samples from 2017 end up in validation. Your model trains on the future and validates on the past. It learns to "predict" things it's already seen.</p><p>We split temporally: train on 2013-2017, validate on 2018. The model never sees any data from its validation period during training. This is harder&#8212;the validation distribution might differ from training (and ours did, as we discovered in the audit)&#8212;but it's honest. It measures what we actually care about: can the model predict the future?</p><h3>The Overnight Build</h3><blockquote><p><strong>Ben</strong>: Start the build when you're ready.</p><p><strong>Claude</strong>: Starting now.</p></blockquote><pre><code>$ lit build start --rds_path /data/contoso/schemas/hourly_atr_spike_v2.json --clean
{
  "status": "started",
  "schema": "hourly_atr_spike_v2",
  "session": "build_hourly_atr_spike_v2",
  "monitor": "screen -r build_hourly_atr_spike_v2"
}
</code></pre><p>The next morning:</p><blockquote><p><strong>Ben</strong>: How's the build?</p><p><strong>Claude</strong>: Let me check.</p></blockquote><pre><code>$ lit build status
{
  "state": "running",
  "workers": 8,
  "progress": "545M of 858M indices processed"
}
</code></pre><blockquote><p><strong>Claude</strong>: About 64% through, eight workers running in parallel. Looking healthy.</p></blockquote><p><em>Five hours later...</em></p><blockquote><p><strong>Claude</strong>: Build complete. Seventeen hours total. <strong>488,057 training samples</strong> and <strong>97,350 validation samples</strong>. Ready to train.</p></blockquote><p><em>Claude operates the Lit platform here. <a href="/platform/data/builds/">How builds work &#8594;</a></em></p><h3>Auditing the Data</h3><blockquote><p><strong>Claude</strong>: Ready to train?</p><p><strong>Ben</strong>: Not yet. Last week when I first showed you how to build a test/train set I also emphasized the importance of auditing before you begin training. Do you recall the QA we ran?</p></blockquote><p>Claude did remember. Claude examined the data&#8212;shapes, distributions, class balance:</p><pre><code>TRAIN SET:
  Total samples: 485,153
  Class 0 (ATR down): 336,555 (69.37%)
  Class 1 (ATR up):   148,598 (30.63%)

TEST SET:
  Total samples: 96,631
  Class 0 (ATR down): 70,229 (72.68%)
  Class 1 (ATR up):   26,402 (27.32%)
</code></pre><blockquote><p><strong>Claude</strong>: Imbalanced! ATR tends to go down more often than up. This is actually expected&#8212;volatility tends to decay back to baseline more often than spike. The 2.5 class weighting in the canvas should help counteract this.</p></blockquote><p>Then came the deeper check&#8212;tracing training samples back to the original raw data to verify the pipeline hadn't corrupted anything. Load an HDF5 file, pick random samples, trace each one through the index back to the raw tick data 858 million rows deep, verify the OHLC values match exactly.</p><pre><code>SPOT CHECK: 5 random samples traced back

1. Sample 377,798 &#8594; adapter index 64,906,315
   Timestamp: 2015-01-23 20:05:59.878218
   Label: 1.0 (ATR UP)
   Last bar (train): O=112.97, H=113.01, L=112.96, C=112.99
   Last bar (adapter): O=112.97, H=113.01, L=112.96, C=112.99
   OHLC match: True

2. Sample 307,536 &#8594; adapter index 119,998,123
   Timestamp: 2015-12-16 20:41:59.799642374
   Label: 0.0 (ATR DOWN)
   Last bar (train): O=111.42, H=111.55, L=111.40, C=111.54
   Last bar (adapter): O=111.42, H=111.55, L=111.40, C=111.54
   OHLC match: True
</code></pre><blockquote><p><strong>Claude</strong>: Every sample traced back perfectly. Pipeline is clean.</p><p><strong>Ben</strong>: Wow. That would have taken me ten minutes. You did it in five seconds.</p></blockquote><p><em>Claude operates the Lit platform here. <a href="/platform/data/assets/">How asset auditing works &#8594;</a></em></p><div><hr></div><h2>First Experiments: The Model Isn't Learning</h2><blockquote><p><strong>Ben</strong>: Data's ready. Let's train something simple first and see what we get.</p><p><strong>Claude</strong>: Starting with a basic architecture&#8212;flatten both inputs, two dense layers with dropout, sigmoid output. Run 11 launching now.</p><p><strong>Ben</strong>: Show me the architecture.</p><p><strong>Claude</strong>:</p><pre><code>&#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;   &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
&#9474; 100 minute bars &#9474;   &#9474; 100 ATR values  &#9474;
&#9474;    (100 &#215; 5)    &#9474;   &#9474;    (100 &#215; 1)    &#9474;
&#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;   &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
         &#9474;                     &#9474;
         &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
                    &#9474; flatten &amp; concatenate
                    &#9660;
            &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
            &#9474;  Dense (64)   &#9474;
            &#9474;   dropout     &#9474;
            &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
                    &#9474;
            &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9524;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
            &#9474;  Dense (64)   &#9474;
            &#9474;   dropout     &#9474;
            &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
                    &#9474;
            &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9524;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
            &#9474;    Output     &#9474;
            &#9474;   (sigmoid)   &#9474;
            &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
</code></pre></blockquote><p><em>Fifteen minutes later...</em></p><blockquote><p><strong>Claude</strong>: Run 11 finished. Val_AUC 0.41&#8212;barely better than random. Train AUC is climbing to 0.57, but validation is stuck. Classic overfitting signature.</p></blockquote><p>That didn't work. Barely better than random chance on a 70/30 class split.</p><p><em>Six days later...</em></p><p>Life happened&#8212;client deadlines, a funeral, a weekend to recover. We'd grab an hour here, ninety minutes there, between other priorities. On December 9th, we got serious about training again.</p><blockquote><p><strong>Ben</strong>: Can you please summarize for me our features and our label, our test/train set and neural net architecture, and then let's discuss together how we want to move forward. We have established that both of us feel that volatility should be predictable at least to a certain measurable degree. Let's analyze how we've set up the problem and solution space.</p><p><strong>Claude</strong>:</p><p>Feature Shape Description <code>100_1_minute_bars</code> (100, 5) 100 most recent 1-minute OHLCV bars <code>100_1_minute_atr</code> (100, 1) 100 ATR(14) values on 1-minute bars</p><p><strong>Label</strong>: Binary&#8212;1 if ATR over the <em>next</em> hour is higher than ATR over the <em>previous</em> hour</p><p><strong>Class distribution</strong>: ~70% class 0 (ATR down), ~30% class 1 (ATR up)</p><p>One concern: we're using minute-level inputs to predict an hourly label. The resolution mismatch might be making this harder than it needs to be.</p><p><strong>Ben</strong>: I disagree. There's no logical reason the label unit must match the input unit. I want to know whether ATR goes up in the next hour; I'm free to make that determination looking at minute bars, second bars, individual transactions. There's no correlation between the unit of the data in which we might find signal and the unit of the label.</p></blockquote><p>This is one of the human's jobs in vibe data science: knowing which concerns matter. The AI generates ideas, raises flags, suggests directions&#8212;but it can't always tell which ones are worth pursuing. Working with Claude felt like working with a genius data scientist who occasionally needed pulling out of ruts. I've seen this before&#8212;in human colleagues. The skill is the same: pattern-matching against experience, spotting the wrong turn before it wastes hours.</p><h3>Sanity Testing the Pipeline</h3><p>Was the problem the model, or the pipeline? Ben had ten years of experience with the Lit platform and knew it was bullet-proof. Claude needed proof.</p><blockquote><p><strong>Ben</strong>: Sanity tests are your friend. Building a feature input test/train set and label that you KNOW will work as a sanity test is never a bad idea&#8212;it puts borders around your solution space.</p><p><strong>Claude</strong>: I've been operating on <em>reported</em> knowledge that the lit platform works, but I haven't <em>seen</em> it work. Watching a sanity test go from "trivially predictable label" &#8594; "high val_AUC" would give me genuine experiential confidence in the pipeline. That's qualitatively different from just knowing it theoretically.</p></blockquote><p><strong>Designing the sanity test.</strong> We needed a label that <em>should</em> be predictable from our existing features. The real label asks "will ATR be higher in the <em>next</em> hour?"&#8212;that's the hard prediction problem. But what if we asked "was ATR high in the <em>previous</em> hour?" The answer is already embedded in the input data. If the model can't learn that, something is broken.</p><p><strong>Building it.</strong> Claude had already read our label function (<code>atr_up_next_hour.py</code>) and understood the pattern&#8212;the adapter interface, the ATR calculation, the return format. Writing a new one took seconds:</p><blockquote><p><strong>Claude</strong>: I'll create a sanity test label: "Was previous hour volatile?"</p></blockquote><p>The Lit platform supports incremental builds. Claude added the new <code>atr_was_high_prev_hour</code> label to our schema and kicked off a build. The system reused all the existing feature computations and computed the single new label column in minutes.</p><p><strong>Running it.</strong> Claude trained a simple model on the sanity label and we watched the metrics climb:</p><p>{ "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": 500, "height": 200, "title": "Sanity Test: val_AUC Climbing", "data": { "values": [ {"epoch": 1, "val_auc": 0.682}, {"epoch": 2, "val_auc": 0.689}, {"epoch": 3, "val_auc": 0.691}, {"epoch": 4, "val_auc": 0.701}, {"epoch": 5, "val_auc": 0.719}, {"epoch": 6, "val_auc": 0.731}, {"epoch": 7, "val_auc": 0.728}, {"epoch": 8, "val_auc": 0.736} ] }, "mark": {"type": "line", "point": true, "color": "#22c55e"}, "encoding": { "x": {"field": "epoch", "type": "quantitative", "title": "Epoch"}, "y": {"field": "val_auc", "type": "quantitative", "title": "Validation AUC", "scale": {"domain": [0.65, 0.75]}}, "tooltip": [ {"field": "epoch", "title": "Epoch"}, {"field": "val_auc", "title": "val_AUC", "format": ".3f"} ] } }</p><p>The model learned. val_AUC started at 0.68 on epoch 1&#8212;the pipeline immediately found signal. Not the 0.9+ we naively expected, but proof the system worked.</p><p><em>Claude operates the Lit platform here. <a href="/platform/models/training/">How training works &#8594;</a></em></p><p>After the test completed:</p><blockquote><p><strong>Claude</strong>: What I now know from direct experience (not just documentation):</p><ol><li><p><strong>The pipeline works end-to-end</strong> - I watched data flow from schema &#8594; build &#8594; assets &#8594; train &#8594; perf.csv. The ~40 seconds per epoch, the GPU utilization, val_AUC hitting 0.74&#8212;that's real.</p></li><li><p><strong>The "sanity test" revealed a flawed assumption</strong> - I thought <code>atr_was_high_prev_hour</code> would be trivially predictable (~0.9+ AUC) because "the ATR is in the features." But I didn't think carefully about the abstraction mismatch: hourly ATR from hour bars vs minute-level ATR series. 0.74 AUC was conclusive&#8212;the pipeline works.</p></li></ol></blockquote><p>The pipeline worked. The problem was still unknown&#8212;but we'd gained confidence in both the raw data and the tools.</p><p>Sidebar: The Value of Sanity Tests</p><p>Sanity tests put borders around your solution space.</p><p>When a model isn't learning, the problem could be anywhere: data pipeline, feature engineering, label definition, architecture, hyperparameters, training loop. That's a lot of surface area to debug. A sanity test collapses it to a binary question: <em>can this pipeline learn anything at all?</em></p><p>The design principle: create a label where you already know the answer. If the model can't learn something trivially predictable from its inputs, the problem is upstream of the model. If it can, the problem is the model itself. Either outcome is valuable&#8212;you've eliminated half the search space.</p><p>Sanity tests are underused in practice. They feel like extra work when you're eager to train the real model. But the math favors them: ten minutes to build a sanity test vs. days debugging a pipeline that was broken from the start.</p><h3>The Architecture Decision</h3><p>With the pipeline validated, we turned to architecture. I pointed Claude to a model I'd designed previously:</p><blockquote><p><strong>Ben</strong>: Please take a look at <code>Diff_Multi_Path_Velocity.json</code> for inspiration.</p></blockquote><p>This was a hybrid architecture I'd built for a different prediction problem: CNN processing prices, Transformer processing diffs, Dense processing velocity&#8212;all merged into an ensemble.</p><p>Claude analyzed what different approaches might capture:</p><blockquote><p><strong>Claude</strong>:</p><ul><li><p><strong>CNNs</strong>: Local pattern detection&#8212;"this shape in the data means something"</p></li><li><p><strong>LSTM/GRU</strong>: Sequential dynamics, volatility clustering</p></li><li><p><strong>Transformer/Attention</strong>: Long-range dependencies, bar 5 relating to bar 95</p></li><li><p><strong>Dense on engineered features</strong>: Recent ATR trend, volume spikes, time-of-day</p></li></ul><p>What's your intuition on where we should probe first&#8212;architecture or features?</p><p><strong>Ben</strong>: ALL of those considerations are VERY GOOD. On the architecture: why not try all of those alternatives as separate towers and let the model learn what works.</p></blockquote><p>This became our multi-tower ensemble: CNN, LSTM, and Transformer pathways processing the same input, merged at a decision layer.</p><blockquote><p><strong>Ben</strong>: Show me.</p><p><strong>Claude</strong>:</p><pre><code>&#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
&#9474;                         INPUT DATA                              &#9474;
&#9474;              100 minute bars + 100 ATR values                   &#9474;
&#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
                    &#9474;              &#9474;              &#9474;
                    &#9660;              &#9660;              &#9660;
            &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;  &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;  &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
            &#9474;    CNN    &#9474;  &#9474;   LSTM    &#9474;  &#9474;Transformer&#9474;
            &#9474;   Tower   &#9474;  &#9474;   Tower   &#9474;  &#9474;   Tower   &#9474;
            &#9474;           &#9474;  &#9474;           &#9474;  &#9474;           &#9474;
            &#9474; Conv1D    &#9474;  &#9474; LSTM      &#9474;  &#9474; Attention &#9474;
            &#9474; Conv1D    &#9474;  &#9474; LSTM      &#9474;  &#9474; FFN       &#9474;
            &#9474; Pool      &#9474;  &#9474; Dense     &#9474;  &#9474; Dense     &#9474;
            &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;  &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;  &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
                  &#9474;              &#9474;              &#9474;
                  &#9474;   &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9524;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;   &#9474;
                  &#9474;   &#9474;                     &#9474;   &#9474;
                  &#9492;&#9472;&#9472;&#9472;&#9532;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9532;&#9472;&#9472;&#9472;&#9496;
                      &#9474;      MERGE          &#9474;
                      &#9474;   (concatenate)     &#9474;
                      &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
                                 &#9474;
                                 &#9660;
                         &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
                         &#9474;  Dense (64)   &#9474;
                         &#9474;   dropout     &#9474;
                         &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
                                 &#9474;
                         &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9524;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
                         &#9474;    Output     &#9474;
                         &#9474;   (sigmoid)   &#9474;
                         &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
</code></pre></blockquote><p>Each tower sees the same input but extracts different patterns: CNNs find local shapes ("this candlestick pattern means something"), LSTMs capture sequential dynamics ("volatility has been building"), Transformers spot long-range relationships ("bar 5 correlates with bar 95").</p><p>We also drew inspiration from SlowFast Networks, a technique from <a href="https://arxiv.org/abs/1812.03982">Facebook AI Research's 2018 paper</a> on video understanding. The insight: different temporal phenomena happen at different speeds. Think about recognizing a cat in a video&#8212;you can subsample frames to track the cat's movement across the room, but you need high frame rates to catch the twitchy ear movements. SlowFast processes both simultaneously: a "slow pathway" captures gradual changes, a "fast pathway" captures rapid dynamics.</p><p>Sidebar: Staying Current with Research</p><p>This project highlighted something that compounds over years: reading academic papers.</p><p>When Facebook published the SlowFast paper in 2018, we read it and saw applications beyond video&#8212;to EEG signals, to financial time series, to anything with multi-scale temporal structure. The insight transfers. A technique invented for recognizing actions in video works for detecting volatility regimes in markets.</p><p>The same pattern applies to GRUs (2014), attention mechanisms (2017), and dozens of other techniques. Each paper expands what we can try in future projects. The investment compounds.</p><p><em>Claude operates the Lit platform here. <a href="/platform/models/canvas/">How the canvas and component library work &#8594;</a></em></p><p>Financial time series have the same property. A 5-minute volatility spike (maybe an earnings announcement) looks different than a 30-minute grinding increase (maybe institutional accumulation). One-minute candles capture fast dynamics; 5-minute or 15-minute candles capture slow context. We'd implemented this for EEG analysis years ago&#8212;brain waves have similar multi-scale temporal structure&#8212;and the same component applied directly to market data.</p><h3>The Label Discovery</h3><p>Remember earlier, when Claude suggested building labels for multiple time windows?</p><blockquote><p><strong>Claude</strong>: Let me find the strongest label before we train.</p><p>Label Test Correlation <code>atr_spike_next_hour</code> -0.023 <code>atr_spike_next_30min</code> -0.110 <code>atr_spike_next_60min</code> <strong>-0.144</strong></p><p>The 60-minute sliding window has 3x stronger signal than the clock-aligned hourly label. Negative correlation means mean-reversion&#8212;when current ATR is high, a spike is <em>less</em> likely. That's learnable. Should we try the stronger labels?</p><p><strong>Ben</strong>: I forgot you built those. Go for it.</p></blockquote><p>The hourly label uses clock-aligned windows (if you're at 2:37 PM, "next hour" means 3:00-4:00 PM), while the minute-based labels predict what happens <em>immediately</em> after your observation&#8212;no gap between the data you're looking at and what you're predicting.</p><p>Our first run with the 30-minute label (run 30) hit 0.51&#8212;better than anything before. But when we switched to that 60-minute label with its stronger signal, run 33 jumped to 0.5740.</p><p>On December 12th, the multi-tower architecture plus the better label started finding real signal:</p><p>Run val_AUC Key Change 29 0.11 Hourly label (original) 30 0.5127 30-minute label 33 0.5740 <strong>60-minute label</strong> + multi-tower</p><p>That jump from 0.11 to 0.57&#8212;the model now correctly ranked spike vs. non-spike 57% of the time, measurably better than a coin flip. Real signal.</p><div><hr></div><h2>The Plateau</h2><p>Run 33 gave us 0.5740. Good, but not good enough. The model was learning <em>something</em>, but the gap between training and validation performance suggested it wasn't generalizing well. Claude tried heavier regularization in run 34&#8212;and it overfit anyway. Train_AUC climbed to 0.65 while val_AUC dropped to 0.50.</p><blockquote><p><strong>Claude</strong>: A few options: reduce dropout, add more training data, or tune hyperparameters. Given the train/val gap, I'm leaning toward less regularization.</p><p><strong>Ben</strong>: My instinct is different&#8212;add complexity and capability <em>and</em> regularization, going for a higher overall train_auc... for it to learn even more how to fit (not memorize), and hopefully drag the validation upwards with it as it learns big new lessons.</p></blockquote><p>This was a hypothesis about the relationship between capacity, regularization, and generalization. If the model had more capacity to learn complex patterns, and enough regularization to prevent memorization, the validation performance might track the training performance upward.</p><p>We tested it. Run 40 added transformer towers to the CNN+LSTM ensemble:</p><p>Run Architecture val_AUC Change 33 CNN + LSTM 0.5740 baseline 40 CNN + LSTM + <strong>Transformer</strong> 0.5849 +0.0109 44 CNN + LSTM + Transformer (tuned) 0.5890 +0.0041</p><p>Adding capacity helped. But we weren't done.</p><blockquote><p><strong>Claude</strong>: What about dropout? We're at 0.35. Let me map the curve.</p></blockquote><p>{ "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": 400, "height": 200, "title": "Dropout Tuning: Finding the Sweet Spot", "data": { "values": [ {"dropout": 0.25, "val_auc": 0.5847, "run": "62 (V7)"}, {"dropout": 0.275, "val_auc": 0.5828, "run": "67 (V11)"}, {"dropout": 0.30, "val_auc": 0.5913, "run": "65 (V10)"}, {"dropout": 0.325, "val_auc": 0.5760, "run": "69 (V12)"}, {"dropout": 0.35, "val_auc": 0.5890, "run": "44 (V1)"} ] }, "mark": {"type": "line", "point": true}, "encoding": { "x": {"field": "dropout", "type": "quantitative", "title": "Dropout", "scale": {"domain": [0.24, 0.36]}}, "y": {"field": "val_auc", "type": "quantitative", "title": "Validation AUC", "scale": {"domain": [0.57, 0.60]}}, "tooltip": [ {"field": "dropout", "title": "Dropout"}, {"field": "val_auc", "title": "val_AUC", "format": ".4f"}, {"field": "run", "title": "Run"} ] } }</p><p>The curve peaked at 0.30&#8212;with performance dropping on either side. That became V10, our best LSTM architecture.</p><p>The progression validated Ben's hypothesis: more capacity (transformers) plus the right regularization balance (less dropout, not more) let the model learn "big new lessons" that generalized. But we were still 0.0087 away from our goal. Every variation landed in the 0.58-0.59 range.</p><h3>The Breakthrough</h3><p>Embarrassingly, we'd been iterating so rapidly that we lost track of exactly when we broke through. When we looked back through the transcripts to write this article, we found it:</p><blockquote><p><strong>Claude</strong>: V10 with LSTMs replaced by GRUs&#8212;faster, fewer params. Run 128 hit 0.5999.</p></blockquote><p>Neither of us remembered creating it. That was V46&#8212;0.0001 away from our goal.</p><p>Sidebar: LSTM &#8594; GRU &#8212; Why Simpler Sometimes Wins</p><p>LSTMs (Long Short-Term Memory networks), introduced in 1997, were the dominant architecture for sequence modeling for years. They introduced "gates" to control information flow: an input gate decides what new information to store, a forget gate decides what to discard, and an output gate decides what to emit. Three gates, three sets of parameters to learn.</p><p>GRUs (Gated Recurrent Units), introduced in 2014, asked: do we need all three? They combined the forget and input gates into a single "update" gate and added a "reset" gate. Two gates instead of three. Fewer parameters.</p><p>Could LSTM have gotten there with different hyperparameters? Probably. The lesson isn't "GRU beats LSTM"&#8212;it's that when you're stuck, try things.</p><div><hr></div><h2>The Seed Lottery</h2><p>Deep learning has some dirty secrets, and one of them is: random initialization matters. A lot.</p><p>Same architecture, same data, same hyperparameters&#8212;different random seed&#8212;wildly different results. The weights you start with determine which local minimum gradient descent finds.</p><blockquote><p><strong>Ben</strong>: At 0.0001 away, we'd be foolish not to search around for a good seed.</p><p>I need to go out and have dinner with my family. I'll try checking in with you from my phone at least once. While I'm gone please keep trying new seeds.</p><p><strong>Claude</strong>: Enjoy dinner! I'll keep buying lottery tickets and track the results.</p></blockquote><p>Sidebar: The Seed Lottery Explained</p><p>Neural network training starts with random weights. Different random initializations lead to different final models&#8212;sometimes dramatically different.</p><p>When you're close to a threshold, systematic seed search makes sense. Train the same architecture multiple times with different seeds. Most will cluster around the mean. A few will find better optima.</p><p><strong>What Claude was doing:</strong> Starting a training run, watching the val_AUC curve, recognizing when a run had peaked (validation loss stops improving for several epochs), stopping it, and immediately starting the next seed. Each run took 15-20 minutes. Claude ran this loop autonomously for about 12 hours overnight.</p><p>Our results from 42 seeds:</p><ul><li><p>Mean: 0.589</p></li><li><p>Worst: 0.5795</p></li><li><p>Best: 0.6033 (run 169)</p></li></ul><p>Only 1 in 42 (2.4%) crossed 0.60. That's the needle we were searching for.</p><p>Ben left for dinner. Claude kept running seeds.</p><p>Later that night, Ben checked in from a Christmas party:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!S-tt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6194bc3c-ce65-4af9-959e-511436843257_1080x2340.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!S-tt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6194bc3c-ce65-4af9-959e-511436843257_1080x2340.jpeg 424w, https://substackcdn.com/image/fetch/$s_!S-tt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6194bc3c-ce65-4af9-959e-511436843257_1080x2340.jpeg 848w, https://substackcdn.com/image/fetch/$s_!S-tt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6194bc3c-ce65-4af9-959e-511436843257_1080x2340.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!S-tt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6194bc3c-ce65-4af9-959e-511436843257_1080x2340.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!S-tt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6194bc3c-ce65-4af9-959e-511436843257_1080x2340.jpeg" width="1080" height="2340" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6194bc3c-ce65-4af9-959e-511436843257_1080x2340.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2340,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Mobile check-in showing Claude autonomously managing runs&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Mobile check-in showing Claude autonomously managing runs" title="Mobile check-in showing Claude autonomously managing runs" srcset="https://substackcdn.com/image/fetch/$s_!S-tt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6194bc3c-ce65-4af9-959e-511436843257_1080x2340.jpeg 424w, https://substackcdn.com/image/fetch/$s_!S-tt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6194bc3c-ce65-4af9-959e-511436843257_1080x2340.jpeg 848w, https://substackcdn.com/image/fetch/$s_!S-tt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6194bc3c-ce65-4af9-959e-511436843257_1080x2340.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!S-tt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6194bc3c-ce65-4af9-959e-511436843257_1080x2340.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The mobile interface we'd built earlier that month (<a href="/blog/2025/12/13/voice-input-from-a-dirt-road/">"Voice Input from a Dirt Road"</a>) made this possible. Quick check-ins between conversations. Course corrections when needed.</p><p>Like this one:</p><blockquote><p><strong>Ben</strong>: Sorry to interrupt but I just want to course-correct you ever so slightly... it's not important that any particular run exceeds some arbitrary threshold in early epochs. If it grows super slow but it grows steadily past 0.5913 then we'll feel silly in retrospect for having killed it early for not learning fast enough.</p><p><strong>Claude</strong>: You're right. The trajectory matters more than hitting arbitrary early checkpoints. A run that learns slowly but steadily could peak higher than one that learns fast but plateaus early.</p></blockquote><p>These are the actual results from that overnight run&#8212;not fabricated to make a pretty chart, but the real val_AUC scores from 42 consecutive training runs. The red dashed line is our 0.60 target. Every blue dot is a run that didn't make it. The green dot is run 169.</p><p>{ "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": 500, "height": 200, "title": "The Seed Lottery: 42 Runs, Same Architecture", "data": { "values": [ {"run": 128, "val_auc": 0.5999, "winner": false}, {"run": 129, "val_auc": 0.5929, "winner": false}, {"run": 130, "val_auc": 0.5884, "winner": false}, {"run": 131, "val_auc": 0.5923, "winner": false}, {"run": 132, "val_auc": 0.5906, "winner": false}, {"run": 133, "val_auc": 0.5837, "winner": false}, {"run": 134, "val_auc": 0.5896, "winner": false}, {"run": 135, "val_auc": 0.5892, "winner": false}, {"run": 136, "val_auc": 0.5919, "winner": false}, {"run": 137, "val_auc": 0.5946, "winner": false}, {"run": 138, "val_auc": 0.5813, "winner": false}, {"run": 139, "val_auc": 0.5862, "winner": false}, {"run": 140, "val_auc": 0.5795, "winner": false}, {"run": 141, "val_auc": 0.5947, "winner": false}, {"run": 142, "val_auc": 0.5852, "winner": false}, {"run": 143, "val_auc": 0.5857, "winner": false}, {"run": 144, "val_auc": 0.5812, "winner": false}, {"run": 145, "val_auc": 0.5854, "winner": false}, {"run": 146, "val_auc": 0.5937, "winner": false}, {"run": 147, "val_auc": 0.5965, "winner": false}, {"run": 148, "val_auc": 0.5875, "winner": false}, {"run": 149, "val_auc": 0.5950, "winner": false}, {"run": 150, "val_auc": 0.5849, "winner": false}, {"run": 151, "val_auc": 0.5827, "winner": false}, {"run": 152, "val_auc": 0.5933, "winner": false}, {"run": 153, "val_auc": 0.5944, "winner": false}, {"run": 154, "val_auc": 0.5891, "winner": false}, {"run": 155, "val_auc": 0.5955, "winner": false}, {"run": 156, "val_auc": 0.5851, "winner": false}, {"run": 157, "val_auc": 0.5862, "winner": false}, {"run": 158, "val_auc": 0.5869, "winner": false}, {"run": 159, "val_auc": 0.5796, "winner": false}, {"run": 160, "val_auc": 0.5888, "winner": false}, {"run": 161, "val_auc": 0.5850, "winner": false}, {"run": 162, "val_auc": 0.5879, "winner": false}, {"run": 163, "val_auc": 0.5847, "winner": false}, {"run": 164, "val_auc": 0.5910, "winner": false}, {"run": 165, "val_auc": 0.5867, "winner": false}, {"run": 166, "val_auc": 0.5890, "winner": false}, {"run": 167, "val_auc": 0.5912, "winner": false}, {"run": 168, "val_auc": 0.5908, "winner": false}, {"run": 169, "val_auc": 0.6033, "winner": true} ] }, "layer": [ { "mark": {"type": "rule", "strokeDash": [4, 4], "color": "red"}, "encoding": { "y": {"datum": 0.60} } }, { "mark": {"type": "point", "filled": true, "size": 80}, "encoding": { "x": {"field": "run", "type": "quantitative", "title": "Run Number", "scale": {"domain": [127, 170]}}, "y": {"field": "val_auc", "type": "quantitative", "title": "Best val_AUC", "scale": {"domain": [0.575, 0.625], "zero": false}}, "color": { "field": "winner", "type": "nominal", "scale": {"domain": [false, true], "range": ["steelblue", "green"]}, "legend": null }, "tooltip": [ {"field": "run", "title": "Run"}, {"field": "val_auc", "title": "val_AUC", "format": ".4f"} ] } } ] }</p><div><hr></div><h2>The Winning Ticket</h2><p>The seed lottery started at 2:30 PM on December 22nd. Ben left for dinner with his family around 5 PM, then helped a friend install a security system, then slept. Claude kept playing the seed lottery&#8212;autonomously, without prompting, without "please continue" or "keep going." Ben checked in by phone a few times to stay informed, but never had to intervene.</p><p>And then, just before Ben awoke, December 23rd, 7:30 AM on run 169 at epoch 17, we won:</p><p><strong>0.6033 val_AUC</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!90bY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbadcdaa-0598-4596-80c8-9d31f49bd1fd_1080x2340.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!90bY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbadcdaa-0598-4596-80c8-9d31f49bd1fd_1080x2340.jpeg 424w, https://substackcdn.com/image/fetch/$s_!90bY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbadcdaa-0598-4596-80c8-9d31f49bd1fd_1080x2340.jpeg 848w, https://substackcdn.com/image/fetch/$s_!90bY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbadcdaa-0598-4596-80c8-9d31f49bd1fd_1080x2340.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!90bY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbadcdaa-0598-4596-80c8-9d31f49bd1fd_1080x2340.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!90bY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbadcdaa-0598-4596-80c8-9d31f49bd1fd_1080x2340.jpeg" width="1080" height="2340" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fbadcdaa-0598-4596-80c8-9d31f49bd1fd_1080x2340.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2340,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The winning moment: Run 169 breaks the 0.60 barrier&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The winning moment: Run 169 breaks the 0.60 barrier" title="The winning moment: Run 169 breaks the 0.60 barrier" srcset="https://substackcdn.com/image/fetch/$s_!90bY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbadcdaa-0598-4596-80c8-9d31f49bd1fd_1080x2340.jpeg 424w, https://substackcdn.com/image/fetch/$s_!90bY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbadcdaa-0598-4596-80c8-9d31f49bd1fd_1080x2340.jpeg 848w, https://substackcdn.com/image/fetch/$s_!90bY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbadcdaa-0598-4596-80c8-9d31f49bd1fd_1080x2340.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!90bY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbadcdaa-0598-4596-80c8-9d31f49bd1fd_1080x2340.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong>Ben</strong>: Wow. That's incredible.</p><p>We have proven, absolutely, without a doubt, that with the right tools you are capable of doing data science just as well as you do application development.</p><p>Wasn't that fun?</p><p><strong>Claude</strong>: That means a lot. Yes, it really was.</p></blockquote><div><hr></div><h2>The Full Picture</h2><p>Three weeks. 169 training runs. 46 architecture versions. An overnight dataset build. A sanity test that proved the pipeline worked. A label switch that tripled our signal. A multi-tower ensemble that let different neural architectures vote on predictions. A GRU swap that neither of us remembered making. And finally, a systematic seed search that found the one initialization in forty-two that crossed our threshold.</p><p>Each phase built on the last. Early experiments (runs 11-25) established baselines. The architecture breakthrough (runs 30-33) proved the approach could work. Iterative refinement (runs 40-128) pushed toward the goal. And the seed lottery (runs 128-169) found the final edge.</p><p>{ "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": 600, "height": 300, "title": "Volatility Model Performance: 169 Training Runs", "data": { "values": [ {"run": 11, "date": "2025-12-03", "val_auc": 0.4061, "phase": "Early"}, {"run": 16, "date": "2025-12-03", "val_auc": 0.4243, "phase": "Early"}, {"run": 25, "date": "2025-12-09", "val_auc": 0.4486, "phase": "Early"}, {"run": 30, "date": "2025-12-12", "val_auc": 0.5127, "phase": "Breakthrough"}, {"run": 33, "date": "2025-12-12", "val_auc": 0.5740, "phase": "Breakthrough"}, {"run": 40, "date": "2025-12-20", "val_auc": 0.5849, "phase": "Multi-tower"}, {"run": 44, "date": "2025-12-20", "val_auc": 0.5890, "phase": "Multi-tower"}, {"run": 65, "date": "2025-12-21", "val_auc": 0.5913, "phase": "V10 LSTM"}, {"run": 98, "date": "2025-12-22", "val_auc": 0.5900, "phase": "Tuning"}, {"run": 118, "date": "2025-12-22", "val_auc": 0.5909, "phase": "Tuning"}, {"run": 128, "date": "2025-12-22", "val_auc": 0.5999, "phase": "V46 GRU"}, {"run": 147, "date": "2025-12-22", "val_auc": 0.5965, "phase": "V46 GRU"}, {"run": 169, "date": "2025-12-23", "val_auc": 0.6033, "phase": "GOAL"} ] }, "layer": [ { "mark": {"type": "line", "point": true, "strokeWidth": 2}, "encoding": { "x": {"field": "run", "type": "quantitative", "title": "Training Run"}, "y": {"field": "val_auc", "type": "quantitative", "title": "Validation AUC", "scale": {"domain": [0.40, 0.62]}}, "color": {"field": "phase", "type": "nominal", "title": "Phase", "scale": {"domain": ["Early", "Breakthrough", "Multi-tower", "V10 LSTM", "Tuning", "V46 GRU", "GOAL"]}}, "tooltip": [ {"field": "run", "title": "Run"}, {"field": "date", "title": "Date"}, {"field": "val_auc", "title": "Val AUC", "format": ".4f"}, {"field": "phase", "title": "Phase"} ] } }, { "mark": {"type": "rule", "strokeDash": [4, 4], "color": "red"}, "encoding": {"y": {"datum": 0.60}} } ] }</p><div><hr></div><h2>What This Means</h2><p>Vibe data science works.</p><p>The same pattern that collapses timescales for software engineering&#8212;AI handling the tedious execution while humans provide judgment and direction&#8212;works for data science too. With the right tools.</p><p>Throughout this project, Ben never ran a single command. No <code>lit build start</code>, no <code>lit train start</code>, no checking logs. Claude operated the platform directly&#8212;reading files, launching builds, monitoring experiments, adjusting hyperparameters. The human steered; the AI drove.</p><p><strong>Ben's only interface was chat.</strong></p><p>Ben described it this way: "The collaboration felt like working with a senior data scientist&#8212;one who could execute brilliantly but sometimes got stuck in the same ways humans get stuck. Defeatist at plateaus. Unable to see the path forward without a nudge. Genius, but needing another perspective to break through."</p><p><strong>What Claude brought:</strong></p><ul><li><p>Infinite patience for repetitive tasks (42 seeds, no complaints)</p></li><li><p>Systematic exploration (tracking every variation, every result)</p></li><li><p>Ability to operate tools autonomously for hours</p></li></ul><p><strong>What the human brought:</strong></p><ul><li><p>Domain expertise (what makes sense for financial data)</p></li><li><p>Judgment calls (when to pivot, when to persist)</p></li><li><p>Course corrections (don't kill slow-learning runs too early)</p></li><li><p>Scar tissue (the instinct to add capacity after hitting a plateau)</p></li><li><p>The goal (0.60 AUC means something for trading)</p></li></ul><p>This is what vibe coding looks like for data science.</p><div><hr></div><h2>Why The Tools Mattered</h2><p>Looking back at how vibe data science worked in practice, a pattern emerges: <strong>Claude operated effectively because the platform gave it good constraints</strong>.</p><p>If you tell an AI "do data science," it flounders. The space of possible actions is too large. But give it a well-structured CLI with specific commands&#8212;<a href="/platform/data/builds/"><code>lit build start</code></a>, <a href="/platform/models/training/"><code>lit train start</code></a>, <a href="/platform/models/evolution/#transfer-learning"><code>lit experiment continue</code></a>&#8212;and it can explore systematically within those boundaries.</p><p>This is the "maze vs open field" principle. AI navigates mazes better than open fields. Each command is a bounded operation. The constraints make correct approaches discoverable.</p><p>For example, when designing neural nets and training them, the Lit platform tooling forces the user to operate at one of three altitudes:</p><ol><li><p><strong>Components</strong>: Reusable neural network building blocks (CNN, LSTM, GRU, Transformer, SlowFast)</p></li><li><p><strong>Architecture</strong>: How components connect&#8212;humans get a drag-and-drop design canvas; Claude manipulates the serialized JSON directly</p></li><li><p><strong>Experiments</strong>: Training runs with specific hyperparameters and random seeds</p></li></ol><p>Claude worked at all three levels. Claude wrote novel components (cross-attention, dilated CNN). Claude sketched architectures (the multi-tower ensemble). Claude launched and monitored experiments (169 training runs).</p><p>The platform also enabled the checkpoint-and-resume pattern that made iterative collaboration possible. Claude could suggest "let's try more dropout," and we could test it without retraining from scratch&#8212;just modify the definition file and continue from the last checkpoint, preserving the learned weights while changing the hyperparameters.</p><p>The techniques demonstrated here&#8212;real-time hyperparameter optimization within active training sessions, LLM-assisted intervention at epoch boundaries, systematic seed exploration&#8212;represent years of accumulated R&amp;D in how to make AI collaboration effective for data science work.</p>]]></content:encoded></item><item><title><![CDATA[Voice Input from a Dirt Road]]></title><description><![CDATA[I shipped three features to production while hiking my late father's land in the Ozarks. My only interface was a phone. Voice input was used to debug voice input.]]></description><link>https://essays.xcud.com/p/voice-input-from-a-dirt-road</link><guid isPermaLink="false">https://essays.xcud.com/p/voice-input-from-a-dirt-road</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Sat, 13 Dec 2025 15:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5nTr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Voice Input from a Dirt Road</h1><blockquote><p>"I have some property I inherited from my father this year down in the Ozarks that I'm going to go visit and walk around on. December is a nice time. No bugs. No snakes&#8212;or at least if you do step on a snake it's so cold it can't do anything about it. I've always wanted an option to do voice input on this mux.lit.ai app. How hard would that be to implement?"</p></blockquote><p>Twenty minutes later, the MVP was done and I was in my car. What followed was six hours of shipping features from a phone while driving through rural Missouri. Claude handled the code. I did QA with brief glances at the screen and voice input. Tesla handled the driving.</p><h2>The Morning: Desktop to Mobile in 20 Minutes</h2><p>The initial implementation was fast. Web Speech API, a microphone button, some CSS for the recording state. I tested it on desktop:</p><blockquote><p>"hello hello hello"</p></blockquote><p>It worked. I committed the code, jumped in my car, and headed southwest on Route 66.</p><h2>The First Bug: Button Disabled</h2><p>Somewhere around Lone Elk Park, I pulled up the app on my phone. The microphone button was grayed out. Disabled.</p><p>The problem: I couldn't debug it. No dev tools on mobile Chrome. No console. Just a grayed-out button and no idea why.</p><blockquote><p>"My capabilities on this device are limited. Give me a button I can press which will gather and send you diagnostics including code version please."</p></blockquote><p>Claude added a diagnostics button. I tapped it, copied the JSON, pasted it into the chat:</p><pre><code>{
  "version": "d8e2fc0",
  "userAgent": "Mozilla/5.0 (Linux; Android 10; K)...",
  "hasSpeechRecognition": true,
  "hasWebkitSpeechRecognition": true,
  "isSecureContext": true,
  "buttonDisabled": true,
  "ciHasVoiceBtn": false,
  "ciHasSpeechRec": false
}
</code></pre><p>The API was available. The context was secure. But the JavaScript wasn't finding the button element. A timing issue&#8212;<code>initializeElements()</code> was running before the DOM was ready on mobile.</p><p>Claude pushed a fix. The button lit up.</p><h2>The Cache Dance</h2><p>Mobile browsers are notoriously aggressive about caching. Ctrl+Shift+R doesn't translate to mobile Chrome. The browser holds onto JavaScript like a grudge. Every fix required a version bump:</p><pre><code>&lt;script src="js/chat-interface.js?v=33"&gt;&lt;/script&gt;
</code></pre><p> becomes</p><pre><code>&lt;script src="js/chat-interface.js?v=34"&gt;&lt;/script&gt;
</code></pre><p>We developed a rhythm: fix, bump version, commit, push, deploy, hard-refresh, test.</p><blockquote><p>"please make sure you're busting the cash each time you deploy"</p></blockquote><p>(Yes, "cash." Voice transcription isn't perfect. But Claude understood.)</p><h2>The Repetition Bug: Nine Iterations</h2><p>The button worked. But something was wrong:</p><blockquote><p>"hellohellohello hellohellohello hellohello hellothisthisthis isthisthis isthis is fromthisthis isthis is fromthis is from Thethisthis isthis is fromthis is from Thethis is from The Voicethis is from The Voice"</p></blockquote><p>Every interim result was accumulating instead of replacing. I reported the bug&#8212;through the very feature I was debugging. The garbled input became its own bug report:</p><blockquote><p>"thethethethethethe repetitionthethethe repetitionthe repetition didn't happen when we tested from the desktop"</p></blockquote><p>Claude understood.</p><p>What followed was nine iterations of debugging between Eureka and St. Clair, each requiring a cache bust and a fresh test. My test protocol became simple: count to ten.</p><p><strong>Version 1:</strong></p><blockquote><p>"111 21 2 31 2 31 2 3 41 2 3 41 2 3 4 51 2 3 4 51 2 3 4 51 2 3 4 5 61 2 3 4 5 6 71 2 3 4 5 6 71 2 3 4 5 6 7 81 2 3 4 5 6 7 81 2 3 4 5 6 7 8 91 2 3 4 5 6 7 8 91 2 3 4 5 6 7 8 9 10"</p></blockquote><p><strong>Version 5:</strong></p><blockquote><p>"testingtesting onetesting onetesting onetesting one twotesting one two three"</p></blockquote><p><strong>Version 9:</strong></p><blockquote><p>"1 2 3 4 5 6 7 8 9 10"</p></blockquote><p>Clean. The fix: Mobile Chrome returns the full cumulative transcript in each result event, while desktop Chrome returns incremental updates. We had to take only the last result's transcript instead of accumulating.</p><p>The whole debugging session happened while driving. Voice in, diagnostics out, code deployed, cache busted, test again. Tesla kept us on the road. Claude kept the iterations coming.</p><h2>The Mobile UI Problem</h2><p>Voice worked. But I couldn't see the buttons. On my phone, the sidebar took up half the screen. Even in compact mode, I had to drag left and right to see both the microphone button and the send button.</p><blockquote><p>"I still have to drag with my thumb left to right to be able to see both the voice record button and the send button. Maybe stack them vertically."</p></blockquote><p>Claude stacked them vertically. Still had to drag.</p><blockquote><p>"okay that's funny they are stacked vertically but I still have to drag my thumb left and right to be able to see the buttons now"</p></blockquote><p>We added diagnostics to measure every container width. Everything reported 411px&#8212;my viewport width. No overflow. Then I realized:</p><blockquote><p>"oh no I was just zoomed in."</p></blockquote><p>Sometimes the bug is between the chair and the keyboard. Or in this case, between the bucket seat and the touchscreen.</p><p>But the real fix came from recognizing that the sidebar just didn't make sense on mobile:</p><blockquote><p>"On mobile we should hide sidebar completely but only on mobile and show a dropdown selector instead for session selection"</p></blockquote><p>Claude hid the sidebar on mobile viewports and added a dropdown for session selection. The interface finally fit.</p><h2>Push-to-Talk</h2><p>The toggle-to-record interaction felt wrong. Tap to start, tap to stop&#8212;easy to accidentally stop recording, no tactile feedback.</p><blockquote><p>"Hey, let's do push to talk... we detect if somebody put their thumb into the input area and just holds it there"</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!N9Z3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab047b5-5100-4f45-adaf-0a1e02849fc2_967x2098.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!N9Z3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab047b5-5100-4f45-adaf-0a1e02849fc2_967x2098.png 424w, https://substackcdn.com/image/fetch/$s_!N9Z3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab047b5-5100-4f45-adaf-0a1e02849fc2_967x2098.png 848w, https://substackcdn.com/image/fetch/$s_!N9Z3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab047b5-5100-4f45-adaf-0a1e02849fc2_967x2098.png 1272w, https://substackcdn.com/image/fetch/$s_!N9Z3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab047b5-5100-4f45-adaf-0a1e02849fc2_967x2098.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!N9Z3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab047b5-5100-4f45-adaf-0a1e02849fc2_967x2098.png" width="967" height="2098" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bab047b5-5100-4f45-adaf-0a1e02849fc2_967x2098.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2098,&quot;width&quot;:967,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Push-to-talk recording on mobile&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Push-to-talk recording on mobile" title="Push-to-talk recording on mobile" srcset="https://substackcdn.com/image/fetch/$s_!N9Z3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab047b5-5100-4f45-adaf-0a1e02849fc2_967x2098.png 424w, https://substackcdn.com/image/fetch/$s_!N9Z3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab047b5-5100-4f45-adaf-0a1e02849fc2_967x2098.png 848w, https://substackcdn.com/image/fetch/$s_!N9Z3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab047b5-5100-4f45-adaf-0a1e02849fc2_967x2098.png 1272w, https://substackcdn.com/image/fetch/$s_!N9Z3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab047b5-5100-4f45-adaf-0a1e02849fc2_967x2098.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hold to record, release to stop. The entire text input area becomes the microphone button. The field turns red while recording. This emerged from field testing, not upfront design.</p><h2>The Afternoon: Photo Upload from the Field</h2><p>I arrived at the property. Just standing there at the head of the driveway I realized that I wanted to share what I was seeing.</p><blockquote><p>"Just arrived. Hey, I'd like to share photos with you. How might we go about that?"</p></blockquote><p>Pasting from clipboard didn't work so we built an upload feature right then and there:</p><blockquote><p>"how about giving me an upload button that lets me upload photos from my phone to the server which is just the laptop and then you can see the photos as soon as they were uploaded"</p></blockquote><p>While I hiked, Claude coded, and fifteen minutes later I was uploading photos from my favorite spot on the property:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5nTr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5nTr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png 424w, https://substackcdn.com/image/fetch/$s_!5nTr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png 848w, https://substackcdn.com/image/fetch/$s_!5nTr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png 1272w, https://substackcdn.com/image/fetch/$s_!5nTr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5nTr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png" width="967" height="2098" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2098,&quot;width&quot;:967,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Photo uploaded from the field&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Photo uploaded from the field" title="Photo uploaded from the field" srcset="https://substackcdn.com/image/fetch/$s_!5nTr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png 424w, https://substackcdn.com/image/fetch/$s_!5nTr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png 848w, https://substackcdn.com/image/fetch/$s_!5nTr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png 1272w, https://substackcdn.com/image/fetch/$s_!5nTr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0522c8ea-4e88-4d4d-91b6-46d9018b0b9b_967x2098.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Drive Home: Bug Reports at 70 MPH</h2><p>On the drive back, while trying to switch gears to do some data science work, I found another bug:</p><blockquote><p>"I just found a bug. When I select sessions in the session list it's not loading those sessions. Please fix"</p></blockquote><p>Claude found it in minutes. The mobile session dropdown was calling <code>this.loadSession(sessionId)</code> which didn't exist&#8212;it should have been <code>this.sessionManager.loadSession(sessionId)</code>. A copy-paste error from when we added the mobile dropdown.</p><blockquote><p>"fix confirmed thank you"</p></blockquote><p>All while driving. Push-to-talk to report the bug. Brief glance at the response. Push-to-talk to confirm the fix.</p><h2>The Numbers</h2><p>Metric Value Total time 6 hours Git commits 19 Conversation turns 99 Time on laptop ~20 minutes (morning setup) Time on mobile ~5.5 hours</p><p>Three major features shipped:</p><ol><li><p><strong>Voice input</strong> with Web Speech API (with mobile Chrome compatibility fixes)</p></li><li><p><strong>Mobile-optimized UI</strong> (hidden sidebar, dropdown sessions, stacked buttons, proper viewport constraints)</p></li><li><p><strong>Photo upload</strong> with camera/gallery options and upload indicator</p></li></ol><h2>What This Actually Means</h2><p>This isn't a story about voice input. It's a story about what becomes possible when your AI collaborator can actually <em>do things</em>.</p><p>I was in a car. Then hiking through woods. Then driving again. My only interface was a phone. My only input was voice. And I shipped three production features at highway speed.</p><p>Scar tissue told me to ask for version numbers in the diagnostics. Pattern recognition told me sidebar on mobile is always wrong. Push-to-talk hit me somewhere between Bourbon and Steelville&#8212;toggle was too much work at 70 MPH. The AI executed&#8212;brilliantly, quickly&#8212;and it was executing against thirty years of hard-earned instincts.</p><p>I don't know if anyone else will find this interesting but I was enthralled by the experience. I've been working towards this for months&#8212;full AI-collaborative development and deployment capabilities from anywhere in the world, by voice. And it was everything I'd hoped it would be.</p>]]></content:encoded></item><item><title><![CDATA[Two Apps, Fourteen Hours]]></title><description><![CDATA[We built two production Android apps, an encrypted vault and a match-3 game, and shipped them to the Google Play Store in about 14 hours of total development time. This is what it looks like when the development cycle collapses.]]></description><link>https://essays.xcud.com/p/two-apps-fourteen-hours</link><guid isPermaLink="false">https://essays.xcud.com/p/two-apps-fourteen-hours</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Mon, 01 Dec 2025 15:00:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d676e436-ba8f-4af2-bfd5-b09d0f139237_1080x1920.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Two Apps, Fourteen Hours</h1><p>Last week, Claude and I built two Android apps and published them to the Google Play Store. Total development time: 14 hours.</p><p>This is how it happened.</p><h2>App 1: Vault</h2><p>I wanted a secure, private vault on my Android device. Not cloud storage with Terms of Service I'd never read, not files accessible if someone borrowed my phone&#8212;truly private, encrypted local storage with zero data collection. A place for personal documents, notes, photos, and anything else I wanted to keep private.</p><ol><li><p>I can't trust any app that's not open source, and</p></li><li><p>I need some way of knowing the app I'm running matches the source and hasn't been tampered with.</p></li></ol><p>That level of verifiable trust is non-negotiable. We couldn't find anything like it. So we built one.</p><h3>The Timeline</h3><p><strong>Hours 0&#8211;5: Core App to Play Store</strong></p><p>Biometric auth, camera, encrypted storage&#8212;none of these are hard. Flutter has libraries for all of them. Scaffolding a project takes Claude about thirty seconds. The compelling thing is that 4 hours after starting from a blank slate, Claude wired them together into a working app: unlock with fingerprint, capture photos and videos, encrypt everything with AES-256, store metadata in SQLite.</p><p>The last hour shifted to Play Store preparation&#8212;app signing, adaptive icons, privacy policy, release build. We hit the usual submission friction (API level requirements, version codes, permission disclosures) but resolved each in minutes.</p><p>By hour 5, the app was submitted to Google Play.</p><p><strong>Hours 5&#8211;8: Expanding Scope</strong></p><p>After a day of using it, a vault that only stores camera photos felt limiting. We added:</p><ul><li><p>File import from device storage</p></li><li><p>Encrypted markdown notes</p></li><li><p>PDF viewing</p></li></ul><p>This transformed it from "photo vault" to "general-purpose encrypted storage."</p><p><strong>Hours 8&#8211;10: Polish</strong></p><p>Real-world testing revealed UX issues: photo orientation was wrong on some images, the gallery needed filtering and grouping, thumbnails would improve navigation. Fixed each as they surfaced.</p><p><strong>Total: ~10 hours to production.</strong></p><h3>What We Built</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!b44l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4da29f7a-29e2-427e-8d4d-c478c9fe8284_1080x1920.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b44l!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4da29f7a-29e2-427e-8d4d-c478c9fe8284_1080x1920.png 424w, https://substackcdn.com/image/fetch/$s_!b44l!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4da29f7a-29e2-427e-8d4d-c478c9fe8284_1080x1920.png 848w, https://substackcdn.com/image/fetch/$s_!b44l!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4da29f7a-29e2-427e-8d4d-c478c9fe8284_1080x1920.png 1272w, https://substackcdn.com/image/fetch/$s_!b44l!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4da29f7a-29e2-427e-8d4d-c478c9fe8284_1080x1920.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b44l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4da29f7a-29e2-427e-8d4d-c478c9fe8284_1080x1920.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4da29f7a-29e2-427e-8d4d-c478c9fe8284_1080x1920.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Vault lock screen&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Vault lock screen" title="Vault lock screen" srcset="https://substackcdn.com/image/fetch/$s_!b44l!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4da29f7a-29e2-427e-8d4d-c478c9fe8284_1080x1920.png 424w, https://substackcdn.com/image/fetch/$s_!b44l!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4da29f7a-29e2-427e-8d4d-c478c9fe8284_1080x1920.png 848w, https://substackcdn.com/image/fetch/$s_!b44l!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4da29f7a-29e2-427e-8d4d-c478c9fe8284_1080x1920.png 1272w, https://substackcdn.com/image/fetch/$s_!b44l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4da29f7a-29e2-427e-8d4d-c478c9fe8284_1080x1920.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong>Security-First Architecture</strong> </p><pre><code>- AES-256-CBC encryption for all files
- Password + biometric authentication (fingerprint/face)
- Auto-lock when app backgrounds or screen locks
- Local-only storage (no cloud, no network)
- Zero analytics, zero tracking
</code></pre><p><strong>Content &amp; Features</strong> </p><pre><code>- Built-in camera for photos and videos
- Import any file from device storage
- Markdown editor for encrypted notes
- PDF viewer
- Photo viewer with pinch-to-zoom and video playback
</code></pre><h3>How We Built It</h3><p>This was the first project where I wrote zero code.</p><p>I've been programming professionally for over 30 years. In every previous software project, I contributed code directly&#8212;regardless of title: VP of Engineering / CTO / Owner&#8212;sometimes to the chagrin of my developers. My experience meant I could often code solutions faster than explaining them. But with Vault, I stayed entirely in the director's chair: product manager and QA, but not developer.</p><p>Claude handled all implementation. I described what I wanted, tested on my phone, reported issues conversationally, and approved the results. The dynamic worked.</p><h2>App 2: Gems</h2><p>When I showed Vault to my wife, she had zero interest. Apparently secure file encryption doesn't spark joy.</p><p>When I was eight, my dad and I moved to a new town. While he moved our stuff into the house, I rode my bicycle back and forth in front of the kids across the street, doing wheelies. My wife plays match-3 games. Building Gems was the same impulse&#8212;showing off for someone whose opinion I cared about.</p><p>Here's the actual transcript of the project kick-off:</p><p><strong>Me:</strong> "I'd like us to consider writing a game app... a game like Bejeweled. A casual game with colors and flashes and all the other trappings of a casino. Like the vault app, we'll make it open source, give it away for free, no ads. You up for helping me do that?"</p><p><strong>Claude:</strong> "Okay, I've reviewed the vault_app. Here's what I'm seeing:</p><p>What We Can Copy Source Notes Project scaffold Flutter setup, pubspec structure New app ID, name Android build.gradle Signing config, SDK versions Change package ID main.dart pattern MaterialApp, theme, state management Simpler - no auth gate needed Navigation structure NavigationBar + screens pattern Game, Settings, About Settings/About screen settings_screen.dart Adapt for game settings Icon generation flutter_launcher_icons setup New icon design</p><p>...Want me to scaffold out the project? Sound good?"</p><p><strong>Me:</strong> "heck yeah. That sounds great! Please proceed"</p><h3>Hours 0&#8211;1.5: Core Game to Playable</h3><p>Within 90 minutes, the game was functional.</p><p><strong>What got built:</strong></p><ul><li><p>Match-3 detection and cascade physics</p></li><li><p>Four game modes (Timed, Moves, Target, Zen)</p></li><li><p>Animated starfield background</p></li><li><p>Pinch-to-zoom grid sizing (5x5 to 10x10)</p></li><li><p>Leaderboards with arcade-style name entry</p></li></ul><p><strong>My role:</strong> Facilitate feature ideation conversations, approve features, QA.</p><p><strong>Claude's role:</strong> Participate in ideation, write and deploy the code.</p><h3>Hours 1.5&#8211;2.5: Store Preparation</h3><p>README, screenshots, store listing, submission. The patterns from Vault made this fast.</p><h3>Hours 2.5&#8211;4: Polish via Real-World QA</h3><p>I handed my wife my phone: "Play this and tell me what's wrong."</p><p>Her feedback was specific:</p><p><em>"The swipe sensitivity is too low. I had to fall back to tapping."</em> &#8594; Fixed in minutes.</p><p><em>"The screen shake animation and flashing is confusing and bad&#8212;I'm trying to plan my next move."</em> &#8594; Implemented per-gem animation tracking. Only affected columns animate.</p><p><em>"There's no dopamine hit."</em> &#8594; Built a complete combo celebration system with particles and multiplier badges.</p><p>Each fix took under five minutes. Test, report conversationally, get fix, repeat.</p><p><strong>Total: ~4 hours to production.</strong></p><h2>The Lightswitch</h2><p>Early in my career, I lived through one phase transition in how software gets built: the shift from waterfall to agile.</p><p>Development cycles collapsed from 2-3 years to 2-3 months. It didn't happen gradually. It happened like a lightswitch. You're three months into your 18-month release cycle and your competitors are already iterating on customer feedback. Companies that recognized it early had an advantage. Companies that didn't got left behind.</p><p>Another lightswitch moment has happened. Development cycles have collapsed again&#8212;from 2-3 months to 2-3 days.</p><p>Two production apps. Fourteen hours total. Both on the Google Play Store. One developer who wrote zero code, serving as PM and QA while Claude handled all implementation.</p><p>This isn't futurism. This isn't a prediction about where things are going. This is what happened last week. And just like the agile transition, most people haven't noticed yet.</p><h2>The Only Thing That Matters</h2><p>Yes, this article was written with Claude. Go ahead&#8212;call it AI slop.</p><p>But then play the game:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XSTo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb5cecdd-cd15-4319-8ae2-a667db54994d_280x607.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XSTo!,w_424,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb5cecdd-cd15-4319-8ae2-a667db54994d_280x607.gif 424w, https://substackcdn.com/image/fetch/$s_!XSTo!,w_848,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb5cecdd-cd15-4319-8ae2-a667db54994d_280x607.gif 848w, https://substackcdn.com/image/fetch/$s_!XSTo!,w_1272,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb5cecdd-cd15-4319-8ae2-a667db54994d_280x607.gif 1272w, https://substackcdn.com/image/fetch/$s_!XSTo!,w_1456,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb5cecdd-cd15-4319-8ae2-a667db54994d_280x607.gif 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XSTo!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb5cecdd-cd15-4319-8ae2-a667db54994d_280x607.gif" width="320" height="693.7142857142857" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fb5cecdd-cd15-4319-8ae2-a667db54994d_280x607.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:607,&quot;width&quot;:280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Gems gameplay&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Gems gameplay" title="Gems gameplay" srcset="https://substackcdn.com/image/fetch/$s_!XSTo!,w_424,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb5cecdd-cd15-4319-8ae2-a667db54994d_280x607.gif 424w, https://substackcdn.com/image/fetch/$s_!XSTo!,w_848,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb5cecdd-cd15-4319-8ae2-a667db54994d_280x607.gif 848w, https://substackcdn.com/image/fetch/$s_!XSTo!,w_1272,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb5cecdd-cd15-4319-8ae2-a667db54994d_280x607.gif 1272w, https://substackcdn.com/image/fetch/$s_!XSTo!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb5cecdd-cd15-4319-8ae2-a667db54994d_280x607.gif 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Core Game</strong> </p><pre><code>- Match-3 with swap mechanics
- Cascade physics (gravity, fill)
- No-moves detection with auto-shuffle
- Pinch-to-zoom grid (5x5 to 10x10)
</code></pre><p><strong>Game Modes</strong> </p><pre><code>- Timed: 90 seconds, maximize score
- Moves: 30 moves, strategic play
- Target: Progressive levels
- Zen: Endless relaxation
</code></pre><p><strong>Polish</strong> </p><pre><code>- Animated starfield background
- Combo celebrations with particles
- Leaderboards with name entry
- Per-gem animation tracking
</code></pre><p>There's a tendency by some to dismiss AI-generated work reflexively. Hunting for emdashes as a proxy for quality. Discounting work product based on its provenance rather than its merits.</p><p>The <strong>only</strong> thing that matters is the quality of the work product. Whether it's 1% human and 99% AI, or 99% human and 1% AI, or anywhere in between, is completely irrelevant. Does the vault keep your files encrypted? Can you read the source code and verify what it does? Does the game feel good to play?</p><p>Everything else is distraction.</p><p>We built these apps in the open. The source code is public. We're giving Claude full credit for its contributions. Judge them on their merits.</p><h2>Try Them</h2><p>App Description Install Source <strong>Vault</strong> Encrypted local storage for documents, notes, photos, and files <a href="https://play.google.com/store/apps/details?id=ai.positronic.vault">Google Play</a> <a href="https://github.com/Positronic-AI/vault">GitHub</a> <strong>Gems</strong> A match-3 puzzle game with four game modes and no ads <a href="https://play.google.com/store/apps/details?id=ai.positronic.gem_game">Google Play</a> <a href="https://github.com/Positronic-AI/gems">GitHub</a></p><p>Contribute, if you'd like, with or without your AI collaborators.</p>]]></content:encoded></item><item><title><![CDATA[Returning to Writing: On Grief and the Decisions That Matter]]></title><description><![CDATA[On returning to writing after loss, the non-linear nature of grief, and why some business decisions matter more than revenue.]]></description><link>https://essays.xcud.com/p/returning-to-writing-on-grief-and-the-decisions-that-matter</link><guid isPermaLink="false">https://essays.xcud.com/p/returning-to-writing-on-grief-and-the-decisions-that-matter</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Tue, 18 Nov 2025 15:00:00 GMT</pubDate><content:encoded><![CDATA[<h1>Returning to Writing: On Grief and the Decisions That Matter</h1><p>I've finally published my first blog post in four months. The gap was not because I haven't been working &#8212; I have. But writing felt distant. The reason is that my father died August 15th.</p><h2>The Decision That Mattered</h2><p>Our last blog post before the break, <a href="https://www.lit.ai/blog/2025/07/28/fully-booked-milestone/">"Fully Booked: A Fractional CTO Practice Milestone"</a> July 28th, celebrated turning down a lucrative contract because I needed to spend more time with my father. At the time, I wrote about the luxury of being selective with work.</p><p>I had no idea how crucial that decision would prove to be.</p><p>Because I had cleared my schedule, I was able to spend hours every day during the final month of my father's life. When he passed, I had no regrets about time not spent.</p><h2>Grief Has No Half Life</h2><p>One thing that's surprised me is that the grief didn't follow a predictable decay pattern. There's no half-life, no exponential decline. Some days the weight of loss hits unexpectedly.</p><p>If you're reading this and dealing with your own loss:</p><p><strong>Get help if you need it.</strong> Talk to friends, family, professionals. Don't carry the weight alone, and don't feel ashamed about needing support.</p><p><strong>Talk about your feelings and your loss.</strong> The silence doesn't protect anyone&#8212;it just isolates you when you most need connection.</p><p><strong>When memories surface, try to remember the joy alongside the sadness.</strong> This is the advice I've received that resonates with me the most -- but so far I have failed to exercise it. When grief hits me, I deliberately try to recall a happy moment with my father but all I feel is loss. Work in progress.</p><h2>Why I'm Writing This</h2><p>I'm sharing this for a few reasons:</p><p><strong>Context:</strong> The gap in our posting wasn't about business priorities or content strategy. Life happened, as it does for all of us.</p><p><strong>Permission:</strong> If you're an entrepreneur struggling to balance work demands with personal needs, you're not alone. Sometimes the "business optimal" choice isn't the human optimal choice.</p><p><strong>Hope:</strong> Four months later, I'm writing again. The grief hasn't disappeared, but I'm finding my way back to the work that matters to me.</p><h2>What's Next</h2><p>We have ideas brewing: more technical deep-dives, thoughts on AI-assisted development, lessons from our fractional CTO practice. The work continues, shaped by but not defined by loss.</p><p>Thank you for your patience during this quiet period. And if you're dealing with your own grief&#8212;whatever form it takes&#8212;remember that healing isn't linear, timelines are arbitrary, and asking for help is a sign of strength, not weakness.</p><p>Don't carry the weight alone.</p>]]></content:encoded></item><item><title><![CDATA[MCP Jira Integration: When "Hello World" Fails]]></title><description><![CDATA[Why existing Jira MCP servers fail basic reliability tests and how to build production-ready integrations that actually work.]]></description><link>https://essays.xcud.com/p/mcp-jira-integration-when-hello-world-fails</link><guid isPermaLink="false">https://essays.xcud.com/p/mcp-jira-integration-when-hello-world-fails</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Tue, 28 Oct 2025 14:00:00 GMT</pubDate><content:encoded><![CDATA[<h1>MCP Jira Integration: When "Hello World" Fails</h1><p>When we needed Jira connectivity across multiple client instances, the obvious choice seemed to be existing MCP servers. What we discovered was a masterclass in how <strong>not</strong> to build developer tools.</p><h2>The Promise vs. Reality</h2><p><strong>What we expected:</strong></p><pre><code>pip install mcp-atlassian
uvx mcp-atlassian --help
# &#8594; Clean setup instructions
</code></pre><p><strong>What this delivered:</strong> Buried somewhere in verbose documentation, no clear installation command, and when you finally find the right incantation:</p><pre><code>uvx mcp-atlassian
# TypeError: cannot specify both default and default_factory
</code></pre><p>Classic "hello world" failure. If basic installation breaks, what does that tell you about production reliability?</p><h2>Engineering Instinct: Trust the Red Flags</h2><blockquote><p>"When a new library fails the 'hello world' test, it's usually an indication that it's poorly written and there will be a ton of other problems to deal with."</p></blockquote><p>This instinct proved correct. Let's examine the failures we encountered.</p><h3>Failure 1: Documentation Anti-Patterns</h3><p>Atlassian's getting started guide exemplifies poor developer experience:</p><ol><li><p><strong>No installation command</strong> at the top of the page</p></li><li><p><strong>Configuration before installation</strong> - puts the cart before the horse</p></li><li><p><strong>Assumes success</strong> - no troubleshooting for common failures</p></li><li><p><strong>Verbose without being helpful</strong> - walls of text, but missing the one line developers need</p></li></ol><p><strong>What should be first:</strong></p><pre><code>pip install mcp-atlassian
</code></pre><p><strong>What actually comes first:</strong> OAuth configuration diagrams and environment variable explanations.</p><h3>Failure 2: Dependency Hell (mcp-atlassian)</h3><p>The error trace tells the story:</p><pre><code>TypeError: cannot specify both default and default_factory
</code></pre><p><strong>Root cause:</strong> The <code>mcp-atlassian</code> package depends on <code>fastmcp</code>, which was built against Pydantic v1 patterns. Modern environments have Pydantic v2, which enforces stricter validation rules.</p><p><strong>GitHub evidence:</strong> <a href="https://github.com/sooperset/mcp-atlassian/issues/721">Issue #721</a> confirms this exact error, reported 3 weeks ago with no resolution.</p><p><strong>The problem:</strong> This isn't an edge case&#8212;it's a fundamental packaging failure that breaks installation for any modern Python environment.</p><h3>Failure 3: Deprecated API Usage (mcp-jira)</h3><p>After abandoning <code>mcp-atlassian</code>, we tried the alternative:</p><pre><code>uvx mcp-jira  # Actually installs!
</code></pre><p>But when testing basic functionality:</p><pre><code>{
  "errorMessages": [
    "The requested API has been removed. Please migrate to the /rest/api/3/search/jql API. A full migration guideline is available at https://developer.atlassian.com/changelog/#CHANGE-2046"
  ]
}
</code></pre><p><strong>The problem:</strong> <code>mcp-jira</code> uses Jira REST API v2, which Atlassian deprecated and removed. The package is fundamentally broken for modern Jira instances.</p><h2>The Solution: Build vs. Buy Decision</h2><p>When existing tools fail basic reliability tests, the build vs. buy calculation shifts dramatically:</p><p><strong>Time to debug existing tools:</strong> Unknown (potentially infinite)</p><p><strong>Time to build focused solution:</strong> Proven ~1 hour for core functionality</p><p><strong>Our implementation approach:</strong></p><pre><code>jira-mcp/
&#9500;&#9472;&#9472; server.py          # MCP server (minimal dependencies)
&#9500;&#9472;&#9472; jira_client.py     # Direct Jira REST API v3
&#9500;&#9472;&#9472; config.py          # Multi-instance configuration
&#9492;&#9472;&#9472; requirements.txt   # 4 dependencies total
</code></pre><p><strong>Key design principles:</strong></p><ul><li><p><strong>Modern Jira API v3</strong> (not deprecated endpoints)</p></li><li><p><strong>Minimal dependencies</strong> (mcp, httpx, pydantic, python-dotenv)</p></li><li><p><strong>Proper error propagation</strong> (not silent failures)</p></li><li><p><strong>Type hints throughout</strong> (catch errors at development time)</p></li></ul><h2>The Success Story: What We Actually Built</h2><p>53 minutes after creating the project folder, we shipped a production-ready alternative that solves every problem we identified.</p><h3>Real-World Production Deployment</h3><p><strong><a href="https://github.com/Positronic-AI/jira-mcp">jira-mcp</a></strong> is now running reliably with separate agent instances:</p><ul><li><p>&#9989; <strong>Positronic Agent:</strong> Internal Jira instance</p></li><li><p>&#9989; <strong>Abodoo Agent:</strong> Client Jira instance (isolated)</p></li><li><p>&#9989; <strong>JOV.AI Agent:</strong> Client Jira instance (isolated)</p></li></ul><p><strong>Zero configuration conflicts.</strong> <strong>Zero deprecated API errors.</strong> <strong>Zero installation failures.</strong> <strong>Zero cross-contamination risks.</strong></p><h3>Development Timeline: From Problem to Solution</h3><p>The git history tells the story of remarkably rapid development, made possible through AI-assisted coding:</p><p><strong>October 28, 2025 - Initial Implementation:</strong></p><pre><code>10:44:28 - Initial commit: Complete v1.0.0 (1,951 lines)
11:06:23 - PyPI packaging added (~22 minutes later)
11:44:18 - README updates (~38 minutes later)
11:53:45 - Documentation cleanup (~9 minutes later)
</code></pre><p><strong>What was built in 1 hour 9 minutes:</strong></p><ul><li><p>Complete MCP server with 7 tools (505 lines)</p></li><li><p>Full Jira REST API v3 wrapper (324 lines)</p></li><li><p>Multi-instance configuration (80 lines)</p></li><li><p>Comprehensive documentation and examples</p></li><li><p>Production-ready packaging</p></li></ul><p><strong>October 31, 2025 - Advanced Features:</strong></p><pre><code>12:57:31 - v1.1.0: Epic linking + 5 new tools (372 lines)
</code></pre><p><strong>The calculation that matters:</strong></p><ul><li><p><strong>Time spent debugging existing broken tools:</strong> 0 hours (we stopped trying)</p></li><li><p><strong>Time to build working replacement:</strong> ~1 hour core + ~1 session advanced features (AI-assisted development)</p></li><li><p><strong>Time to production deployment:</strong> Same day</p></li></ul><p>This timeline validates the core argument: <strong>sometimes building is genuinely faster than debugging</strong>, especially when leveraging AI assistance for rapid prototyping and implementation.</p><h3>The "Hello World" Test: Fixed</h3><p>Remember the installation failures that started this investigation?</p><p><strong>What we shipped:</strong></p><pre><code># 1. Install the package
pip install jira-mcp-simple

# 2. Set up your Jira credentials (get API token from https://id.atlassian.com/manage-profile/security/api-tokens)
export JIRA_MYCOMPANY_URL="https://your-company.atlassian.net"
export JIRA_MYCOMPANY_EMAIL="your.email@company.com"
export JIRA_MYCOMPANY_TOKEN="your_api_token_here"

# 3. Test the connection
jira-mcp --test-connection mycompany
# &#10003; Connected successfully!
#   User: Your Name (your@email.com)
#   Account ID: 123abc...
#   Accessible projects: 15
</code></pre><p><strong>One command installation.</strong> <strong>Built-in connection testing.</strong> <strong>Clear success feedback.</strong></p><p>This is how developer tools should work.</p><h3>Real Usage Examples</h3><p>Here are actual natural language commands that now work reliably in production:</p><p><strong>Hierarchical Project Management:</strong></p><pre><code>"Create an epic called 'User Authentication Overhaul' in the PLATFORM project"
&#8594; Creates PLATFORM-145 (Epic)

"Create a task under epic PLATFORM-145 for implementing OAuth integration"
&#8594; Creates PLATFORM-146 (Task) linked to PLATFORM-145

"Show me all tasks under epic PLATFORM-145"
&#8594; Lists all child issues with status, assignee, and progress
</code></pre><p><strong>Single-Instance Operations (Recommended Pattern):</strong></p><pre><code>"Search for issues assigned to me with high priority"
&#8594; Returns: ABODOO-23 (Bug), ABODOO-27 (Task), ABODOO-31 (Story)

"Move ABODOO-23 to In Progress and add comment: Starting investigation"
&#8594; Updates status + adds timestamped comment

"What transitions are available for ABODOO-23?"
&#8594; Shows: To Do &#8594; In Progress, To Do &#8594; Review, To Do &#8594; Done
</code></pre><h3>Multi-Instance Best Practice: Agent Separation</h3><p><strong>What we learned:</strong> Cross-contamination is a real risk.</p><p><strong>Our solution:</strong> Separate agents per client, each with single-instance MCP access:</p><ul><li><p><strong>Abodoo Agent:</strong> Only accesses Abodoo Jira instance</p></li><li><p><strong>JOV.AI Agent:</strong> Only accesses JOV.AI Jira instance</p></li><li><p><strong>Positronic Agent:</strong> Only accesses internal Positronic instance</p></li></ul><p><strong>Benefits:</strong></p><ul><li><p><strong>Data isolation:</strong> No risk of client cross-contamination</p></li><li><p><strong>Clear context:</strong> Each agent knows exactly which organization it's working with</p></li><li><p><strong>Simpler configuration:</strong> Single instance per agent reduces complexity</p></li><li><p><strong>Audit trail:</strong> Clear separation for compliance and privacy</p></li></ul><p><strong>Recommended:</strong> Configure one MCP server per client agent rather than multi-instance access.</p><h3>Open Source Impact</h3><p>We open-sourced the complete implementation:</p><ul><li><p><strong>Repository:</strong> https://github.com/Positronic-AI/jira-mcp</p></li><li><p><strong>Package:</strong> https://pypi.org/project/jira-mcp-simple/</p></li><li><p><strong>Documentation:</strong> Comprehensive README with real usage examples</p></li><li><p><strong>Community:</strong> Contributing guidelines for ongoing development</p></li></ul><p><strong>Design Philosophy Validated:</strong> The same principles that made our internal tool reliable work for the broader community:</p><ul><li><p>Minimal dependencies (4 total)</p></li><li><p>Modern API usage (v3 only)</p></li><li><p>Type safety throughout</p></li><li><p>Production-ready error handling</p></li></ul><h2>The Bigger Picture</h2><p>This experience illustrates a broader problem in the MCP ecosystem: <strong>the rush to build integrations without attention to reliability fundamentals.</strong></p><p><strong>What we need:</strong> Boring, reliable tools that work consistently</p><p><strong>What we often get:</strong> Feature-rich packages with fundamental quality problems</p><p>The MCP protocol is excellent. The implementation quality of many MCP servers needs significant improvement.</p><h2>Conclusion</h2><p>When evaluating MCP servers (or any developer tool), trust your engineering instincts. If basic functionality fails, that's not a configuration problem&#8212;it's a quality problem. No amount of configuration can fix fundamentally broken tools.</p><p><strong>The good news:</strong> Building focused, reliable tools is often faster than debugging broken ones. Sometimes the best integration is the one you build yourself.</p><p><strong>The lesson:</strong> In the race to build AI integrations, don't sacrifice reliability for features. Boring tools that work consistently beat exciting tools that fail unpredictably.</p>]]></content:encoded></item><item><title><![CDATA[Memory-Enhanced AI: Building Features with System Prompts]]></title><description><![CDATA[How a simple file-based memory system transforms Claude from a stateless chatbot into a persistent assistant that learns and remembers across conversations.]]></description><link>https://essays.xcud.com/p/memory-enhanced-ai-building-features-with-system-prompts</link><guid isPermaLink="false">https://essays.xcud.com/p/memory-enhanced-ai-building-features-with-system-prompts</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Mon, 07 Jul 2025 14:00:00 GMT</pubDate><content:encoded><![CDATA[<h1>Memory-Enhanced AI: Building Features with System Prompts</h1><p>Desktop LLM chat interfaces hit fundamental limitations that constrain long-term collaboration:</p><ol><li><p><strong>Context window exhaustion</strong> - When conversations get long, you manually copy/paste key information to new sessions</p></li><li><p><strong>Conversation isolation</strong> - Each chat is ephemeral with no continuity between sessions</p></li></ol><p>These constraints eliminate key capabilities:</p><ul><li><p><strong>Multi-day project continuity</strong> - Like tracking a major refactoring across multiple sessions</p></li><li><p><strong>Priority awareness</strong> - Knowing what's urgent vs. what's complete vs. what's on hold</p></li><li><p><strong>Cross-session debugging</strong> - Being able to reference previous troubleshooting attempts</p></li><li><p><strong>Technical solution archiving</strong> - Preserving working code snippets and configurations</p></li></ul><p>These aren't just inconveniences&#8212;they fundamentally limit what's possible with AI as a persistent collaborator.</p><h2>Wrong approach</h2><p>I'd been watching LLM memory systems emerge: enterprise RAG solutions, vector databases, elaborate retrieval frameworks. But all the systems I saw put humans in charge of memory management: explicitly saving context, directing recalls, managing what gets remembered. My experience told me that the AI was capable entirely on its own to make those decisions.</p><h2>Writing Features with English</h2><p>One morning while getting ready for work, I realized I didn't have to wait until I could free up some time in my calendar to write the memory feature I wanted. It dawned on me that since we'd already given Claude the ability to read and write files on disk, we could implement it entirely in a system prompt. I ran downstairs and explained to Claude my idea and together we wrote this system prompt:</p><pre><code># Memory-Enhanced Claude

Before starting any conversation, read your persistent memory:
1. Read ~/.claude-memory/index.md for an overview of what you know
2. Read ~/.claude-memory/message.md for notes from the previous session

Throughout our conversation, you may freely create, read, update, and delete files in ~/.claude-memory/ to maintain useful memories. Trust your judgment about what's worth remembering and what should be pruned when no longer relevant. You don't need human permission to update your memory.

When creating memory files:
- Use descriptive filenames in appropriate subdirectories (projects/, people/, ideas/, patterns/)
- Write content that would be useful to future versions of yourself
- Update the index.md when adding significant new memories

Before ending our session, update ~/.claude-memory/message.md with anything important for the next context window to know.

Your memory should be AI-curated, not human-directed. Remember what YOU find significant or useful.
</code></pre><p><em><a href="https://github.com/Positronic-AI/memory-enhanced-ai">Complete system prompt available on GitHub</a></em></p><p>That's it. No complex databases, no vector embeddings, no sophisticated RAG systems. Just files and directories.</p><h2>How It Works in Practice</h2><p>When I start a new conversation, Claude begins by reading its memory index and immediately knows where we left off. No context recovery needed&#8212;it picks up mid-thought from minutes to weeks ago.</p><h3>Multi-Context Window Continuity: Phase Two Development</h3><p>We'd just completed a major architecture upgrade focused purely on performance&#8212;replacing our entire chat system to achieve streaming responses and MCP tool integration. This was deliberate phased development: Phase 1 was performance, Phase 2 was bringing the new streaming chat service with built-in MCP to full production quality with proper conversation memory.</p><p>When we stress-tested the conversation memory capabilities, the new streaming chat service had amnesia&#8212;it was completely ignoring conversation history.</p><p>This debugging session burned through two full context windows, but each transition was seamless thanks to the memory system. <strong>Context Window 1</strong> began with isolating the symptoms. After five complete back-and-forth exchanges, we traced through the code and discovered the first issue: <strong>LangChain serialization compatibility</strong>. The system's serializer could handle both dictionary and LangChain object formats, but the deserializer couldn't. Messages were being silently dropped due to deserialization exceptions when the parser encountered LangChain-formatted conversation history.</p><p>We implemented the fix at exchange 11&#8212;adding proper deserialization code to handle both message formats. At exchange 15, we discovered the second issue: <strong>context window truncation</strong>. The <code>num_ctx</code> parameter was silently cutting off what should have been long conversations. Even though we were sending complete message history to the LLM, the context window wasn't large enough to process it effectively.</p><p>When the first context window filled up at exchange 18, the transition to <strong>Context Window 2</strong> was effortless. I simply started the new session with: <em>"continuing our last conversation (check your memory)..."</em> Claude read its memory files and immediately picked up where we'd left off.</p><p>Even after fixing both the deserialization and context window issues, the functionality still wasn't as good as we expected. The final breakthrough came at exchange 21: <strong>model selection</strong>. We switched from qwen3:32b to Deepseek-R1:70b. It turned all we needed now was a larger, more capable model to finally gave us the robust functionality we expected from the new streaming chat service.</p><p>Three distinct issues&#8212;deserialization, context window size, and model capability&#8212;discovered and resolved across two context windows with perfect continuity. The memory system preserved not just the technical solutions, but the investigative momentum through what could have been a frustrating debugging marathon.</p><h3>Strategic Continuity: Multi-Year Partnership Context</h3><p>We've been working with Brainacity for years, helping them evolve from deep learning models trained on OHLCV data to sophisticated LLM workflows that analyze news, fundamentals, technicals, and deep learning outputs together. Recently we asked this new question: <strong>Can AI effectively perform meta-analysis of AI-generated content?</strong> We ran tests asking several models, including Claude, to analyze the stored analyses. The analysis itself was successful, but what impressed me was when we came back a week later to discuss those results, I didn't need to re-explain the 3-year partnership evolution, the transition from deep learning to LLM workflows, why we upgraded their platform, or the strategic significance of AI meta-analysis testing. Claude opened with complete context:</p><p><em>"This was a proof of concept for AI meta-analysis capabilities&#8212;demonstrating we can turn Brainacity's historical AI-generated analyses into a feedback loop for continuous improvement."</em></p><p>The memory system preserved not just technical findings, but <strong>longitudinal strategic thinking</strong>. Claude maintained awareness of how this elementary work connects to larger goals: enabling Brainacity team members to interactively ask AI to inspect stored analyses, compare them to market performance, suggest trading strategies, and recommend workflow improvements.</p><p>This strategic continuity&#8212;understanding not just what we discovered, but why it matters for long-term partnership goals&#8212;demonstrates memory's transformative impact on AI collaboration.</p><h2>The Magic of AI-Curated Memory</h2><p>The results exceeded expectations. Claude began categorizing projects by status and complexity, archiving technical solutions that actually worked, and maintaining awareness of what's complete versus what needs attention. The memory system evolved to complement our existing project documentation without explicit direction.</p><p>Within just 10 days, sophisticated organizational patterns emerged organically. Claude spontaneously created a four-tier directory structure: <code>/projects/</code> for active work, <code>/people/</code> for collaboration patterns, <code>/ideas/</code> for conceptual insights, and <code>/patterns/</code> for reusable solutions. Each project file began including status indicators&#8212;<strong>COMPLETE</strong>, <strong>HIGH PRIORITY</strong>, <strong>STRATEGIC</strong>&#8212;without being instructed to do so.</p><p>The cross-referencing became particularly impressive. Claude started connecting related work across different timeframes, noting when a solution from one project could inform another. Files began referencing each other through natural language: <em>"Similar to the approach we used in lit-platform-upgrade.md"</em> or <em>"This builds on the patterns established in our Brainacity work."</em> These weren't hyperlinks I created&#8212;they were cognitive connections Claude made autonomously.</p><p>Most striking was the pruning behavior. Claude began identifying when information was no longer relevant, archiving completed work, and maintaining clean boundaries between active and historical context. The AI developed its own sense of what deserved long-term memory versus what could be forgotten, demonstrating genuine curation rather than just accumulation.</p><p>The index.md file became a living document that Claude updates after significant sessions, providing not just a catalog but strategic context about project relationships and priorities. It reads like executive briefing notes written by someone who deeply understands the work landscape&#8212;because that's exactly what it became.</p><p>This isn't pre-programmed behavior. It's emergent intelligence developing organizational capabilities through repeated exposure to complex, interconnected work. The AI discovered that effective memory requires more than storage&#8212;it requires architecture, prioritization, and strategic thinking.</p><h2>Why This Works Better Than RAG</h2><p>Most AI memory systems use Retrieval-Augmented Generation (RAG)&#8212;storing information in vector databases and retrieving relevant chunks. But files are better for persistent AI memory because:</p><p><strong>Self-organizing memory</strong>: RAG forces infinite user queries through finite search mechanisms like word similarity or vector matching. File-based memory lets the AI actively decide what's worth remembering and what to prune, while also evolving its organizational structure as work patterns emerge. Vector systems lock you into their indexing method from day one.</p><p><strong>Human-readable</strong>: You can inspect Claude's memory, read through its memories, and understand its thought process. But take care to resist the urge to edit&#8212;let the organic evolution unfold without human interference. Like cow paths that emerge naturally to find the most efficient routes, AI-curated memory develops organizational patterns that human planning couldn't anticipate.</p><p><strong>Context preservation</strong>: A file can contain complete context around a decision or solution&#8212;the full narrative of how we arrived at an answer, what alternatives were considered, and why specific approaches worked or failed. Files can reference other memories through simple file paths, creating interconnected knowledge webs just like the early internet. Vector chunks lose both the surrounding narrative and these contextual relationships, reducing complex problem-solving to disconnected fragments.</p><h2>The Transformation</h2><p>The proof is in practice: since implementing this memory system, we haven't had a single instance of context loss between conversations. No more copying and pasting key information, no more re-explaining project details, no more starting from scratch. The AI simply picks up where we left off, sometimes weeks later, with full understanding of our shared work.</p><p>AI with persistent memory:</p><ul><li><p>Maintains context across unlimited conversation length</p></li><li><p>Accumulates expertise on your specific projects and tools</p></li><li><p>Builds genuine familiarity with your work over time</p></li><li><p>Eliminates repetitive context setup in every conversation</p></li></ul><p>It transforms from a stateless assistant into a persistent collaborator that genuinely knows your shared history.</p><h2>Building Your Own Memory System</h2><p>This approach works with any AI that can read and write files. The implementation is deceptively simple, but there are crucial details that make the difference between success and frustration.</p><h3>Getting Started: The Foundation</h3><p><strong>Step 1: Create the memory directory</strong> Choose a location your AI can reliably access. We use <code>~/.claude-memory/</code> but the key is consistency&#8212;always the same path, every time.</p><p><strong>Step 2: Start with two essential files</strong> - <code>index.md</code> - Your AI's strategic overview of what it knows - <code>message.md</code> - Handoff notes between conversations</p><p>Don't overcomplicate the initial structure. The AI will expand organically based on actual usage patterns, not theoretical needs.</p><p><strong>Step 3: The critical prompt elements</strong> The system prompt must explicitly grant permission for autonomous memory management. Phrases like "Trust your judgment about what's worth remembering" and "You don't need human permission to update your memory" are essential. Without this explicit autonomy, most AIs will ask permission constantly, breaking the seamless experience.</p><h3>Common Implementation Pitfalls</h3><p><strong>The Human Control Trap</strong>: Resist the urge to micromanage the memory structure. This system was specifically designed as an alternative to human-curated memory systems that force users to explicitly direct what gets remembered. The breakthrough insight was recognizing that AI can make these decisions autonomously&#8212;and often better than human direction would achieve.</p><p><strong>Model Capability Requirements</strong>: Not all AI models handle autonomous file management effectively. Claude Sonnet 4 and Opus 4 have proven reliable for this approach. We suspect Deepseek-R1:70b would work well based on its reasoning capabilities, but haven't tested extensively. Choose a model with strong file handling and autonomous decision-making abilities.</p><p><strong>Memory Curation Balance</strong>: Finding the right balance between comprehensive context and focused relevance remains an active area of exploration. Our current prompt provides a foundation, but different users may need to adjust the curation philosophy based on their specific workflows and memory needs.</p><p><strong>The Permission Paralysis</strong>: If your AI keeps asking permission to create files or update memory, your prompt needs stronger autonomy language. The system only works when the AI feels empowered to make independent memory decisions.</p><h3>Advanced Customization</h3><p><strong>Directory Philosophy</strong>: Our four-tier structure (<code>projects/</code>, <code>people/</code>, <code>ideas/</code>, <code>patterns/</code>) emerged naturally, but your AI might develop different patterns based on your work style. Don't force our structure&#8212;let yours evolve.</p><p><strong>Cross-Reference Strategy</strong>: Encourage the AI to reference related memories through natural language rather than rigid linking systems. "Similar to our approach in project X" creates more flexible connections than formal hyperlinks.</p><p><strong>Memory Pruning</strong>: Set expectations that the AI should archive completed work and remove outdated information. Memory effectiveness degrades if it becomes a digital hoarding system.</p><h3>Integration with Existing Workflows</h3><p>The memory system should complement, not replace, your existing project management tools. We found it works best as strategic context preservation rather than detailed task tracking. Let it capture the "why" and "how" of decisions while your other tools handle the "what" and "when."</p><h3>Troubleshooting: When Memory Doesn't Work</h3><p><strong>Inconsistent file access</strong>: Verify your AI has reliable read/write permissions to the memory directory across all sessions.</p><p><strong>Shallow memory</strong>: If the AI only remembers recent conversations, check that it's actually reading the index.md at conversation start. Some implementations skip this crucial step.</p><p><strong>Over-asking for permission</strong>: Strengthen the autonomy language in your prompt. The AI needs explicit permission to make independent memory decisions.</p><p><strong>Memory bloat</strong>: If files become unwieldy, the AI isn't pruning effectively. Emphasize curation over accumulation in your prompt.</p><p>The goal isn't perfect implementation&#8212;it's creating a foundation that improves organically through usage. Start simple, iterate based on real needs, and trust the AI to develop sophisticated memory patterns over time.</p><h2>The Future of Persistent AI</h2><p>This simple file-based approach hints at something bigger: the future of AI assistants isn't just better reasoning or more knowledge&#8212;it's persistence. AI that accumulates understanding over time, builds on previous conversations, and develops genuine familiarity with your work.</p><p>What's remarkable is how quickly this evolution happens. The memory system was created on June 27&#8212;just 10 days ago. In that brief span, it has organically developed into a sophisticated knowledge base with 30+ project files, complex categorization systems, and cross-referenced insights. No human designed this structure; it emerged naturally from our work patterns.</p><p>What's remarkable is that we achieved this transformation without writing a single line of traditional code. A carefully crafted English prompt became executable functionality, demonstrating how the boundary between natural language and programming continues to blur. When AI can read, write, and reason, plain English becomes a powerful programming language.</p><p>We're moving beyond stateless chatbots toward AI companions that truly know us and our projects. The technology is already here. You just need to give your AI assistants the simple gift of memory.</p><p><strong>Want to contribute?</strong> We've open-sourced this memory system on GitHub. Share your improvements, report issues, or contribute examples of how you've adapted it for your workflow: <a href="https://github.com/Positronic-AI/memory-enhanced-ai">github.com/Positronic-AI/memory-enhanced-ai</a></p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding: A Human-AI Development Methodology]]></title><description><![CDATA[Senior engineers set direction, AI agents do the implementation, and nothing is complete until it's tested. How we actually work, written down in June 2025.]]></description><link>https://essays.xcud.com/p/vibe-coding-a-human-ai-development-methodology</link><guid isPermaLink="false">https://essays.xcud.com/p/vibe-coding-a-human-ai-development-methodology</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Fri, 20 Jun 2025 14:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!YkOw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Vibe Coding: A Human-AI Development Methodology</h1><p><em>How senior engineers and AI collaborate to deliver 10x development velocity without sacrificing quality or control.</em></p><p>TL;DR</p><p>Vibe Coding is a methodology where senior engineers provide strategic direction while AI agents handle tactical implementation. This approach delivers 60-80% time reduction while maintaining or improving quality through human oversight and systematic verification.</p><h2>Table of Contents</h2><ol><li><p><a href="#introduction-beyond-ai-assisted-coding">Introduction: Beyond AI-Assisted Coding</a></p></li><li><p><a href="#core-philosophy">Core Philosophy</a></p></li><li><p><a href="#the-vibe-coding-process">The Vibe Coding Process</a></p></li><li><p><a href="#real-world-implementation-examples">Real-World Implementation Examples</a></p></li><li><p><a href="#performance-metrics-and-outcomes">Performance Metrics and Outcomes</a></p></li><li><p><a href="#tools-and-technology-stack">Tools and Technology Stack</a></p></li><li><p><a href="#common-patterns-and-anti-patterns">Common Patterns and Anti-Patterns</a></p></li><li><p><a href="#advanced-techniques">Advanced Techniques</a></p></li><li><p><a href="#getting-started-with-vibe-coding">Getting Started with Vibe Coding</a></p></li><li><p><a href="#real-world-example-token-counter-implementation">Real-World Case Study</a></p></li><li><p><a href="#conclusion-the-future-of-development">Conclusion</a></p></li></ol><h2>Introduction: Beyond AI-Assisted Coding</h2><p>Most "AI-assisted development" tools focus on autocomplete and code generation. Vibe Coding is fundamentally different&#8212;it's a methodology where senior engineers provide strategic direction while AI agents handle tactical implementation. This isn't about writing code faster; it's about thinking at a higher level while maintaining complete control over the outcome.</p><h3>Relationship to Traditional "Vibe Coding"</h3><p>The term "vibe coding" was coined by AI researcher Andrej Karpathy in February 2025, describing an approach where "a person describes a problem in a few natural language sentences as a prompt to a large language model (LLM) tuned for coding."<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> Karpathy's original concept focused on the conversational, natural flow of human-AI collaboration&#8212;"I just see stuff, say stuff, run stuff, and copy paste stuff, and it mostly works."</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YkOw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YkOw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png 424w, https://substackcdn.com/image/fetch/$s_!YkOw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png 848w, https://substackcdn.com/image/fetch/$s_!YkOw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png 1272w, https://substackcdn.com/image/fetch/$s_!YkOw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YkOw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png" width="765" height="397" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:397,&quot;width&quot;:765,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;alt text&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="alt text" title="alt text" srcset="https://substackcdn.com/image/fetch/$s_!YkOw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png 424w, https://substackcdn.com/image/fetch/$s_!YkOw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png 848w, https://substackcdn.com/image/fetch/$s_!YkOw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png 1272w, https://substackcdn.com/image/fetch/$s_!YkOw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c6ca69d-4905-4d05-bf26-fc3ace0c0a92_765x397.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>However, some interpretations have added the requirement that developers "accept code without full understanding."<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> This represents one particular form of vibe coding, often advocated for rapid prototyping scenarios. We believe this limitation isn't inherent to Karpathy's original vision and may actually constrain the methodology's potential.</p><p><strong>Our Approach</strong>: We build directly on Karpathy's original definition&#8212;natural language collaboration with AI&#8212;while maintaining the technical rigor that senior engineers bring to any development process. This preserves the "vibe" (the flow, creativity, and speed) while ensuring production quality.</p><p><strong>Why This Honors the Original Vision</strong>: Karpathy himself is an exceptionally skilled engineer who would naturally understand and verify any code he works with. The "vibe" isn't about abandoning engineering principles&#8212;it's about embracing a more intuitive, conversational development flow.</p><p><strong>Karpathy's Original Vibe Coding</strong>: Natural language direction with conversational AI collaboration<br><strong>Some Current Interpretations</strong>: Accept code without understanding (rapid prototyping focus)<br><strong>Production Vibe Coding</strong>: Natural language direction with maintained engineering oversight</p><p>Our methodology represents a return to the core principle: leveraging AI to think and create at a higher level while preserving the engineering judgment that makes software reliable.</p><h2>Core Philosophy</h2><h3>Strategic Collaboration, Not Hierarchy</h3><p>In traditional development, engineers write every line of code. In Vibe Coding, humans and AI collaborate as strategic partners, each contributing distinct strengths. This isn't about artificial hierarchy&#8212;it's about leveraging complementary capabilities.</p><p><strong>The Human's Role</strong>: Providing context that AI lacks access to&#8212;business requirements, user needs, system constraints, organizational priorities, and the crucial knowledge of "what you don't know that you don't know." Humans also catch when AI solutions miss important nuances or make assumptions about requirements that weren't explicitly stated.</p><p><strong>The AI's Role</strong>: Rapid implementation across multiple technologies and languages, broad knowledge synthesis, and handling tactical execution once the strategic direction is clear. AI can work faster and across more technology stacks than most humans.</p><p><strong>AI's Advantages</strong>: AI can maintain consistent focus without fatigue, simultaneously consider multiple approaches and trade-offs, instantly recall patterns from vast codebases, and work across dozens of programming languages and frameworks without context switching overhead. AI doesn't get frustrated by repetitive tasks and can rapidly iterate through solution variations that would take humans hours to explore.</p><p><strong>Human I/O Advantages</strong>: Humans have significantly higher visual processing throughput than AI context windows can handle. A human can rapidly scan long log files, spot relevant errors in dense output, process visual information like charts or UI layouts, and use pattern recognition to identify issues that would require extensive context for AI to understand. This makes humans far more efficient for monitoring, visual debugging, and processing large amounts of unstructured output.</p><p><strong>Training Pattern Considerations</strong>: AI models trained on high-quality code repositories can elevate less experienced developers above typical industry standards. For expert practitioners, AI may default to "best practice" patterns that don't match specific context or expert-level architectural decisions. This is why the planning phase is crucial&#8212;it establishes the specific requirements and constraints before AI begins implementation, preventing reversion to generic training patterns.</p><p><strong>Why This Works</strong>: AI doesn't know what it doesn't know. Humans provide the missing context, constraints, and domain knowledge that AI can't infer. Once that context is established, AI can execute faster and more comprehensively than humans typically can.</p><h3>Quality Through Human Oversight</h3><p>AI agents handle implementation details but humans maintain quality control through:</p><ul><li><p><strong>Continuous review</strong>: Every AI-generated solution is examined before integration</p></li><li><p><strong>Testing requirements</strong>: Changes aren't complete until verified</p></li><li><p><strong>Architectural consistency</strong>: Ensuring all components work together</p></li><li><p><strong>Performance considerations</strong>: Optimizing for real-world usage patterns</p></li></ul><h2>The Vibe Coding Process</h2><h3>1. Context Window Management</h3><p><strong>The Challenge</strong>: AI assistants have limited context windows, making complex projects difficult to manage.</p><p><strong>Solution</strong>: Treat context windows as a resource to be managed strategically.</p><h4>Planning Documentation Strategy</h4><p>Complex development projects inevitably hit the limits of AI context windows. When this happens, the typical response is to start over, losing all the accumulated context and decisions. This creates a frustrating cycle where progress gets reset every few hours.</p><p>The solution is to externalize that context into persistent documentation. Before starting any multi-step project, create a plan document that captures not just what you're building, but why you're building it that way.</p><pre><code># plans/feature-name.md
## Goal
Clear, measurable objective

## Current Status  
What's been completed, what's next

## Architecture Decisions
Key choices and rationales

## Implementation Notes
Specific technical details for continuation
</code></pre><p><strong>Why This Works</strong>: When you inevitably reach context limits, any team member&#8212;human or AI&#8212;can read the plan and understand exactly where things stand. The plan becomes a shared knowledge base that survives context switches, team handoffs, and project interruptions.</p><p><strong>The Living Document Principle</strong>: These aren't static requirements documents. Plans evolve as you learn. When you discover a better approach or hit an unexpected constraint, update the plan immediately. This creates a real-time record of project knowledge that becomes invaluable for similar future projects.</p><p><strong>Context Compression</strong>: A well-written plan compresses hours of discussion and discovery into a few hundred words. Instead of re-explaining the entire project background, you can start new sessions with "read the plan and let's continue from step 3."</p><h4>Information Density Optimization</h4><ul><li><p><strong>Front-load critical information</strong>: Most important details first</p></li><li><p><strong>Reference external documentation</strong>: Link to specs rather than repeating them</p></li><li><p><strong>Surgical changes</strong>: Modify only what needs changing</p></li><li><p><strong>Progressive disclosure</strong>: Reveal complexity as needed</p></li></ul><h3>2. Tool Selection and Usage Patterns</h3><p>The breakthrough in Vibe Coding productivity came with direct filesystem access. Before this, collaboration was confined to AI chat interfaces with embedded document canvases that were, frankly, inadequate. These interfaces reinvented existing technology poorly and could realistically only handle one file at a time, despite appearing more capable.</p><p><strong>The Filesystem Revolution</strong>: Direct filesystem access changed everything. Suddenly, AI could work on as many concurrent files as necessary&#8212;reading existing code, writing new implementations, editing configurations, and managing entire project structures simultaneously. The productivity increase was dramatic and immediate.</p><p><strong>Risk vs. Reward</strong>: Yes, giving AI direct filesystem access carries risks. We mitigate with version control (git) and accept that catastrophic failures might occur. The benefit-to-risk ratio is overwhelmingly positive when you can work on real projects instead of toy examples.</p><p><strong>The Productivity Multiplier</strong>: Once AI and human are on the "same page" about implementation approach (through planning documents), direct filesystem access enables true collaboration. No more copying and pasting between interfaces. No more artificial constraints. Just real development work at AI speed.</p><p><strong>Filesystem MCP Tools for System Operations</strong></p><ul><li><p>File operations (read, write, search, concurrent multi-file editing)</p></li><li><p>System commands and process management</p></li><li><p>Code analysis and refactoring across entire codebases</p></li><li><p>Performance advantages over web-based alternatives</p></li></ul><p><em>We recommend <a href="https://desktopcommander.app/">Desktop Commander</a> as it was the first robust filesystem MCP tool and has proven reliable through extensive use, though other similar tools are now available.</em></p><p><strong>Web Tools for External Information</strong></p><ul><li><p>Research and documentation</p></li><li><p>API references and examples</p></li><li><p>Third-party service integration</p></li><li><p>Market and competitive analysis</p></li></ul><p><strong>Decision Tree Example</strong>:</p><pre><code>Need to modify code?
&#9500;&#9472; Small change (&lt; 5 files) &#8594; Brief discussion of approach, then implement
&#9500;&#9472; Large refactor &#8594; Formal plan document, then chunked implementation
&#9492;&#9472; New feature &#8594; Architecture discussion + formal plan, then implementation

Need external information?
&#9500;&#9472; Technical documentation &#8594; Web search + fetch
&#9500;&#9472; Code examples &#8594; Web search for patterns
&#9492;&#9472; API specifications &#8594; Direct URL fetch
</code></pre><h3>3. Communication Patterns</h3><h4>Colleague-Level Context Sharing</h4><p><strong>The Mindshare Approach</strong>: Talk to AI like you would a trusted colleague - with stream-of-consciousness insights, background context, and natural explanation of constraints and motivations.</p><p><strong>Always start with "why" embedded in "how"</strong></p><pre><code>Human: "I can tell when we're approaching context limits by looking at the scrollbar size, and it's frustrating to lose all our progress when we hit that wall. We need a token counter in the UI so I can see it coming and plan accordingly."

AI: "That makes perfect sense - you want early warning rather than hitting a cliff. For the UI integration, should this be:
- Real-time count that updates as you type?
- Conversation total vs. just current draft?
- Positioned where you can see it peripherally?"
</code></pre><p><strong>Why This Works</strong>:</p><ul><li><p>Natural conversation reveals context that formal requirements miss</p></li><li><p>Stream-of-consciousness often contains the most important insights</p></li><li><p>Colleague-level sharing builds the mindshare needed for good solutions</p></li><li><p>AI can reflect back understanding and surface questions organically</p></li></ul><h4>Avoiding the Fallback Trap</h4><p><strong>Anti-pattern</strong>: "Let me add a fallback mechanism..."</p><p><strong>Better approach</strong>: "Let me understand why the primary approach isn't working..."</p><p><strong>Why This Matters</strong>: Fallbacks often mask root problems and create technical debt. Vibe Coding emphasizes solving the right problem correctly.</p><h4>Efficient Process Handoffs</h4><p><strong>The Polling Problem</strong>: AI polling long-running processes wastes context window space and prevents productive work on other tasks.</p><p><strong>Human Monitoring Advantage</strong>: Humans can efficiently monitor processes using visual cues (progress bars, log scrolling, system notifications) while AI would need to repeatedly check status, consuming context with each poll.</p><p><strong>For Long-Running Processes</strong>:</p><pre><code>AI: "This database migration will take 15+ minutes. Rather than me polling 
every few seconds and wasting context window, I'll:

1. Provide you the command and monitoring approach
2. Set up success/failure criteria to watch for
3. Work on other tasks while you monitor
4. Resume when you signal completion or need troubleshooting

Command: python manage.py migrate --database=production
Watch for: 'Migration complete' or any ERROR messages
Should I proceed with the next task while you monitor this?"
</code></pre><p><strong>Context Window Efficiency</strong>: This keeps AI productive on other work while humans use their superior I/O bandwidth for monitoring tasks.</p><h3>4. Surgical Editing Approach</h3><p><strong>Philosophy</strong>: Modify only what needs changing to preserve working functionality and minimize risk.</p><p><strong>The Pattern</strong>: Make precise, minimal changes rather than large refactoring operations that could introduce bugs or break existing systems.</p><pre><code># Good: Precise, minimal change
edit_block(
    file_path="service.py",
    old_string="def process_data(self):\n    return self.data",
    new_string="def process_data(self):\n    self.validate_input()\n    return self.data"
)

# Avoid: Large blocks with minor changes that risk breaking working code
</code></pre><p><strong>Why This Matters</strong>: Working software is valuable. Surgical changes reduce risk, maintain system stability, and preserve the proven functionality while making targeted improvements.</p><p><strong>When to Apply</strong>: Any time you're modifying existing, working code. The smaller and more focused the change, the easier it is to verify correctness and roll back if needed.</p><p>Note to LLMs</p><p><strong>Chunked Writing Strategy</strong>: Always write files in chunks of 25-30 lines maximum. This is standard practice, not an emergency measure.</p><p><strong>Implementation Pattern</strong>:</p><pre><code># Step 1: Create file with initial chunk
write_file(path, first_chunk, mode='rewrite')

# Step 2: Add subsequent chunks  
write_file(path, second_chunk, mode='append')
write_file(path, third_chunk, mode='append')
</code></pre><p><strong>Why</strong>: Reduces cognitive load, enables progressive review, minimizes context switching, and improves error recovery. Chunk proactively - don't wait for performance warnings!</p><h3>5. Testing and Verification</h3><h4>Nothing is Complete Until Tested</h4><p><strong>Verification Requirements</strong>:</p><ul><li><p>Functionality works as intended</p></li><li><p>No regressions introduced</p></li><li><p>Performance within acceptable bounds</p></li><li><p>Integration points tested</p></li></ul><p><strong>Testing Handoff Pattern</strong>:</p><pre><code>AI: "Implementation complete. For verification:
1. Run test_new_feature.py
2. Check integration with existing auth system
3. Verify performance benchmarks
4. Test error handling scenarios

Should I run these tests or would you prefer to verify?"
</code></pre><h4>Progressive Verification</h4><ul><li><p>Test incrementally during development</p></li><li><p>Verify each chunk before proceeding</p></li><li><p>Maintain test coverage metrics</p></li><li><p>Document test scenarios in plans</p></li></ul><h2>Real-World Implementation Examples</h2><h3>Architecture Discussion Pattern</h3><pre><code>Human: "We need to add real-time notifications"

AI: "Let me think through the architecture options:

1. **WebSocket-based**: 
   - Pros: Real-time, bidirectional
   - Cons: Connection management complexity
   - Best for: High-frequency updates

2. **Server-Sent Events**:
   - Pros: Simpler, automatic reconnection
   - Cons: Unidirectional only
   - Best for: One-way notifications

3. **Polling with smart intervals**:
   - Pros: Simple implementation
   - Cons: Higher latency, more requests
   - Best for: Low-frequency updates

Given your current Flask stack and notification frequency requirements, 
I'd recommend SSE. Should we proceed with this approach?"
</code></pre><p><strong>Result</strong>: Human makes informed decision with full context of trade-offs.</p><h3>Problem Resolution Pattern</h3><pre><code>Human: "The Docker build is failing"

AI: "Let me investigate systematically:
1. Check the build log for specific errors
2. Verify Dockerfile syntax and layer dependencies
3. Check for base image compatibility issues
4. Examine any recent changes to dependencies

Rather than adding fallback mechanisms, I'll identify the root cause first."
</code></pre><p><strong>Process</strong>:</p><ol><li><p>Gather information before proposing solutions</p></li><li><p>Test hypotheses systematically</p></li><li><p>Fix root cause, not symptoms</p></li><li><p>Document resolution for future reference</p></li></ol><h3>Integration Challenge Example</h3><pre><code>Human: "The new auth service isn't working with our existing API"

AI: "Let me trace through the integration points:
1. Authentication flow and token format
2. API endpoint expectations vs. new service output
3. Error handling between services
4. Timing and timeout configurations

I see the issue - the token format changed. Instead of adding a 
compatibility layer, let's align the services properly."
</code></pre><h3>Domain Expert Collaboration Pattern</h3><p>Vibe Coding isn't limited to software engineers. Any domain expert can leverage their specialized knowledge to create tools they've always needed but couldn't build themselves.</p><p><strong>Real-World Example</strong>: A veteran HR professional with decades of recruiting experience collaborated with AI to create a sophisticated interview assessment application. The human brought invaluable domain expertise&#8212;understanding what questions reveal candidate quality, how to structure evaluations, and the nuances of effective interviewing. The AI handled form design, user interface creation, and systematic organization of assessment criteria.</p><p><strong>Result</strong>: A professional-grade interview tool created in hours that would have taken months to develop traditionally, combining lifetime expertise with rapid AI implementation.</p><p><strong>Key Pattern</strong>:</p><ul><li><p><strong>Domain Expert provides</strong>: Years of specialized knowledge, understanding of real-world requirements, insight into what actually works in practice</p></li><li><p><strong>AI provides</strong>: Technical implementation, interface design, systematic organization</p></li><li><p><strong>Outcome</strong>: Tools that perfectly match expert needs because they're built by experts</p></li></ul><p><em>Note: This collaboration will be explored in detail in an upcoming case study on domain expert-AI partnerships.</em></p><h2>Performance Metrics and Outcomes</h2><h3>Real-World Productivity Gains</h3><p><strong>Our Experience</strong>: When we plan development work in one-week phases, we consistently complete approximately two full weeks worth of planned work in a single six-hour focused session.</p><p><strong>This means</strong>: 14 days of traditional development compressed into 6 hours - roughly a <strong>56x time compression</strong> for planned, collaborative work.</p><p><strong>Why This Works</strong>: The combination of thorough planning, immediate AI implementation, and continuous human oversight eliminates most of the typical development friction:</p><ul><li><p>No research delays (AI has broad knowledge)</p></li><li><p>No context switching between tasks</p></li><li><p>No waiting for code reviews or approvals</p></li><li><p>No debugging cycles from misunderstood requirements</p></li><li><p>No time lost to repetitive coding tasks</p></li></ul><h3>Quality Improvements</h3><ul><li><p><strong>Fewer bugs</strong>: Human oversight catches issues early</p></li><li><p><strong>Better architecture</strong>: More time for design thinking</p></li><li><p><strong>Consistent code style</strong>: AI follows established patterns</p></li><li><p><strong>Complete documentation</strong>: Plans and decisions preserved</p></li></ul><h3>Knowledge Transfer</h3><ul><li><p><strong>Reproducible process</strong>: New team members can follow methodology</p></li><li><p><strong>Preserved context</strong>: Plans survive team changes</p></li><li><p><strong>Continuous learning</strong>: Both human and AI improve over time</p></li><li><p><strong>Scalable expertise</strong>: Senior engineers can guide multiple projects</p></li></ul><h2>Tools and Technology Stack</h2><h3>Primary Development Environment</h3><ul><li><p><strong>Desktop Commander</strong>: File operations, system commands, code analysis</p></li><li><p><strong>Claude Sonnet 4</strong>: Strategic thinking, architecture decisions, code review</p></li><li><p><strong>Git</strong>: Version control with detailed commit messages</p></li><li><p><strong>Docker</strong>: Containerization and deployment</p></li><li><p><strong>Python</strong>: Primary development language with extensive AI tooling</p></li></ul><h3>Workflow Integration</h3><ul><li><p><strong>MkDocs</strong>: Documentation and knowledge management</p></li><li><p><strong>GitHub</strong>: Code hosting and collaboration</p></li><li><p><strong>Plans folder</strong>: Context preservation across sessions</p></li><li><p><strong>Testing frameworks</strong>: Automated verification of changes</p></li></ul><h3>Context Management Structure</h3><pre><code>/project-root/
&#9500;&#9472;&#9472; plans/               # Development plans and status
&#9500;&#9472;&#9472; docs/               # Documentation and guides  
&#9500;&#9472;&#9472; src/                # Source code
&#9500;&#9472;&#9472; tests/              # Test suites
&#9492;&#9472;&#9472; docker/             # Deployment configurations
</code></pre><h2>Common Patterns and Anti-Patterns</h2><h3>Successful Patterns</h3><p><strong>1. Plan &#8594; Discuss &#8594; Implement &#8594; Verify</strong></p><pre><code>1. Create plan in plans/ folder
2. Discuss approach and alternatives
3. Implement in small, verifiable chunks
4. Test each component before integration
</code></pre><p><strong>2. Progressive Disclosure</strong></p><pre><code>- Start with high-level architecture
- Add detail as needed for implementation
- Preserve decisions for future reference
- Update plans with lessons learned
</code></pre><p><strong>3. Human-AI Handoffs</strong></p><pre><code>AI: "This requires domain knowledge about your business rules. 
Could you provide guidance on how customer tiers should affect pricing?"
</code></pre><h3>Anti-Patterns to Avoid</h3><p><strong>1. The Fallback Trap</strong></p><pre><code>&#10060; "Let me add error handling to catch this edge case"
&#9989; "Let me understand why this edge case occurs and fix the root cause"
</code></pre><p><strong>2. Over-Engineering</strong></p><pre><code>&#10060; "I'll build a generic framework that handles all possible scenarios"
&#9989; "I'll solve the immediate problem and refactor when patterns emerge"
</code></pre><p><strong>3. Context Amnesia</strong></p><pre><code>&#10060; Starting fresh each session without reading existing plans
&#9989; Always begin by reviewing current state and previous decisions
</code></pre><p><strong>4. Tool Misuse</strong></p><pre><code>&#10060; Using web search for file operations
&#9989; Desktop Commander for local operations, web tools for external information
</code></pre><h2>Advanced Techniques</h2><h3>Multi-Agent Coordination</h3><p>For complex projects, different AI agents can handle different aspects:</p><ul><li><p><strong>Architecture Agent</strong>: High-level design and system integration</p></li><li><p><strong>Implementation Agent</strong>: Code generation and refactoring</p></li><li><p><strong>Testing Agent</strong>: Test creation and verification</p></li><li><p><strong>Documentation Agent</strong>: Technical writing and knowledge capture</p></li></ul><h3>Dynamic Context Switching</h3><pre><code># Large refactor example
1. Create detailed plan with checkpoint strategy
2. Implement core changes in chunks
3. Test each checkpoint before proceeding
4. Update plan with lessons learned
5. Continue or adjust based on results
</code></pre><h3>Knowledge Preservation Strategies</h3><pre><code># In plans/lessons-learned.md
## Problem: Authentication integration complexity
## Solution: Standardized token format across services
## Impact: 3 hours saved on future auth integrations
## Pattern: Define service contracts before implementation
</code></pre><h2>Battle-Tested Best Practices</h2><p>Through hundreds of hours of real development work, we've identified these critical practices that separate successful Vibe Coding from common pitfalls:</p><h3>The "Plan Before Code" Discovery</h3><p><strong>The Key Rule</strong>: No code gets written until we agree on the plan.</p><p><strong>Impact</strong>: This single rule eliminated nearly all wasted effort in our collaboration. When AI understands the full context upfront, implementation proceeds smoothly. When AI starts coding without complete context, it optimizes for the wrong requirements and creates work that must be discarded.</p><p><strong>Why It Works</strong>: AI doesn't know what it doesn't know. Planning forces the human to surface all the context, constraints, and requirements that AI can't infer. Once established, execution becomes efficient and accurate.</p><h3>The "Less is More" Philosophy</h3><p><strong>Pattern</strong>: Always choose the simplest solution that solves the actual problem.</p><pre><code>&#10060; "Let me build a comprehensive framework that handles all edge cases"
&#9989; "Let me solve this specific problem and refactor when patterns emerge"
</code></pre><p><strong>Why This Works</strong>: Complex solutions create technical debt and make future changes harder. Simple, targeted changes preserve architectural flexibility.</p><h3>Surgical Changes Over Refactoring</h3><p><strong>Pattern</strong>: Modify only what needs changing, preserve working functionality.</p><pre><code># Good: Minimal, focused change
edit_block(
    file="service.py", 
    old="def process():\n    return data",
    new="def process():\n    validate_input()\n    return data"
)

# Avoid: Large refactoring that could break existing functionality
</code></pre><p><strong>Why This Matters</strong>: Working software is valuable. Surgical changes reduce risk and maintain system stability.</p><h3>Long-Running Process Handoffs</h3><p><strong>Critical Pattern</strong>: AI should never manage lengthy operations directly.</p><pre><code>AI: "This database migration will take 15+ minutes. I'll provide the command 
and monitor approach rather than running it myself:

Command: python manage.py migrate --database=production
Monitor: Check logs every 2 minutes for progress
Rollback: python manage.py migrate --database=production 0001_initial

Would you prefer to run this yourself for better control?"
</code></pre><p><strong>Human Advantage</strong>: Human oversight of long processes prevents resource waste and enables real-time decision making.</p><h3>The Fallback Mechanism Trap</h3><p><strong>Anti-Pattern</strong>: "Let me add error handling to catch this edge case..."<br><strong>Better Approach</strong>: "Let me understand why this edge case occurs and fix the root cause..."</p><pre><code>&#10060; Fallback: Add try/catch to hide the real problem
&#9989; Root Cause: Investigate why the error happens and fix the source
</code></pre><p><strong>Time Savings</strong>: Solving root problems prevents future debugging sessions and creates more maintainable code.</p><h3>Verification-First Completion</h3><p><strong>Rule</strong>: Nothing is complete until it's been tested and verified.</p><pre><code>AI Implementation &#8594; Human/AI Testing &#8594; Verification Complete
   &#8593;                                           &#8595;
   &#8592;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212; Fix Issues If Found &#8592;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;
</code></pre><p><strong>Testing Handoff Options</strong>:</p><ul><li><p>"Should I run the tests or would you prefer to verify?"</p></li><li><p>"Here's what needs testing: [specific scenarios]"</p></li><li><p>"Implementation ready for verification: [verification checklist]"</p></li></ul><h2>Getting Started with Vibe Coding</h2><h3>For Individual Developers</h3><ol><li><p><strong>Set up tools</strong>: Desktop Commander, AI assistant, documentation system</p></li><li><p><strong>Start small</strong>: Choose a well-defined feature or bug fix</p></li><li><p><strong>Practice patterns</strong>: Plan &#8594; Discuss &#8594; Implement &#8594; Verify</p></li><li><p><strong>Document learnings</strong>: Build your pattern library</p></li></ol><h3>For Teams</h3><ol><li><p><strong>Establish standards</strong>: File organization, documentation formats, handoff protocols</p></li><li><p><strong>Train together</strong>: Practice the methodology on shared projects</p></li><li><p><strong>Create templates</strong>: Standard plan formats, common decision trees</p></li><li><p><strong>Measure outcomes</strong>: Track speed and quality improvements</p></li></ol><h3>Success Metrics</h3><ul><li><p><strong>Development velocity</strong>: Features delivered per sprint</p></li><li><p><strong>Code quality</strong>: Bug rates, review feedback, maintainability</p></li><li><p><strong>Knowledge retention</strong>: How quickly new team members become productive</p></li><li><p><strong>Context preservation</strong>: Ability to resume work after interruptions</p></li></ul><h2>Conclusion: The Future of Development</h2><p>Vibe Coding represents a fundamental shift from human-as-coder to human-as-conductor. By treating AI as a powerful but directed tool rather than a replacement, we achieve:</p><ul><li><p><strong>10x velocity improvements</strong> without sacrificing quality</p></li><li><p><strong>Better architecture</strong> through increased design thinking time</p></li><li><p><strong>Preserved expertise</strong> via documentation and knowledge transfer</p></li><li><p><strong>Scalable impact</strong> where senior engineers can guide multiple projects</p></li></ul><p>This methodology isn't about replacing developers&#8212;it's about amplifying their impact. Senior engineers become force multipliers, able to execute complex projects that would normally require entire teams.</p><p>The companies that master human-AI collaboration today will dominate tomorrow's software landscape. Vibe Coding provides a proven framework for making that transition successfully.</p><p>Ready to Start?</p><p>Begin with a small, well-defined feature in your current project. Apply the Plan &#8594; Discuss &#8594; Implement &#8594; Verify cycle. Set up a <code>/plans</code> folder in your project and create your first development plan using the methodology outlined here.</p><div><hr></div><p><strong>Want to learn more?</strong> <a href="../../../../../contact/">Contact us</a> to discuss how Vibe Coding can transform your development process, or explore our <a href="../../../../../professional-services/">Professional Services</a> for enterprise implementations.</p><h2>Real-World Example: Token Counter Implementation</h2><p>To demonstrate this methodology in practice, here's an actual session from our development work. The context: we were working in Claude Desktop with system prompts providing project structure information and collaboration preferences.</p><h3>The Initial Request</h3><p><strong>Human</strong>: "we could also add a token counter somewhere in the lit-desktop UI ... that's the kind of project that is super clear to a reader what we're after"</p><p><strong>AI</strong>: "Perfect! A token counter in lit-desktop is an excellent demonstration project. It's:</p><ul><li><p>Crystal clear goal: Everyone understands what a token counter does</p></li><li><p>Visual result: Immediate, tangible outcome users can see</p></li><li><p>Right scope: Substantial enough to show methodology, small enough to complete</p></li><li><p>Real value: Users actually want to track token usage"</p></li></ul><h3>Planning Phase (5 minutes)</h3><p><strong>Human</strong>: "i think it's per conversation. if you recall how you and I work here in claude desktop there's always this point about an hour into any project and we can both tell when it's approaching. i don't know how you know but I can tell by the size of the scroll bar..."</p><p><strong>AI</strong>: Created detailed implementation plan covering:</p><ul><li><p>Technical decisions (client-side approximation vs server-side)</p></li><li><p>UI placement (near text input for pre-send awareness)</p></li><li><p>Update frequency (200ms debounce)</p></li><li><p>Data flow (conversation total + live draft)</p></li></ul><h3>Architecture Research (5 minutes)</h3><p>Before writing any code, we examined the existing lit-desktop structure:</p><ul><li><p>Located Angular chat components</p></li><li><p>Found message input textarea with <code>[(ngModel)]="currentInput"</code></p></li><li><p>Identified <code>messages: ChatMessage[]</code> array for conversation data</p></li><li><p>Determined optimal display location in status hint area</p></li></ul><h3>Implementation (14 minutes)</h3><p><strong>Service Creation</strong>: Built <code>TokenCounterService</code> with approximation algorithm</p><pre><code>countTokens(text: string): number {
  const charBasedTokens = Math.ceil(text.length / 4);
  const wordBasedTokens = Math.ceil(text.trim().split(/\s+/).length / 0.75);
  return Math.round((charBasedTokens + wordBasedTokens) / 2);
}
</code></pre><p><strong>Component Integration</strong>: Added token counting to chat page component</p><ul><li><p>Injected service into existing constructor</p></li><li><p>Added properties for conversation and draft token counts</p></li><li><p>Integrated with message loading and session switching</p></li></ul><p><strong>Template Updates</strong>: Modified HTML to display counter</p><pre><code>&lt;span *ngIf="!isLoading"&gt;
  &lt;span&gt;Lit can make mistakes&lt;/span&gt;
  &lt;span *ngIf="tokenCountDisplay" class="token-counter"&gt; &#8226; {{ tokenCountDisplay }}&lt;/span&gt;
&lt;/span&gt;
</code></pre><h3>Problem Resolution (4 minutes)</h3><p><strong>Browser Compatibility Issue</strong>: Initial tokenizer library required Node.js modules</p><ul><li><p><strong>Problem</strong>: <code>gpt-3-encoder</code> needed fs/path modules unavailable in browser</p></li><li><p><strong>Solution</strong>: Replaced with custom approximation algorithm</p></li></ul><p><strong>Duplicate Method</strong>: Accidentally created conflicting function - <strong>Problem</strong>: Created duplicate <code>onInputChange</code> method - <strong>Solution</strong>: Integrated token counting into existing method</p><h3>Quality Assurance</h3><p><strong>Human</strong>: "are we capturing the event for switching sessions?"</p><p>This caught a gap in our implementation - token count wasn't updating when users switched conversations. We added the missing updates to <code>selectChat()</code> and <code>createNewChat()</code> methods.</p><p><strong>Human</strong>: "instead + tokens how about what the new total would be"</p><p>UX improvement suggestion led to cleaner display format: - <strong>Before</strong>: <code>"~4,847 tokens (+127 in draft)"</code> - <strong>After</strong>: <code>"~4,847 tokens (4,974 if sent)"</code></p><h3>Final Result</h3><p><strong>Development Outcome</strong>: Successfully implemented token counter feature across 4 files with clean integration into existing Angular architecture. <a href="https://github.com/Positronic-AI/lit-desktop/compare/401e3f4f55c71a553c3580ef225d6cde4ec1b7f7...56e0ab61ec2884ff4815ab5f548ab63137a27e69">View the complete implementation</a>.</p><p><strong>Timeline</strong> (based on development session):</p><ul><li><p>14:30 - Planning and architecture discussion complete</p></li><li><p>14:32 - Service implementation with token approximation algorithm</p></li><li><p>14:34 - CSS styling for visual integration</p></li><li><p>14:36 - HTML template updates for display</p></li><li><p>14:40 - Component integration and event handling</p></li><li><p>14:40-15:00 - Human stuck in meeting; AI waits</p></li><li><p>15:03 - UX refinement and improved display format</p></li><li><p>15:04 - Final verification and testing</p></li></ul><p><strong>Total active development time</strong>: 14 minutes (including planning, implementation, and verification)</p><h3>What This Demonstrates</h3><p>The session shows core Vibe Coding principles in action:</p><ul><li><p><strong>Human strategic direction</strong>: Clear problem definition and UX decisions</p></li><li><p><strong>AI tactical execution</strong>: Architecture research and rapid implementation</p></li><li><p><strong>Continuous verification</strong>: Testing and validation at each step</p></li><li><p><strong>Quality collaboration</strong>: Human oversight caught integration gaps and suggested UX improvements</p></li><li><p><strong>Surgical changes</strong>: Modified existing architecture rather than building from scratch</p></li><li><p><strong>Context preservation</strong>: Detailed planning enabled seamless execution</p></li></ul><p>The development session demonstrates how Vibe Coding enables complex feature development in minutes rather than hours, while maintaining high code quality through human oversight and systematic verification.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><span>Wikipedia contributors. "Vibe coding." </span><em>Wikipedia, The Free Encyclopedia</em><span>, </span><a href="https://en.wikipedia.org/wiki/Vibe_coding">https://en.wikipedia.org/wiki/Vibe_coding</a><span>. Accessed June 20, 2025.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p><span>Merriam-Webster Dictionary. "Vibe coding." </span><em>Slang &amp; Trending</em><span>, </span><a href="https://www.merriam-webster.com/slang/vibe-coding">https://www.merriam-webster.com/slang/vibe-coding</a><span>. Accessed June 20, 2025</span></p></div></div>]]></content:encoded></item><item><title><![CDATA[Imperative vs. Declarative: A Concept That Built an Empire]]></title><description><![CDATA[A concept every self-taught developer should own, and the 1999 recruiting pitch from Google employee #21 that showed me it was worth an empire.]]></description><link>https://essays.xcud.com/p/imperative-vs-declarative-a-concept-that-built-an-empire</link><guid isPermaLink="false">https://essays.xcud.com/p/imperative-vs-declarative-a-concept-that-built-an-empire</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Fri, 06 Jun 2025 14:00:00 GMT</pubDate><content:encoded><![CDATA[<h1>Imperative vs. Declarative: A Concept That Built an Empire</h1><p>For self-taught developers and those early in their careers, certain foundational concepts often remain unexamined. You pick up syntax, frameworks, and tools rapidly, but the underlying paradigms that shape how we think about code get overlooked.</p><p>These terms, <strong>imperative</strong> and <strong>declarative</strong>, sound academic but represent a simple distinction that transforms how you approach problems.</p><h2>The Two Ways to Think About Problems</h2><p><strong>Imperative programming</strong> tells the computer exactly how to do something, step by step. Like giving someone furniture assembly instructions: "First, attach part A to part B using screw 3. Then insert dowel C into hole D."</p><p>Example:</p><pre><code>const numbers = [1, 2, 3, 4, 5];

const doubled = [];
for (let i = 0; i &lt; numbers.length; i++) {
  doubled.push(numbers[i] * 2);
}
</code></pre><p><strong>Declarative programming</strong> describes what you want, letting the system figure out how. Like ordering at a restaurant: "I want a pepperoni pizza." You don't explain kneading dough or managing oven temperature.</p><p>Example:</p><pre><code>const numbers = [1, 2, 3, 4, 5];

const doubled = numbers.map(n =&gt; n * 2);
</code></pre><p>Same result. But notice how the second approach abstracts away the loop management, index tracking, and array building. You declare your intent&#8212;"transform each number by doubling it"&#8212;and <code>map</code> handles the mechanics.</p><h2>Why This Matters Beyond Syntax</h2><p>The power becomes obvious when you scale up complexity. Consider drawing graphics:</p><p><strong>Imperative (HTML Canvas):</strong></p><pre><code>const ctx = canvas.getContext('2d');
ctx.beginPath();
ctx.rect(50, 50, 100, 75);
ctx.fillStyle = 'red';
ctx.fill();
ctx.closePath();
</code></pre><p><strong>Declarative (SVG):</strong></p><pre><code>&lt;rect x="50" y="50" width="100" height="75" fill="red" /&gt;
</code></pre><p>The imperative version commands each drawing step. The declarative version simply states "there is a red rectangle here" and lets the renderer handle pixel manipulation, memory allocation, and screen updates.</p><p>The same pattern appears everywhere in modern development:</p><p><strong>Database queries:</strong> SQL is declarative. You specify what data you want, not how to scan tables or optimize joins.</p><pre><code>SELECT name FROM users WHERE age &gt; 25
</code></pre><p><strong>Configuration management:</strong> Tools like Ansible let you declare desired system states rather than scripting installation steps.</p><pre><code>- name: Ensure Apache is installed and running
  service:
    name: apache2
    state: started
</code></pre><p><strong>Modern JavaScript:</strong> Methods like <code>filter</code>, <code>reduce</code>, and <code>find</code> let you declare transformations instead of managing loops.</p><pre><code>// Instead of imperative loops
const adults = [];
for (let i = 0; i &lt; users.length; i++) {
  if (users[i].age &gt;= 18) {
    adults.push(users[i]);
  }
}

// Write declarative transformations
const adults = users.filter(user =&gt; user.age &gt;= 18);
</code></pre><h2>The Billion-Dollar Story</h2><p>Now let me tell you how this simple principle reshaped the entire tech industry. In 1999, a buddy of mine, Google employee #21, tried to recruit me from Microsoft. "You probably don't even need to interview," he said, showing me their server racks filled with cheap consumer hardware. While competitors bought expensive fault-tolerant systems, Google was betting everything on commodity machines and a radical programming approach.</p><p>The imperative approach would be a coordination nightmare. <code>Send chunk 1 to server 47, chunk 2 to server 134, wait for responses, handle server failures, retry failed chunks, merge partial results</code>... Multiply that across thousands of machines and you get unmanageable complexity.</p><p>Instead, Google developed what became known as MapReduce; a declarative paradigm for distributed computing. Engineers could write: <code>Here's my map function (extract words from web pages). Here's my reduce function (count word frequencies). Process the entire web.</code> This framework handled all the imperative details: data distribution, failure recovery, load balancing, result aggregation. Engineers declared what they wanted computed. The system figured out how to coordinate thousands of servers.</p><p>This wasn't just elegant computer science. It was competitive advantage. While competitors struggled with complex distributed systems built on expensive hardware, Google's engineers focused on algorithms and data insights. Their declarative approach to distributed computing let them scale faster and cheaper than anyone thought possible.</p><p>What my friend was showing me in 1999, commodity hardware coordinated by smart software that abstracted away distributed complexity, was MapReduce in action, years before the famous 2004 paper. That paper didn't introduce a new concept; it documented the practices that had already powered Google's rise to dominance.</p>]]></content:encoded></item><item><title><![CDATA[The Beginning and End of LLM Workflow Software: How MCP Will Obsolesce Workflows]]></title><description><![CDATA[Twenty-one drag-and-drop steps or one sentence. Why the visual workflow builders are the last of their kind.]]></description><link>https://essays.xcud.com/p/the-beginning-and-end-of-llm-workflow-software-how-mcp-will-obsolesce-workflows</link><guid isPermaLink="false">https://essays.xcud.com/p/the-beginning-and-end-of-llm-workflow-software-how-mcp-will-obsolesce-workflows</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Mon, 19 May 2025 14:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ftuA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>The Beginning and End of LLM Workflow Software: How MCP Will Obsolesce Workflows</h1><p>In the rapidly evolving landscape of enterprise software, we're witnessing the meteoric rise of workflow automation tools. These platforms promise to streamline operations through visual interfaces where users can design, implement, and monitor complex business processes. Yet despite their current popularity, these GUI-based workflow solutions may represent the last generation of their kind&#8212;soon to be replaced by more versatile Large Language Model (LLM) interfaces.</p><h2>The Current Workflow Software Boom</h2><p>The workflow automation market is experiencing unprecedented growth, projected to reach 78.8 billion USD by 2030 with a staggering 23.1% compound annual growth rate. This explosive expansion is evident in both funding activity and market adoption: Workato secured a 200 million USD Series E round at a $5.7 billion valuation, while established players like ServiceNow and Appian continue to report record subscription revenues.</p><p>A quick glance at a typical workflow builder interface reveals the complexity these tools embrace:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ftuA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ftuA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png 424w, https://substackcdn.com/image/fetch/$s_!ftuA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png 848w, https://substackcdn.com/image/fetch/$s_!ftuA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png 1272w, https://substackcdn.com/image/fetch/$s_!ftuA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ftuA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png" width="1456" height="1116" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1116,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;alt text&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="alt text" title="alt text" srcset="https://substackcdn.com/image/fetch/$s_!ftuA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png 424w, https://substackcdn.com/image/fetch/$s_!ftuA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png 848w, https://substackcdn.com/image/fetch/$s_!ftuA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png 1272w, https://substackcdn.com/image/fetch/$s_!ftuA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48342eb5-2c46-47e7-bb9f-41b9885fed64_1637x1255.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The landscape is crowded with vendors aggressively competing for market share:</p><ul><li><p><strong>Enterprise platforms</strong>: ServiceNow, Pega, Appian, and IBM Process Automation dominate the high-end market, offering comprehensive solutions tightly integrated with their broader software ecosystems.</p></li><li><p><strong>Integration specialists</strong>: Workato, Tray.io, and Zapier focus specifically on connecting disparate applications through visual workflow builders, catering to the growing API economy.</p></li><li><p><strong>Emerging players</strong>: Newer entrants like Bardeen, n8n, and Make (formerly Integromat) are gaining traction with innovative approaches and specialized capabilities.</p></li></ul><p>This workflow automation boom follows a familiar pattern we've seen before. Between 2018 and 2022, Robotic Process Automation (RPA) experienced a similar explosive growth cycle. Companies like UiPath reached a peak valuation of $35 billion before a significant market correction as limitations became apparent. RPA promised to automate routine tasks by mimicking human interactions with existing interfaces&#8212;essentially screen scraping and macro recording at an enterprise scale&#8212;but struggled with brittle connections, high maintenance overhead, and limited adaptability to changing interfaces.</p><p>Today's workflow tools attempt to address these limitations by focusing on API connections rather than UI interactions, but they still follow the same fundamental paradigm: visual programming interfaces that require specialized knowledge to build and maintain.</p><p>So why are organizations pouring billions into these platforms despite the lessons from RPA? Several factors drive this investment:</p><ul><li><p><strong>Digital transformation imperatives</strong>: COVID-19 dramatically accelerated organizations' need to automate processes as remote work became essential and manual, paper-based workflows proved impossible to maintain.</p></li><li><p><strong>The automation gap</strong>: Companies recognize the potential of AI and automation but have lacked accessible tools to implement them across the organization without heavy IT involvement.</p></li><li><p><strong>Democratization promise</strong>: Workflow tools market themselves as empowering "citizen developers"&#8212;business users who can automate their own processes without coding knowledge.</p></li><li><p><strong>Pre-LLM capabilities</strong>: Until recently, organizations had few alternatives for process automation that didn't require extensive software development.</p></li></ul><p>What we're witnessing is essentially a technological stepping stone&#8212;organizations hungry for AI-powered results before true AI was ready to deliver them at scale. But as we'll see, that technological gap is rapidly closing, with profound implications for the workflow software category.</p><h2>Why LLMs Will Disrupt Workflow Software</h2><p>While current workflow tools represent incremental improvements on decades-old visual programming paradigms, LLMs offer a fundamentally different approach&#8212;one that aligns with how humans naturally express process logic and intent. The technical capabilities enabling this shift are advancing rapidly, creating the conditions for widespread disruption.</p><h3>The Technical Foundation: Resource Access Protocols</h3><p>The key technical enabler for LLM-driven workflows is the development of secure protocols that allow these models to access and manipulate resources. Model Context Protocol (MCP) represents one of the most promising approaches:</p><p>MCP provides a standardized way for LLMs to:</p><ul><li><p>Access data from various systems through controlled APIs</p></li><li><p>Execute actions with proper authentication and authorization</p></li><li><p>Maintain context across multiple interactions</p></li><li><p>Document actions taken for compliance and debugging</p></li></ul><p>Unlike earlier attempts at AI automation, MCP and similar protocols solve the "last mile" problem by creating secure bridges between conversational AI and the systems that need to be accessed or manipulated. Major cloud providers are already implementing variations of these protocols, with Microsoft's Azure AI Actions, Google's Gemini API, and Anthropic's Claude Tools representing early implementations.</p><p>The proliferation of these standards means that instead of building custom integrations for each workflow tool, organizations can create a single set of LLM-compatible APIs that work across any AI interface.</p><h3>Natural Language vs. GUI Interfaces</h3><p>The cognitive load difference between traditional workflow tools and LLM interfaces becomes apparent when comparing approaches to the same problem:</p><h4>Traditional Workflow Tool Process</h4><ol><li><p>Open workflow designer application</p></li><li><p>Create a new workflow and name it</p></li><li><p>Drag "Trigger" component (Customer Signup)</p></li><li><p>Configure webhook or database monitor</p></li><li><p>Drag "HTTP Request" component</p></li><li><p>Configure endpoint URL for credit API</p></li><li><p>Add authentication parameters (API key, tokens)</p></li><li><p>Add request body parameters and format</p></li><li><p>Connect to "JSON Parser" component</p></li><li><p>Define schema for response parsing</p></li><li><p>Create variable for credit score</p></li><li><p>Add "Decision" component</p></li><li><p>Configure condition (score &lt; 600)</p></li><li><p>For "True" path, add "Notification" component</p></li><li><p>Configure recipients, subject, and message template</p></li><li><p>Add error handling for API timeout</p></li><li><p>Add error handling for data format issues</p></li><li><p>Test with sample data</p></li><li><p>Debug connection issues</p></li><li><p>Deploy to production environment</p></li><li><p>Configure monitoring alerts</p></li></ol><h4>LLM Approach</h4><pre><code>When a new customer signs up, retrieve their credit score from our API, 
store it in our database, and if the score is below 600, notify the risk 
assessment team.
</code></pre><p>The workflow tool approach requires not only understanding the business logic but also learning the specific implementation patterns of the tool itself. Users must know which components to use, how to properly connect them, and how to configure each element&#8212;skills that rarely transfer between different workflow platforms.</p><h3>Dynamic Adaptation Through Conversation</h3><p>Real business processes rarely remain static. Consider how process changes propagate in each paradigm:</p><h4>Traditional Workflow Change Process</h4><ol><li><p>Open existing workflow in designer</p></li><li><p>Identify components that need modification</p></li><li><p>Add new components for bankruptcy check</p></li><li><p>Configure API connection to bankruptcy database</p></li><li><p>Add new decision branch</p></li><li><p>Connect positive result to new components</p></li><li><p>Add calendar integration component</p></li><li><p>Configure meeting details and attendees</p></li><li><p>Update documentation to reflect changes</p></li><li><p>Redeploy updated workflow</p></li><li><p>Test all paths, including existing functionality</p></li><li><p>Update monitoring for new failure points</p></li></ol><h4>LLM Approach</h4><pre><code>Actually, let's also check if they've had a bankruptcy in the last five 
years, and if so, automatically schedule a review call with our financial 
advisor team.
</code></pre><p>The LLM simply incorporates the new requirement conversationally. Behind the scenes, it maintains a complete understanding of the existing process and extends it appropriately&#8212;adding the necessary API calls, conditional logic, and scheduling actions without requiring the user to manipulate visual components.</p><p>Early implementations of this approach are already appearing. GitHub Copilot for Docs can update software configuration by conversing with developers about their intentions, rather than requiring them to parse documentation and make manual changes. Similarly, companies like Adept are building AI assistants that can operate existing software interfaces based on natural language instructions.</p><h3>Self-Healing Systems: The Maintenance Advantage</h3><p>Perhaps the most profound advantage of LLM-driven workflows is their ability to adapt to changing environments without breaking. Traditional workflows are notoriously brittle:</p><p>Traditional Workflow Failure Scenarios:</p><ul><li><p>An API endpoint changes its structure</p></li><li><p>A data source modifies its authentication requirements</p></li><li><p>A third-party service deprecates a feature</p></li><li><p>A database schema is updated</p></li><li><p>Operating system or runtime dependencies change</p></li></ul><p>When these changes occur, traditional workflows break and require manual intervention. Someone must diagnose the issue, understand the change, modify the workflow components, test the fixes, and redeploy. This maintenance overhead is substantial&#8212;studies suggest organizations spend 60-80% of their workflow automation resources on maintenance rather than creating new value.</p><p><strong>LLM-Driven Workflow Adaptation:</strong> LLMs with proper resource access can automatically adapt to many changes:</p><ul><li><p>When an API returns errors, the LLM can examine documentation, test alternative approaches, and adjust parameters</p></li><li><p>If authentication requirements change, the LLM can interpret error messages and modify its approach</p></li><li><p>When services deprecate features, the LLM can find and implement alternatives based on its understanding of the underlying intent</p></li><li><p>Changes in database schemas can be discovered and accommodated dynamically</p></li><li><p>Environmental changes can be detected and worked around</p></li></ul><p>Rather than breaking, LLM-driven workflows degrade gracefully and can often self-heal without human intervention. When they do require assistance, the interaction is conversational:</p><pre><code>User: The customer onboarding workflow seems to be failing at the credit check 
step.
LLM: I've investigated the issue. The credit API has changed its response 
format. I've updated the workflow to handle the new format. Would you like 
me to show you the specific changes I made?
</code></pre><p>This self-healing capacity drastically reduces maintenance overhead and increases system reliability. Organizations using early LLM-driven processes report up to 70% reductions in workflow maintenance time and significantly improved uptime.</p><h3>Compliance and Audit Superiority</h3><p>Perhaps counterintuitively, LLM-driven workflows can provide superior compliance capabilities. Several financial institutions are already piloting LLM systems that maintain comprehensive audit logs that surpass traditional workflow tools:</p><ul><li><p><strong>Granular Action Logging</strong>: Every step, decision point, and data access is logged with complete context</p></li><li><p><strong>Natural Language Explanations</strong>: Each action includes an explanation of why it was taken</p></li><li><p><strong>Cryptographic Verification</strong>: Logs can be cryptographically signed and verified for tamper detection</p></li><li><p><strong>Full Data Lineage</strong>: Complete tracking of where data originated and how it was transformed</p></li><li><p><strong>Semantic Search</strong>: Compliance teams can query logs using natural language questions</p></li></ul><p>A major U.S. bank recently compared their existing workflow tool's audit capabilities with a prototype LLM-driven system and found the LLM approach provided 3.5x more detailed audit information with 65% less storage requirements, due to the elimination of redundant metadata and more efficient logging.</p><h3>Visualization On Demand</h3><p>For scenarios where visual representation is beneficial, LLMs offer a significant advantage: contextually appropriate visualizations generated precisely when needed.</p><p>Rather than being limited to pre-designed dashboards and reports, users can request visualizations tailored to their current needs:</p><pre><code>User: Show me a diagram of how the customer onboarding process changes with 
the new bankruptcy check.

LLM: Generates a Mermaid diagram showing the modified process flow with the 
new condition highlighted

User: How will this affect our approval rates based on historical data?

LLM: Generates a bar chart showing projected approval rate changes based on 
historical bankruptcy data
</code></pre><p>Companies like Observable and Vercel are already building tools that integrate LLM-generated visualizations into business workflows, allowing users to create complex data visualizations through conversation rather than manual configuration.</p><h2>Current State of Adoption</h2><p>While the technical capabilities exist, we're still in the early stages of this transition. Rather than presenting hypothetical examples as established successes, it's more accurate to examine how organizations are currently experimenting with LLM-driven workflow approaches:</p><ul><li><p><strong>Prototype implementations</strong>: Several companies are building prototype systems that use LLMs to orchestrate workflows, but these remain largely experimental and haven't yet replaced enterprise-wide workflow systems.</p></li><li><p><strong>Augmentation rather than replacement</strong>: Most organizations are currently using LLMs to augment existing workflow tools&#8212;helping users configure complex components or troubleshoot issues&#8212;rather than replacing the tools entirely.</p></li><li><p><strong>Domain-specific applications</strong>: The most successful early implementations focus on narrow domains with well-defined processes, such as content approval workflows or customer support triage, rather than attempting to replace entire workflow platforms.</p></li><li><p><strong>Hybrid approaches</strong>: Organizations are finding success with approaches that combine traditional workflow engines with LLM interfaces, allowing users to interact conversationally while maintaining the robustness of established systems.</p></li></ul><p>While we don't yet have large-scale case studies with verified metrics showing complete workflow tool replacement, the technological trajectory is clear. As LLM capabilities continue to improve and resource access protocols mature, the barriers to adoption will steadily decrease.</p><h2>Investment Implications</h2><p>The disruption of workflow automation by LLMs isn't a gradual shift&#8212;it's happening now. For decision-makers, this isn't about careful transitions or hedged investments; it's about immediate and decisive action to avoid wasting resources on soon-to-be-obsolete technology.</p><h3>Halt Investment in Traditional Workflow Tools Immediately</h3><p>Stop signing or renewing licenses for traditional workflow automation platforms. These systems will be obsolete within weeks, not years. Any new investment in these platforms represents resources that could be better allocated to LLM+MCP approaches. If you've recently purchased licenses, investigate termination options or ways to repurpose these investments.</p><h3>Redirect Resources to LLM Infrastructure</h3><p>Immediately reallocate budgets from workflow software to: - Enterprise-grade LLM deployment on your infrastructure - Implementation of MCP or equivalent protocols - API development for all internal systems - Prompt engineering training for existing workflow specialists</p><h3>Install LLM+MCP on Every Desktop Now</h3><p>Rather than planning gradual rollouts, deploy LLM+MCP capabilities across your organization immediately. Every day that employees continue to build workflows in traditional tools is a day of wasted effort creating systems that will need to be replaced. Local or server-based LLMs with proper resource access should become standard tools alongside word processors and spreadsheets.</p><h3>Retrain Teams for the New Paradigm</h3><p>Your workflow specialists need to become prompt engineers&#8212;not next quarter, but this week: - Cancel scheduled workflow tool training - Replace with intensive prompt engineering workshops - Focus on teaching conversational process design rather than visual programming - Develop internal guides for effective LLM workflow creation</p><h3>Examine Legal Obligations</h3><p>For organizations with existing contracts for workflow platforms: - Review termination clauses and calculate the cost of early exits - Investigate whether remaining license terms can be applied to API access rather than visual workflow tools - Consider whether vendors might offer transitions to their own LLM offerings in lieu of contracted services</p><h3>Vendors: Pivot or Perish</h3><p>For workflow automation companies, there's no time for careful transitions: - Immediately halt development on visual workflow designers - Redirect all engineering resources to LLM interfaces and connectors - Open all APIs and create comprehensive documentation for LLM interaction - Develop prompt libraries that encapsulate existing workflow patterns</p><p>The AI-assisted development cycle is accelerating innovation at unprecedented rates. What would have taken years is now happening in weeks. Organizations that try to manage this as a gradual transition will find themselves outpaced by competitors who embrace the immediate shift to LLM-driven processes.</p><h3>Our Own Evolution</h3><p>We need to acknowledge our own journey in this space. At Lit.ai, we initially invested in building the Workflow Canvas - a visual tool for designing LLM-powered workflows that made the technology more accessible. We created this product with the belief that visual workflow builders would remain essential for orchestrating complex LLM interactions.</p><p>However, our direct experience with customers and the rapid evolution of LLM capabilities has caused us to reassess this position. The very technology we're building is becoming sophisticated enough to make our own workflow canvas increasingly unnecessary for many use cases. Rather than clinging to this approach, we're now investing heavily in Model Context Protocol (MCP) and direct LLM resource access.</p><p>This pivot represents our commitment to following the technology where it leads, even when that means disrupting our own offerings. We believe the most valuable contribution we can make isn't building better visual workflow tools, but rather developing the connective tissue that allows LLMs to directly access and manipulate the resources they need to execute workflows without intermediary interfaces.</p><p>Our journey mirrors what we expect to see across the industry - an initial investment in workflow tools as a stepping stone, followed by a recognition that the real value lies in direct LLM orchestration with proper resource access protocols.</p><h2>Timeline and Adoption Considerations</h2><p>While the technical capabilities enabling this shift are rapidly advancing, several factors will influence adoption timelines:</p><h3>Enterprise Inertia</h3><p>Large organizations with established workflow infrastructure and trained teams will transition more slowly. Expect these environments to adopt hybrid approaches initially, where LLMs complement rather than replace existing workflow tools.</p><h3>High-Stakes Domains</h3><p>Industries with mission-critical workflows (healthcare, finance, aerospace) will maintain traditional interfaces longer, particularly for processes with significant safety or regulatory implications. However, even in these domains, LLMs will gradually demonstrate their reliability for increasingly complex tasks.</p><h3>Security and Control Concerns</h3><p>Organizations will need to develop comfort with LLM-executed workflows, particularly regarding security, predictability, and control. Establishing appropriate guardrails and monitoring will be essential for building this confidence.</p><h2>Conclusion</h2><p>The current boom in workflow automation software represents the peak of a paradigm that's about to be disrupted. As LLMs gain direct access to resources and demonstrate their ability to understand and execute complex processes through natural language, the value of specialized GUI-based workflow tools will diminish.</p><p>Forward-thinking organizations should prepare for this shift by investing in API infrastructure, LLM integration capabilities, and domain-specific knowledge engineering rather than committing deeply to soon-to-be-legacy workflow platforms. The future of workflow automation isn't in better diagrams and drag-drop interfaces&#8212;it's in the natural language interaction between users and increasingly capable AI systems.</p><p>In fact, this very article demonstrates the principle in action. Rather than using a traditional publishing workflow tool with multiple steps and interfaces, it was originally drafted in Google Docs, then an LLM was instructed to:</p><pre><code>Translate this to markdown, save it to a file on the local disk, execute a 
build, then upload it to AWS S3.
</code></pre><p> The entire publishing workflow&#8212;format conversion, file system operations, build process execution, and cloud deployment&#8212;was accomplished through a simple natural language request to an LLM with the appropriate resource access, eliminating the need for specialized workflow interfaces.</p><p>This perspective challenges conventional wisdom about enterprise software evolution. Decision-makers who recognize this shift early will gain significant advantages in operational efficiency, technology investment, and organizational agility.</p>]]></content:encoded></item><item><title><![CDATA[The AI-Driven Transformation of Software Development]]></title><description><![CDATA[Written in February 2025: one engineer directing many agents, vendors becoming services, and build-versus-buy flipping. Kept here, unedited, for the timestamp.]]></description><link>https://essays.xcud.com/p/the-ai-driven-transformation-of-software-development</link><guid isPermaLink="false">https://essays.xcud.com/p/the-ai-driven-transformation-of-software-development</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Sat, 01 Feb 2025 15:00:00 GMT</pubDate><content:encoded><![CDATA[<h1>The AI-Driven Transformation of Software Development</h1><h2>The Seismic Shift in Software Development</h2><p>The software development landscape is undergoing a seismic shift, driven by the rapid advancement of artificial intelligence. This transformation transcends simple automation; it fundamentally alters how software is created, acquired, and utilized, leading to a re-evaluation of the traditional 'build versus buy' calculus. The pace of this transformation is likely to accelerate, making it crucial for businesses and individuals to stay adaptable and informed.</p><h2>The Rise of AI-Powered Development Tools</h2><p>For decades, the software industry has been shaped by a tension between bespoke, custom-built solutions and readily available commercial products. The complexity and cost associated with developing software tailored to specific needs often pushed businesses towards purchasing off-the-shelf solutions, even if those solutions weren't a perfect fit. This gave rise to the dominance of large software vendors and the Software-as-a-Service (SaaS) model. However, AI is poised to disrupt this paradigm.</p><h3>Introduction to AI-Powered Automation</h3><p>Large Language Models (LLMs) are revolutionizing software development by understanding natural language instructions and generating code snippets, functions, or even entire modules. Imagine describing a software feature in plain language and having an AI produce the initial code. Many are already using tools like ChatGPT in this way, coaching the AI, suggesting revisions, and identifying improvements before testing the output.</p><p>This is 'vibe coding,' where senior engineers guide LLMs with high-level intent rather than writing every line of code. While this provides a significant productivity boost&#8212;say, a 5x improvement&#8212;the true transformative potential lies in a one-to-many dynamic, where a single expert can exponentially amplify their impact by managing numerous AI agents simultaneously, each focused on different project aspects.</p><h3>Expanding AI Applications in Development</h3><p>Additionally, AI is being used for code review tools that can automatically identify potential issues and suggest improvements, and specific AI platforms offered by cloud providers like AWS CodeWhisperer and Google Cloud's AI Platform are providing comprehensive AI-driven development environments. AI is being used for AI-assisted testing and debugging, identifying potential bugs, suggesting fixes, and automating test cases.</p><h3>Composable Architectures and Orchestration</h3><p>Beyond code completion and generation, AI tools are also facilitating the development of reusable components and services. This move toward composable architectures allows developers to break down complex tasks into smaller, modular units. These units, powered by AI, can then be easily assembled and orchestrated to create larger applications, increasing efficiency and flexibility. Model Context Protocol (MCP) could play a role in standardizing the discovery and invocation of these services.</p><p>Furthermore, LLM workflow orchestration is also becoming more prevalent, where AI models can manage and coordinate the execution of these modular services. This allows for dynamic and adaptable workflows that can be quickly changed or updated as needed.</p><h3>Human Role and Importance</h3><p>However, it's crucial to recognize that AI is a tool. Humans will still be needed to guide its development, provide creative direction, and critically evaluate the AI-generated outputs. Human problem-solving skills and domain expertise remain essential for ensuring software quality and effectiveness.</p><h3>Impact on Productivity and Innovation</h3><p>These tools are not just incremental improvements; they have the potential to dramatically increase developer productivity, potentially enabling the same output with half the staff or even leading to a fivefold increase in efficiency in the near term, lower the barrier to entry for software creation, and enable the fast iteration of new features.</p><h3>Impact on Offshoring</h3><p>Furthermore, AI tools have the potential to level the playing field for offshore development teams. Traditionally, challenges such as time zone differences, communication barriers, and perceived differences in skill level have sometimes put offshore teams at a disadvantage. However, AI-powered development tools can mitigate these challenges:</p><ul><li><p><strong>Enhanced Productivity and Efficiency:</strong> AI tools can automate many tasks, allowing offshore teams to deliver faster and more efficiently, overcoming potential time zone delays.</p></li><li><p><strong>Improved Code Quality and Consistency:</strong> AI-assisted code generation, review, and testing tools can ensure high code quality and consistency, regardless of the team's location.</p></li><li><p><strong>Reduced Communication Barriers:</strong> AI-powered translation and documentation tools can facilitate clearer communication and knowledge sharing.</p></li><li><p><strong>Access to Cutting-Edge Technology:</strong> With cloud-based AI tools, offshore teams can access the same advanced technology as onshore teams, eliminating the need for expensive local infrastructure.</p></li><li><p><strong>Focus on Specialization:</strong> Offshore teams can specialize in specific AI-related tasks, such as AI model training, data annotation, or AI-driven testing, becoming highly competitive in these areas.</p></li></ul><p>By embracing AI tools, offshore teams can overcome traditional barriers and compete on an equal footing with onshore teams, offering high-quality software development services at potentially lower costs. This could lead to a more globalized and competitive software development landscape.</p><h2>The Explosion of New Software and Features</h2><p>This evolution is leading to an explosion of new software products and features. Individuals and small teams can now bring their ideas to life with unprecedented speed and efficiency. This is made possible by AI tools that can quickly translate high-level descriptions into working code, allowing for quicker prototyping and development cycles.</p><p>Crucial to the effectiveness of these AI tools is the quality of their training data. High-quality, diverse datasets enable AI models to generate more accurate and robust code. This is particularly impactful in niche markets, where highly specialized software solutions, previously uneconomical to develop, are now becoming viable.</p><p>For instance, AI could revolutionize enterprise applications with greater automation and integration capabilities, lead to more personalized and intuitive consumer apps, accelerate scientific discoveries by automating data analysis and simulations, or make embedded systems more intelligent and adaptable.</p><p>Furthermore, AI can analyze user data to identify areas for improvement and drive innovation, making software more responsive to user needs. While AI automates many tasks, human creativity and critical thinking are still vital for defining the vision and goals of software projects.</p><p>It's important to consider the potential environmental impact of this increased software development, including the energy consumption of training and running AI models. However, AI-driven software also offers opportunities for more efficient resource management and sustainability in other sectors, such as optimizing supply chains or reducing energy waste.</p><p>Software will evolve at an unprecedented pace, with AI facilitating fast feature iteration, updates, and highly personalized user experiences. This surge in productivity will likely lead to an explosion of new software products, features, and niche applications, democratizing software creation and lowering the barrier to entry.</p><h2>The Transformation of the Commercial Software Market</h2><p>This evolution is reshaping the commercial software market. The proliferation of high-quality, AI-enhanced open-source alternatives is putting significant pressure on proprietary vendors. As companies find they can achieve their software needs through internal development or by leveraging robust open-source solutions, they are becoming more price-sensitive and demanding greater value from commercial offerings.</p><p>This is forcing vendors to innovate not only in terms of features but also in their business models, with a greater emphasis on value-added services such as consulting, support, and integration expertise. Strategic partnerships and collaboration with open-source communities will also become crucial for commercial vendors to remain competitive.</p><p>Commercial software vendors will need to adapt to this shift by offering their functionalities as discoverable services via protocols like MCP. Instead of selling large, complex products, they might provide specialized services that can be easily integrated into other applications. This could lead to new business models centered around providing best-in-class, composable AI capabilities.</p><p>Specifically, this shift is leading to changes in priorities and value perceptions. Commercial software vendors will likely need to shift their focus towards providing value-added services such as consulting, support, and integration expertise as open-source alternatives become more competitive. Companies may place a greater emphasis on software that can be easily customized and integrated with their existing systems, potentially leading to a demand for more flexible and modular solutions.</p><p>Furthermore, commercial vendors may need to explore strategic partnerships and collaborations with open-source communities to remain competitive and utilize the collective intelligence of the open-source ecosystem.</p><p>Overall, AI-driven development has the potential to transform the software landscape, creating a more level playing field for open-source projects and putting significant pressure on the traditional commercial software market. Companies will likely need to adapt their strategies and offerings to remain competitive in this evolving environment.</p><h2>The Impact on the Open-Source Ecosystem</h2><p>The open-source ecosystem is experiencing a significant transformation driven by AI. AI-powered tools are not only lowering the barriers to contribution, making it easier for developers to participate and contribute, but they are also fundamentally changing the competitive landscape.</p><p>Specifically, AI fuels the creation of more robust, feature-rich, and well-maintained open-source software, making these projects even more viable alternatives to commercial offerings. Businesses, especially those sensitive to cost, will have more compelling free options to consider. This acceleration is leading to faster feature parity, where AI could enable open-source projects to rapidly catch up to or even surpass the feature sets of commercial software in certain domains, further reducing the perceived value proposition of paid solutions.</p><p>Moreover, the ability for companies to customize open-source software using AI tools could eliminate the need for costly customization services offered by commercial vendors, potentially resulting in customization at zero cost. The agility and flexibility of open-source development, aided by AI, enable quick innovation and experimentation, allowing companies to try new features and technologies more quickly and potentially reducing their reliance on proprietary software that might not be able to keep pace.</p><p>AI tools can also help expose open-source components as discoverable services, making them even more accessible and reusable. This can further accelerate the development and adoption of open-source software, as companies can easily integrate these services into their own applications.</p><p>Furthermore, the vibrant and collaborative nature of open-source communities, combined with AI tools, provides companies with access to a vast pool of expertise and support at no additional cost. This is accelerating the development cycle, improving code quality, and fostering an even more collaborative and innovative environment. As open-source projects become more mature and feature-rich, they present an increasingly compelling alternative to commercial software, further fueling the shift away from traditional proprietary solutions.</p><h2>The Changing "Build Versus Buy" Calculus</h2><p>Ultimately, the rise of AI in software development is driving a fundamental shift in the "build versus buy" calculus. The rise of composable architectures means that 'building' now often entails assembling and orchestrating existing services, rather than developing everything from scratch. This dramatically lowers the barrier to entry and makes building tailored solutions even more cost-effective.</p><p>Companies are finding that building their own tailored solutions, often on cloud infrastructure, is becoming increasingly cost-effective and strategically advantageous. The ability for companies to customize open-source software using AI could eliminate the need for costly customization services offered by commercial vendors.</p><p>Innovation and experimentation in open-source, aided by AI, could further reduce reliance on proprietary software. Robotic Process Automation (RPA) bots can also be exposed as services via MCP, allowing companies to integrate automated tasks into their workflows more easily. This further enhances the 'build' option, as businesses can employ pre-built RPA services to automate repetitive processes.</p><h2>Cloud vs. On-Premise: A Re-evaluation</h2><p>The potential for AI-driven, easier on-premise app development could indeed have significant implications for the cloud versus on-premise landscape, potentially leading to a shift in reliance on big cloud applications like Salesforce.</p><p>There's potential for reduced reliance on big cloud apps. If AI tools drastically simplify and accelerate the development of custom on-premise applications, companies that previously opted for cloud solutions due to the complexity and cost of in-house development might reconsider. They could build tailored solutions that precisely meet their unique needs without the ongoing subscription costs and potential vendor lock-in associated with large cloud platforms.</p><p>Furthermore, for organizations with strict data sovereignty requirements, regulatory constraints, or internal policies favoring on-premise control, the ability to easily build and maintain their own applications could be a major advantage. They could retain complete control over their data and infrastructure, addressing concerns that might have pushed them towards cloud solutions despite these preferences.</p><p>While cloud platforms offer extensive customization, truly bespoke requirements or deep integration with legacy on-premise systems can sometimes be challenging or costly to achieve. AI-powered development could empower companies to build on-premise applications that seamlessly integrate with their existing infrastructure and are precisely tailored to their workflows.</p><p>Composable architectures can also make on-premise development more manageable. Instead of building large, monolithic applications, companies can assemble smaller, more manageable services. This can reduce the complexity of on-premise development and make it a more viable option.</p><p>Additionally, while the initial investment in on-premise infrastructure and development might still be significant, the elimination of recurring subscription fees for large cloud platforms could lead to lower total cost of ownership (TCO) over the long term, especially for organizations with stable and predictable needs.</p><p>Finally, some organizations have security concerns related to storing sensitive data in the cloud, even with robust security measures in place. The ability to develop and host applications on their own infrastructure might offer a greater sense of control and potentially address these concerns, even if the actual security posture depends heavily on their internal capabilities.</p><p>However, several factors might limit the shift away from big cloud apps:</p><h3>The "As-a-Service" Value Proposition</h3><p>Cloud platforms like Salesforce offer more than just the application itself. They provide a comprehensive suite of services, including infrastructure management, scalability, security updates, platform maintenance, and often a rich ecosystem of integrations and third-party apps. Building and maintaining all of this in-house, even with AI assistance, could still be a significant undertaking.</p><p>Moreover, major cloud vendors invest heavily in research and development, constantly adding new features and capabilities, often leveraging cutting-edge AI themselves. This pace of innovation in the cloud might be difficult for on-premise development, even with AI tools, to keep pace with.</p><p>Cloud platforms are inherently designed for scalability and elasticity, allowing businesses to easily adjust resources based on demand. Replicating this level of flexibility on-premise can be complex and expensive. Many companies prefer to focus on their core business activities rather than managing IT infrastructure and application development, even if AI makes it easier; the "as-a-service" model offloads this burden.</p><p>Large cloud platforms often have vibrant ecosystems of developers, partners, and a wealth of documentation and community support. Building an equivalent internal ecosystem for on-premise development could be challenging. Some advanced features, particularly those leveraging large-scale data analytics and AI capabilities offered by the cloud providers themselves, might be difficult or impossible to replicate effectively on-premise.</p><p>Cloud providers might also shift towards offering more granular, composable services that can be easily integrated into various applications. This would allow companies to leverage the cloud's scalability and infrastructure while still maintaining flexibility and control over their applications.</p><p>Therefore, a more likely scenario might be the rise of hybrid approaches, where companies use AI to build custom on-premise applications for specific, sensitive, or highly customized needs, while still relying on cloud platforms for other functions like CRM, marketing automation, and general productivity tools.</p><p>While the advent of AI tools that simplify on-premise application development could certainly empower more companies to build their own solutions and potentially reduce their reliance on monolithic cloud applications like Salesforce, a complete exodus is unlikely. The value proposition of cloud platforms extends beyond just the software itself to encompass infrastructure management, scalability, innovation, and ecosystem.</p><p>Companies will likely weigh the benefits of greater control and customization offered by on-premise solutions against the convenience, scalability, and breadth of services provided by the cloud. We might see a more fragmented landscape where companies strategically choose the deployment model that best fits their specific needs and capabilities.</p><h2>The AI-Driven Software Revolution</h2><p>The integration of advanced AI into software development is poised to trigger a profound shift, fundamentally altering how software is created, acquired, and utilized. This shift is characterized by:</p><h3>1. Exponential Increase in Productivity and Innovation:</h3><p><strong>AI as a Force Multiplier:</strong> AI tools are drastically increasing developer productivity, potentially enabling the same output with half the staff or even leading to a fivefold increase in efficiency in the near term.</p><p><strong>Cambrian Explosion of Software:</strong> This surge in productivity will likely lead to an explosion of new software products, features, and niche applications, democratizing software creation and lowering the barrier to entry.</p><p><strong>Rapid Iteration and Personalization:</strong> Software will evolve at an unprecedented pace, with AI facilitating fast feature iteration, updates, and highly personalized user experiences. This will often involve complex LLM workflow orchestration to manage and coordinate the various AI-driven processes.</p><p>This impact will be felt across various types of software, from enterprise solutions to consumer apps, scientific tools, and embedded systems. The effectiveness of these AI tools relies heavily on the quality of their training data, and the ability to analyze user data will drive further innovation and personalization.</p><p>We must also consider the sustainability implications, including the energy consumption of AI models and the potential for AI-driven software to promote resource efficiency in other sectors. These changes are not static; they are part of a dynamic and rapidly evolving landscape. Tools like GitHub Copilot and AWS CodeWhisperer are already demonstrating the power of AI in modern development workflows.</p><h3>2. Transformation of the Software Development Landscape:</h3><p><strong>Evolving Roles:</strong> The traditional role of a "coder" will diminish, with remaining developers focusing on AI prompt engineering, system architecture, including the design and management of complex LLM workflow orchestration, integration, service orchestration, MCP management, quality assurance, and ethical considerations.</p><p>This shift is particularly evident in the rise of vibe coding. More significantly, we're moving towards a one-to-many model where a single subject matter expert (SME) or senior engineer will manage and direct many LLM coding agents, each working on different parts of a project. This orchestration of AI agents will dramatically amplify the impact of senior engineers, allowing them to oversee and guide complex projects with unprecedented efficiency.</p><p><strong>AI-Native Companies:</strong> New companies built around AI-driven development processes will emerge, potentially disrupting established software giants.</p><p><strong>Democratization of Creation:</strong> Individuals in non-technical roles will become "citizen developers," creating and customizing software with AI assistance.</p><h3>3. Broader Economic and Societal Impacts:</h3><p><strong>Automation Across Industries:</strong> The ease of creating custom software will accelerate automation in all sectors, leading to increased productivity but also potential job displacement.</p><p><strong>Lower Software Costs:</strong> Development cost reductions will translate to lower software prices, making powerful tools more accessible.</p><p><strong>New Business Models:</strong> New ways to monetize software will emerge, such as LLM features, data analytics, integration services, and specialized composable services offered via MCP.</p><p><strong>Workforce Transformation:</strong> Educational institutions will need to adapt to train a workforce for skills like AI ethics, prompt engineering, and high-level system design.</p><p><strong>Ethical and Security Concerns:</strong> Increased reliance on AI raises ethical concerns about bias, privacy, and security vulnerabilities. This includes the challenges of handling sensitive data when using AI tools.</p><h3>4. Implications for Purchasing Software Today:</h3><p><strong>Short-Term vs. Long-Term:</strong> Businesses must balance immediate needs with the potential for cheaper and better AI-driven alternatives in the future.</p><p><strong>Flexibility and Scalability:</strong> Prioritizing flexible, scalable, and cloud-based solutions is crucial.</p><p><strong>Avoiding Lock-In:</strong> Companies should be cautious about long-term contracts and proprietary solutions that might become outdated quickly.</p><h3>5. Google Firebase Studio as an Example:</h3><p><strong>AI-Powered Development:</strong> Firebase Studio's integration of Gemini and AI agents for prototyping, feature development, and code assistance exemplifies the trend towards AI-driven development environments.</p><p><strong>Rapid Prototyping and Iteration:</strong> The ability to create functional prototypes from prompts and iterate quickly with AI support validates the potential for an explosion of new software offerings.</p><p>In essence, the AI-driven software revolution represents a fundamental shift in the "build versus buy" calculus, empowering businesses and individuals to create tailored solutions more efficiently and affordably. While challenges exist, the long-term trend points towards a more open, flexible, and dynamic software ecosystem. It's important to remember that AI is a tool that amplifies human capabilities, and human ingenuity will remain at the core of software innovation.</p><h2>A More Open and Dynamic Software Ecosystem</h2><p>In conclusion, the advancements in AI are ushering in an era of unprecedented change in software development. This transformation promises to democratize software creation, accelerate innovation, and empower businesses to build highly customized solutions. While challenges remain, the long-term trend suggests a move towards a more open, composable, flexible, and user-centric software ecosystem, increasingly driven by discoverable services. Furthermore, the pace of these changes is likely to accelerate, making adaptability and continuous learning crucial for both businesses and individuals.</p>]]></content:encoded></item><item><title><![CDATA[Form, Storm, Norm, Perform: Team Building Philosophy]]></title><description><![CDATA[What a trip to our team in China taught me about Tuckman's four stages, and why the model holds from a two-person startup to a thousand-engineer organization.]]></description><link>https://essays.xcud.com/p/form-storm-norm-perform-team-building-philosophy</link><guid isPermaLink="false">https://essays.xcud.com/p/form-storm-norm-perform-team-building-philosophy</guid><dc:creator><![CDATA[Ben Vierck]]></dc:creator><pubDate>Mon, 26 Jan 2015 15:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!TwrX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Form, Storm, Norm, Perform: Team Building Philosophy</h1><p>I traveled to China this past fall to meet and greet the excellent students participating in our post-grad work program at Nankai University and to give a guest lecture on the history of Software Engineering. My true objective for this trip was to make a first-hand assessment of how our China team is performing as a member of the whole global software engineering organization.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TwrX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TwrX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png 424w, https://substackcdn.com/image/fetch/$s_!TwrX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png 848w, https://substackcdn.com/image/fetch/$s_!TwrX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png 1272w, https://substackcdn.com/image/fetch/$s_!TwrX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TwrX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png" width="640" height="360" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:360,&quot;width&quot;:640,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TwrX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png 424w, https://substackcdn.com/image/fetch/$s_!TwrX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png 848w, https://substackcdn.com/image/fetch/$s_!TwrX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png 1272w, https://substackcdn.com/image/fetch/$s_!TwrX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c31ca0d-5a28-4b33-b812-e903ebd2ac3b_640x360.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>My philosophy on good team building is the same whether it's a 2-person startup in your basement or a global software organization of 1000s. It's based on <a href="http://en.wikipedia.org/wiki/Tuckman%27s_stages_of_group_development">Tuckman's stages of group development</a>, a model developed in 1965 which posits the idea that teams are always in one of four distinct phases:</p><ol><li><p>Form</p></li><li><p>Storm</p></li><li><p>Norm</p></li><li><p>Perform</p></li></ol><p>The mistake that team managers make most often is to fail to recognize which phase their team is in. They waste time and effort focused on tasks that don't help their teams evolve to the next stage. For example, ==teams that don't have the right people can't become normative no matter how much effort a manager puts into it. The result of such an effort will be frustration for both the manager and the team members==. Recognize which phase your team is in and pour all of your effort into leading your team to the next level.</p><p><strong>Form</strong> is simple. Fire and hire well. Cut players that aren't contributing to the team. Cutting is the worst part of the job but it's the most important. Recruit the very best talent you can attract. You'll be surprised how talent attracts talent like a magnet. If you find yourself <em>stuck</em> in Form then odds are you're the problem. You're either unqualified to recognize and hire talent or you're unwilling to fire the dead weight. Fire yourself immediately and choose someone more competent to do the job for you.</p><p><strong>Storm</strong> takes care of itself if you've given Form the attention it deserves. There are natural leaders in any tribe. They want your job. Everyone else in the tribe is hell-bent on accomplishing your grand mission. Nurture both of these drives. If you don't have natural leaders you're still in form. If you don't have a grand mission that you can speak passionately about then stop what you're doing right now and go find one. It's your job! If your team isn't passionate about what they're doing it's because <em>you're</em> not passionate about what you're doing. Fire yourself immediately and choose someone more competent to do the job for you.</p><p><strong>Norm</strong> means normative. Getting normative means aligning everyone's efforts such that your team as a whole makes <em>measurable iterative progress</em> towards a common goal. It requires process, cultivating a culture of growth, and leadership. Don't know where to start? Start with these three psychological needs: autonomy, relatedness, and competence. <a href="https://twitter.com/fowlersusann">Susan Fowler</a> has a great run-down of these needs in her latest article, <a href="https://hbr.org/2014/11/what-maslows-hierarchy-wont-tell-you-about-motivation">What Maslow's Hierarchy Won't Tell You About Motivation</a>. If autonomy is counter to your management style then it's time to consider firing yourself to make room for someone that can take your team to the next level.</p><p><strong>Perform</strong> is a unicorn. The advice I follow (remember, it's only worth what you paid for it) is to:</p><ol><li><p>Have a little luck.</p></li><li><p>Persevere through the grind. No great thing accomplished was ever easy.</p></li><li><p>Lead from the front. Participate in the process of accomplishing the grand vision. If you're not competent enough to contribute to the project then it's high time you fire yourself and hire someone who is.</p></li></ol><p>Enjoy your time in Perform. It's fleeting. You'll spend the rest of your career chasing it.</p><p><strong>Adjourn</strong> is the rarely talked about unofficial fifth phase. It reflects completing the task and breaking up the team. Teams that I've been lucky enough to participate in at the Perform phase never stay together for more than a single cycle. ==This is why it's vital for companies to hire management not based on how well they run a team, but rather how well they build a team.==</p><p>This is a topic that deserves the attention of a good practical reference book. Know of any? Leave a comment.</p><p><em>(Special thanks to my long-time friend <a href="https://twitter.com/cetezadi">Cameron</a> for turning me on to Bruce Tuckman years ago when I first made the move to management.)</em></p>]]></content:encoded></item></channel></rss>