<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <title>yous.dev — field notes</title>
  <subtitle>Agentic AI, AI security, tools, tutorials and talks — by José Antonio Cordón Muñoz (yous).</subtitle>
  <link href="https://yous.dev/feed.xml" rel="self"/>
  <link href="https://yous.dev/blog/"/>
  <id>https://yous.dev/blog/</id>
  <updated>2026-10-04T00:00:00Z</updated>
  <author><name>José Antonio Cordón Muñoz (yous)</name><uri>https://yous.dev/</uri></author>
  <icon>https://yous.dev/icon.png</icon>
  <entry>
    <title>Experimenting with Apple's Foundation Models</title>
    <link href="https://yous.dev/blog/experimenting-with-apple-foundation-models"/>
    <id>https://yous.dev/blog/experimenting-with-apple-foundation-models</id>
    <published>2026-10-04T00:00:00Z</published>
    <updated>2026-10-04T00:00:00Z</updated>
    <category term="tools"/>
    <summary>Why pay a subscription for AI notes when your phone already has the models? A Shortcut that turns any recording into a structured note with a diagram — on-device, on Private Cloud Compute, or with local open models.</summary>
    <content type="html">&lt;p class="post-sub"&gt;There's a whole category of gadgets that record your meetings and charge you a subscription to turn them into notes. I wanted to know whether the phone I already carry could do the same job with the language models Apple ships inside it. It can — and it ended up as a Shortcut that turns any recording into a structured note, diagram included, without an account, a subscription or another device.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/foundation-models-cover.jpg" alt="An iPhone on a desk: a sound wave flows into it and comes out the other side as note cards, a checklist and a diagram" loading="lazy" width="1600" height="900"&gt;
&lt;figcaption&gt;A recording in, a note out — all on the phone.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The idea started with the &lt;a href="https://www.plaud.ai/" target="_blank" rel="noopener"&gt;Plaud Note Pro&lt;/a&gt;: a slim recorder that clips to your phone and promises AI notes from everything you record. The hardware costs close to €200. Then, past a free tier of minutes, the notes themselves come with a &lt;a href="https://eu.plaud.ai/collections/plaud-ai-membership" target="_blank" rel="noopener"&gt;membership&lt;/a&gt; — around a hundred a year for the Pro plan, more for unlimited. I couldn't get past it. The phone in your pocket already has a good microphone, already transcribes speech, and since iOS 26 it has a language model you can call from anywhere. Why would anyone pay twice for something they already own?&lt;/p&gt;
&lt;p&gt;So I started tinkering. What came out is the most useful experiment I've done with small language models so far — not because it's clever, but because it's the kind of thing you end up using every day.&lt;/p&gt;
&lt;h2&gt;What Apple actually ships&lt;/h2&gt;
&lt;p&gt;The &lt;a href="https://developer.apple.com/documentation/foundationmodels" target="_blank" rel="noopener"&gt;Foundation Models framework&lt;/a&gt; opens Apple Intelligence's models to developers: a compact model of roughly three billion parameters that runs on the device itself, and a larger one that runs on &lt;a href="https://security.apple.com/blog/private-cloud-compute/" target="_blank" rel="noopener"&gt;Private Cloud Compute&lt;/a&gt;, Apple's servers built so that your request isn't kept or visible to anyone, Apple included. Apps get structured output, tool calling and streaming. And — the part that matters here — Shortcuts got a &lt;a href="https://www.macstories.net/notes/i-have-many-questions-about-apples-updated-foundation-models-and-the-great-use-model-action-in-shortcuts/" target="_blank" rel="noopener"&gt;Use Model action&lt;/a&gt; that lets anyone send a prompt to the on-device model, the cloud one, or ChatGPT, and feed the answer into the rest of the shortcut.&lt;/p&gt;
&lt;p&gt;It's worth being clear about what these models are not. They're small, their knowledge has an old cut-off, and their context window is short. Ask one about the world and you'll be disappointed. But that's the wrong job. A small model isn't an oracle; it's a &lt;em&gt;text-transformation engine&lt;/em&gt;: summarize this, restructure that, pull out the decisions, the dates, the tasks. Which is exactly what a set of notes is.&lt;/p&gt;
&lt;h2&gt;El Apuntador Pro&lt;/h2&gt;
&lt;p&gt;The shortcut is called &lt;strong&gt;&lt;a href="https://atajos.yous.dev/atajos/apuntador-pro/" target="_blank" rel="noopener"&gt;El Apuntador Pro&lt;/a&gt;&lt;/strong&gt; — roughly, "the note-taker" — and the flow is simple from the outside:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Ingest&lt;/strong&gt; a recording — record on the spot, or hand it an existing audio file.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Transcribe&lt;/strong&gt; it on the device.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pick a mode&lt;/strong&gt;: a class, a meeting, a brainstorm, or a quick summary. Each one has its own prompt and its own shape of note — a class wants concepts and definitions, a meeting wants decisions and owners, a brainstorm wants the ideas grouped, not flattened.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Generate&lt;/strong&gt; the note: a title, a TL;DR, the key points, the action items, and a Mermaid diagram that draws the structure of what was said.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Render and save&lt;/strong&gt;: the diagram becomes an image, and everything lands as a new note in Apple Notes.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;It launches from wherever you happen to be: the Action button, Siri, a widget or Control Center. Drawn as the kind of diagram it produces, the pipeline looks like this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;flowchart LR
A[Recording] --&amp;gt; B[Transcript]
B --&amp;gt; C{Mode}
C --&amp;gt;|class · meeting · brainstorm · quick| D[Prompt]
D --&amp;gt; E[Model]
E --&amp;gt; F[Sections + Mermaid]
F --&amp;gt; G[Note in Apple Notes]&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;From the inside, it's a different story. Most of the shortcut isn't the model call — it's everything around it: cleaning the transcript, a prompt per mode, parsing what comes back into sections, separating the diagram code from the prose, rendering it, and assembling a note that reads well on any device. That's where the real work went, and it's also where you learn what a small model needs. Narrow instructions and an explicit output format, and it behaves remarkably well; something open-ended, and it wanders.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/apuntador-pro-note.jpg" alt="A note in Apple Notes generated by El Apuntador Pro in class mode from a YouTube tutorial on extracting audio with VLC: a glossary, detailed step-by-step notes and a rendered Mermaid flowchart" loading="lazy" width="1600" height="964"&gt;
&lt;figcaption&gt;Class mode, from a short YouTube tutorial: a glossary, the detailed notes and the diagram — untouched.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2&gt;One shortcut, three kinds of model&lt;/h2&gt;
&lt;p&gt;What surprised me most is how interchangeable the brain turned out to be. The same shortcut runs on three very different engines:&lt;/p&gt;
&lt;div class="table-wrap"&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Engine&lt;/th&gt;&lt;th&gt;Where it runs&lt;/th&gt;&lt;th&gt;When I use it&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Apple on-device&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;On the phone, offline&lt;/td&gt;&lt;td&gt;Quick summaries, anything private, anywhere with no signal. Fast, free, always there.&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Private Cloud Compute&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Apple's private servers&lt;/td&gt;&lt;td&gt;The bigger model, for longer recordings and richer notes — still private.&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Open models, local&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;On the phone, through &lt;a href="https://apps.apple.com/app/locally-ai-local-ai-chat/id6741426692" target="_blank" rel="noopener"&gt;Locally&lt;/a&gt;&lt;/td&gt;&lt;td&gt;Comparing models, or when I want something other than Apple's — still without leaving the device.&lt;/td&gt;&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;/div&gt;
&lt;p&gt;That third row is the one I didn't expect to work as well as it does. Locally runs open models on Apple silicon with MLX and exposes them to Shortcuts, so swapping the brain is a matter of changing one action. On the iPhone 18 Pro, with its 12 GB of RAM, I've had good results with &lt;a href="https://prismml.com/news/bonsai-27b" target="_blank" rel="noopener"&gt;Bonsai 27B&lt;/a&gt; — PrismML's extreme quantization of a 27-billion-parameter model down to under 4 GB — as well as Qwen 3.5 4B and Gemma 4 E4B. A 27B-class model writing meeting notes in your pocket, offline, was not on my list of things I'd see this year.&lt;/p&gt;
&lt;p&gt;A note on privacy, because it's half the point. The audio and the transcript never leave the phone with the on-device or local models, and with Private Cloud Compute they go only to Apple's private servers. The one thing that goes out is the diagram's code — a few lines of text — to &lt;a href="https://mermaid.ink/" target="_blank" rel="noopener"&gt;mermaid.ink&lt;/a&gt;, which turns it into an image so the note looks the same everywhere. The note itself lives in Notes, like any other.&lt;/p&gt;
&lt;h2&gt;What the experiment taught me&lt;/h2&gt;
&lt;p&gt;I've written before about &lt;a href="https://yous.dev/blog/deploy-huggingface-models-with-ollama"&gt;running models on your own hardware&lt;/a&gt;, and my master's thesis was about small, fine-tuned models beating much bigger ones on narrow tasks. This was the same lesson from the other side: the most useful model isn't the biggest one, it's the one that's already there. No subscription, no login, no extra gadget to charge, no meter running per minute. It works on a train, in a lecture hall with no Wi-Fi, in a meeting where nothing should leave the room.&lt;/p&gt;
&lt;p&gt;Small language models stopped being a curiosity the moment they shipped inside the phone. The interesting question now isn't whether they're good enough — for transforming text, they are — but what we build around them. A recorder with a subscription is one answer. A shortcut is another.&lt;/p&gt;
&lt;p&gt;El Apuntador Pro is free, along with the rest of the shortcuts I make, at &lt;strong&gt;&lt;a href="https://atajos.yous.dev/" target="_blank" rel="noopener"&gt;atajos.yous.dev&lt;/a&gt;&lt;/strong&gt;. It needs iOS or iPadOS 26 and a device with Apple Intelligence turned on.&lt;/p&gt;
&lt;p&gt;Shortcut: &lt;strong&gt;&lt;a href="https://atajos.yous.dev/atajos/apuntador-pro/" target="_blank" rel="noopener"&gt;atajos.yous.dev/atajos/apuntador-pro&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>Read more, not scroll more</title>
    <link href="https://yous.dev/blog/read-more-not-scroll-more"/>
    <id>https://yous.dev/blog/read-more-not-scroll-more</id>
    <published>2026-09-01T00:00:00Z</published>
    <updated>2026-09-01T00:00:00Z</updated>
    <category term="misc"/>
    <summary>ReadIt is the reading app that wants you off your phone: real shelves, reading in company without spoilers, a candle instead of the scroll, and a whole reading life in one place. No ads, no trackers, no Amazon.</summary>
    <content type="html">&lt;p class="post-sub"&gt;I love reading, and I'm tired of apps that would rather I scrolled. So I built &lt;strong&gt;ReadIt&lt;/strong&gt; — a reading app, in Spanish and made with love, whose whole reason to exist is to get you off the phone and back inside a book. If you read, this one is for you.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/readit-cover.jpg" alt="A minimal illustration of a person reading by an arched window, a lit candle and a small stack of books beside them" loading="lazy" width="1536" height="1024"&gt;
&lt;figcaption&gt;A candle, a book, and nothing else asking for your attention.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Somewhere along the way, the apps meant to help us read turned into places to &lt;em&gt;not&lt;/em&gt; read. You open one to log a book, and twenty minutes later you're deep in a feed — a stranger's status, a hot take, an ad for a box set you'll never buy. The reading app became one more thing standing between you and the reading. And the one everyone lands on, Goodreads, is owned by Amazon, tuned for time-in-app and data, and it buries the single number I most want to see: how many people gave the book up.&lt;/p&gt;
&lt;p&gt;So I built the one I wished existed: &lt;strong&gt;&lt;a href="https://readit.es" target="_blank" rel="noopener"&gt;ReadIt&lt;/a&gt;&lt;/strong&gt; — &lt;em&gt;tu biblioteca en el bolsillo&lt;/em&gt;, your library in your pocket. Mobile-first, installable as an app, and built on one stubborn principle: the point is to read more, not to keep you here. &lt;em&gt;Hecho con cariño para quien lee&lt;/em&gt; — made with care, for people who read.&lt;/p&gt;
&lt;h2&gt;A shelf, not a list&lt;/h2&gt;
&lt;p&gt;Your library isn't a list — it's a wall of shelves with spines you pick up and move with a finger. You decorate it: put a cat on it, change the wallpaper, make it yours. There are four shelves to start — Reading, Read, Wishlist, and &lt;strong&gt;Gave up&lt;/strong&gt; — and you invent the rest. Yes, "Gave up" is a shelf, on purpose: abandoning a book also says something about who you are. And if you need one called "changed my life," you just make it.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/readit-shelf.jpg" alt="A ReadIt bookshelf, the 'Salón principal', with rows of book spines you flip to reveal each cover" loading="lazy" width="854" height="560"&gt;
&lt;figcaption&gt;A real shelf — spines you turn to see the cover, on a screen.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2&gt;Read in company, without spoilers&lt;/h2&gt;
&lt;p&gt;This is my favourite part. Walk into another reader's open library, pick up their copy of a book, and you'll see their highlights and the notes they left in the margins — the same book, a little different in every home. And when you read your own, you don't read alone: you see what other people underlined &lt;em&gt;up to where you are&lt;/em&gt;, and everything further on stays sealed until you reach it. Company without spoilers — a quiet book club around every page.&lt;/p&gt;
&lt;h2&gt;The candle, not the scroll&lt;/h2&gt;
&lt;p&gt;There is no infinite scroll here. There's a &lt;strong&gt;candle&lt;/strong&gt;. You light it and it burns down while you read — a warm little session of twenty-five minutes with nothing moving on the screen. And when you've been away from your books too long, ReadIt does the opposite of every other app: it nudges you to put the phone down. Its notifications say things like &lt;em&gt;"the algorithm has had enough of you today — your book doesn't track you,"&lt;/em&gt; or &lt;em&gt;"nobody is going to like this, and that's exactly the point."&lt;/em&gt; It is the rare app whose best possible outcome is that you close it.&lt;/p&gt;
&lt;h2&gt;A whole reading life, in one place&lt;/h2&gt;
&lt;p&gt;Around all of that is everything else a reader's life has, built for readers and no one else:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Book clubs&lt;/strong&gt; with a shared reading checkpoint, so no one spoils what lies ahead.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Events and a calendar&lt;/strong&gt; — readings, meet-ups, launches.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Challenges and badges&lt;/strong&gt; you earn by reading, not by opening the app — including the one I'm fondest of, the &lt;em&gt;Candle of 25&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Streaks&lt;/strong&gt; gentle enough to actually keep, and a reader "type" that slowly takes shape from what you read and underline.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;marketplace&lt;/strong&gt; to buy, sell and swap physical books with other readers, and a way to see what's being read near you.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Search by a line&lt;/strong&gt; — find a book by a sentence someone underlined.&lt;/li&gt;
&lt;li&gt;One-click import of your &lt;strong&gt;Goodreads and Kindle&lt;/strong&gt; highlights, and a small blog for the longer thoughts.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;What it refuses to be&lt;/h2&gt;
&lt;p&gt;No ads. No trackers — &lt;em&gt;una cookie de sesión y ya&lt;/em&gt;, one session cookie and nothing more. No Amazon behind it. It isn't a social network with books bolted on; the people are there to help you find the next great read, not to farm your attention. Or, as the site itself puts it: &lt;em&gt;leer más, no hacer amigos&lt;/em&gt; — read more, not make friends.&lt;/p&gt;
&lt;h2&gt;Come in early&lt;/h2&gt;
&lt;p&gt;ReadIt is invitation-only for now, while it finds its feet. But if you've read this far, you're exactly the reader it was built for — so here's a door in. The first &lt;strong&gt;20 people&lt;/strong&gt; to use this code get early access:&lt;/p&gt;
&lt;blockquote&gt;Invite code: &lt;strong&gt;&lt;code&gt;4HBK-WMQH&lt;/code&gt;&lt;/strong&gt; — 20 uses, first come, first served.&lt;/blockquote&gt;
&lt;p&gt;Redeem it at &lt;strong&gt;&lt;a href="https://readit.es" target="_blank" rel="noopener"&gt;readit.es&lt;/a&gt;&lt;/strong&gt;, bring three books, and your shelf will already start to look like you.&lt;/p&gt;
&lt;p&gt;I built ReadIt because reading is the quietest, most stubborn antidote to the feed I know, and the tools around it had forgotten that. If you love reading, this is the shelf I made for you.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>How long can I walk away?</title>
    <link href="https://yous.dev/blog/how-long-can-i-walk-away"/>
    <id>https://yous.dev/blog/how-long-can-i-walk-away</id>
    <published>2026-08-12T00:00:00Z</published>
    <updated>2026-08-12T00:00:00Z</updated>
    <category term="agentic ai"/>
    <summary>claudestimate gives Claude Code an honest, self-correcting ETA — "away until 16:45" — by splitting the plan from the clock, so you can spend the wait on the work that actually needs you.</summary>
    <content type="html">&lt;p class="post-sub"&gt;When you hand a long task to a coding agent, the honest question isn't "is it working?" — it's "how long have I got?" This is a small skill that gives Claude Code a live, self-correcting ETA — &lt;em&gt;away until 16:45&lt;/em&gt; — so you can spend the wait on the work that actually needs you instead of watching the screen. It pings you only when that time moves.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/claudestimate-cover.jpg" alt="A laptop showing claudestimate's ETA — away until 16:45 — beside a coffee and a notebook that reads 'Next problem.'" loading="lazy" width="1600" height="878"&gt;
&lt;figcaption&gt;Away until 16:45 — long enough to start the next problem.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The more you let a coding agent run on its own — &lt;a href="https://yous.dev/blog/prompts-became-loops"&gt;a loop that keeps going&lt;/a&gt; until the goal is met — the more you find yourself with time on your hands while it works. The worst thing to do with that time is spend it watching the screen. &lt;a href="https://yous.dev/blog/ai-writes-the-code-who-steers-the-system"&gt;Implementation stopped being the hard part&lt;/a&gt;; sitting there supervising the typing is dead time. The point of handing the work off is to &lt;em&gt;parallelize&lt;/em&gt; — to pick up the part that actually needs you (the next problem, the design, the framing, another project entirely) while the machine handles the part that doesn't.&lt;/p&gt;
&lt;p&gt;Which turns one small question into the most useful one you can ask: &lt;em&gt;how long?&lt;/em&gt; Not mainly so I can take a break — though you can — but so I know how big a block of real work I can start before I need to be back. Knowing the finish time lets you plan your own hours as deliberately as you planned the agent's, instead of burning them idling in front of a progress log.&lt;/p&gt;
&lt;p&gt;The trouble is that Claude can't tell you. Ask an agent "how long will this take?" and you get a number pulled from thin air — and worse, it has no idea how much time has actually gone by while it works. A language model has no clock.&lt;/p&gt;
&lt;p&gt;So I built a small skill to give it one, honestly: &lt;strong&gt;&lt;a href="https://github.com/jcordon5/claudestimate" target="_blank" rel="noopener"&gt;claudestimate&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;
&lt;h2&gt;Split the plan from the clock&lt;/h2&gt;
&lt;p&gt;The trick is a division of labour, and it's the whole reason the number is trustworthy instead of invented.&lt;/p&gt;
&lt;div class="table-wrap"&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Who&lt;/th&gt;&lt;th&gt;Provides&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Claude&lt;/strong&gt; (the skill)&lt;/td&gt;&lt;td&gt;The &lt;strong&gt;plan&lt;/strong&gt; — steps, sub-agents, estimated minutes — and it re-plans when it discovers work: a red test, a bug, a deploy slower than expected. It never writes a real timestamp; that's the one rule.&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;The engine&lt;/strong&gt; (bash + awk)&lt;/td&gt;&lt;td&gt;The &lt;strong&gt;clock&lt;/strong&gt; — it stamps how long each step actually took, recalibrates the estimate from what already happened, and decides when to alert you.&lt;/td&gt;&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;/div&gt;
&lt;p&gt;Claude does what it's good at — decomposing the task and judging effort — and stays out of what it can't do, which is measure time. The result isn't a stopwatch; it's a &lt;strong&gt;self-correcting forecast&lt;/strong&gt;. The first estimate is rough, and by the second or third step it has converged on real data.&lt;/p&gt;
&lt;p&gt;And the unit is deliberately not "22 minutes left" — a countdown you'd have to keep re-checking — but &lt;strong&gt;"away until 16:45"&lt;/strong&gt;, a clock time you can plan around and then forget.&lt;/p&gt;
&lt;h2&gt;Where you see it&lt;/h2&gt;
&lt;p&gt;Two places, no ceremony. A &lt;strong&gt;statusline&lt;/strong&gt; in Claude Code, always in view:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;⏱ 22m → 16:45  ▓▓▓░░░ 45%  ·  Implement callbacks (6/12m)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And a &lt;strong&gt;web dashboard&lt;/strong&gt; with the progress bar, the step timeline, every time the estimate was revised, and a task history — so that after the fact you can see whether the forecast was worth trusting. It's configurable from one panel: English or Spanish, alert thresholds, update cadence, and the model.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/claudestimate-dashboard.svg" alt="The claudestimate web dashboard: an 'away until 16:45' hero, the step timeline, estimate revisions, and a recent-tasks history" loading="lazy" width="840" height="560"&gt;
&lt;figcaption&gt;The dashboard: the finish time, the steps, and every time the estimate moved.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Because you're going to be away from the keyboard anyway, it closes the loop the way &lt;a href="https://yous.dev/blog/talk-to-your-agents-from-telegram"&gt;telegram-bridge&lt;/a&gt; does: point it at a Telegram bot and it &lt;strong&gt;pings you only when the finish time actually moves&lt;/strong&gt; — come back sooner if it sped up, extend your break if it slipped. Not a notification every thirty seconds; one message when the plan changes.&lt;/p&gt;
&lt;h2&gt;Why "honest" is the whole point&lt;/h2&gt;
&lt;p&gt;The easy version of this tool would print a confident number and be wrong. Two things keep it from that.&lt;/p&gt;
&lt;p&gt;First, it estimates in &lt;strong&gt;Claude's wall-clock, not human effort&lt;/strong&gt;. Early on I estimated each step in "how long would a person take," and the ETA overshot by roughly 2× — a five-step toy project I instrumented came in at 6.9 real minutes against a 14-minute guess. Estimating in the agent's own pace fixed it.&lt;/p&gt;
&lt;p&gt;Second, it &lt;strong&gt;calibrates per step-type&lt;/strong&gt;. Exploring, editing code, running tests and waiting on a deploy have very different rhythms, and a deploy shouldn't be re-timed against how fast code gets written. With that split, the "away until" time converges to under a fifth of a minute of error after two of five steps — once you're about 40% in, you know the finish to the quarter-minute. It's model-aware, too: it detects whether you're on Opus or Sonnet — Opus reasons more, so it's slower — and seeds the first guess accordingly, before real data takes over.&lt;/p&gt;
&lt;p&gt;When a task ends, it prints the drift between the initial estimate and reality, in minutes and percent, so you can decide whether the forecast earns its place:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;── Estimate deviation ──
Task:      Add OAuth login and deploy
Estimated: 20 min   ·   Real: 18.4 min
Deviation: -1.6 min (-8%)   Very accurate ✓&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;A 20-minute task that lands within 2 is worth it; one that drifts 15 isn't — and now you can tell the difference at a glance, task after task, in the history.&lt;/p&gt;
&lt;h2&gt;Install&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;bash&lt;/code&gt;, &lt;code&gt;awk&lt;/code&gt; and &lt;code&gt;curl&lt;/code&gt; are already on any Mac or Linux, so there is nothing to &lt;code&gt;npm install&lt;/code&gt;; the dashboard rides on &lt;code&gt;python3&lt;/code&gt;, also already there.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;git clone https://github.com/jcordon5/claudestimate.git
cd claudestimate
./install.sh&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It adds the statusline, drops in the skill and a &lt;code&gt;/claudestimate&lt;/code&gt; command, and opens the dashboard with a short tutorial. Restart Claude Code once, and from then on — in any project — Claude fills in the plan by itself on long tasks, and you get a finish time you can actually plan around. It pairs naturally with the way I already &lt;a href="https://yous.dev/blog/make-an-agent-finish-the-whole-job"&gt;hand big jobs to an agent&lt;/a&gt;: the plan that drives the work can also tell you when it'll be done.&lt;/p&gt;
&lt;p&gt;The thread running through these notes is that the agent increasingly works without you at the keyboard — and the point of that isn't idleness, it's spending your attention on the work that needs it while the machine handles the work that doesn't. If it's going to run on its own, the least it can do is tell you — honestly — how long you've got.&lt;/p&gt;
&lt;p&gt;Repo: &lt;strong&gt;&lt;a href="https://github.com/jcordon5/claudestimate" target="_blank" rel="noopener"&gt;github.com/jcordon5/claudestimate&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>How to make an AI agent finish the whole job</title>
    <link href="https://yous.dev/blog/make-an-agent-finish-the-whole-job"/>
    <id>https://yous.dev/blog/make-an-agent-finish-the-whole-job</id>
    <published>2026-08-01T00:00:00Z</published>
    <updated>2026-08-01T00:00:00Z</updated>
    <category term="agentic ai"/>
    <summary>Hand an agent a big, messy pile of work and get all of it back: a plan, a task table it can't lose, and a clean-context worker per task — plus the three prompts to steal.</summary>
    <content type="html">&lt;p class="post-sub"&gt;The last essay argued that the engineer's job is to steer. This is the operational half: how to hand a big, messy pile of work to an agent and get &lt;em&gt;all&lt;/em&gt; of it back — with a plan, a task table it can't lose, and a clean room per task.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/finish-the-job.jpg" alt="A lone figure on a winding path between a cloud of scattered shapes and a completed checklist" loading="lazy" width="1600" height="900"&gt;
&lt;figcaption&gt;From a scattered brief to a checklist that's actually done.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In &lt;a href="https://yous.dev/blog/ai-writes-the-code-who-steers-the-system"&gt;&lt;em&gt;AI writes the code. Who steers the system?&lt;/em&gt;&lt;/a&gt; I argued that once the machine handles implementation, the human's leverage moves to framing, judgment and direction. This post is about the part we hand over: once you've decided &lt;em&gt;what&lt;/em&gt; to build, how do you make the agent actually build all of it — correctly, and without quietly dropping half the work?&lt;/p&gt;
&lt;p&gt;Because that's the real failure mode. Give an agent a big brief — "do all of this" — and one of two things happens. Either it works in a single conversation until the context fills up, the output degrades, and it starts forgetting what it did at the start; or it declares victory having done maybe 60% of the list, and because the list was long, you don't notice the gaps until they bite you later. The naïve "just do everything" prompt is the problem, not the agent.&lt;/p&gt;
&lt;h2&gt;The first prompt's job is not to do the work&lt;/h2&gt;
&lt;p&gt;The move that fixes this: the first prompt you send is not "implement all of this." It's "turn all of this into a plan and a task table, ask me anything ambiguous, and don't start building until I've approved it."&lt;/p&gt;
&lt;p&gt;That reframing does a lot of work, and it isn't just my habit. &lt;a href="https://www.anthropic.com/research/building-effective-agents" target="_blank" rel="noopener"&gt;Anthropic's own guidance on building agents&lt;/a&gt; calls it the &lt;strong&gt;orchestrator-workers&lt;/strong&gt; pattern: a central agent breaks a complex task into subtasks, delegates them, and synthesizes the results. The orchestrator's first output isn't code — it's a decomposition. So the first prompt asks the lead agent for two artifacts:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;an &lt;strong&gt;implementation plan&lt;/strong&gt; — how it intends to approach the work, written after reading the whole brief;&lt;/li&gt;
&lt;li&gt;a &lt;strong&gt;segmented task table&lt;/strong&gt; in plain markdown — the work broken into small, well-scoped tasks, each with a status field: &lt;code&gt;todo / doing / done&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;And before it writes either, it has to &lt;strong&gt;ask you the questions&lt;/strong&gt; it needs to remove ambiguity. A good brief already carries most of the answers, but the agent surfaces the gaps you didn't notice. You answer, you correct its assumptions, you approve — and only then does anyone start building. No worker gets launched until the plan is validated; I've learned to be strict about that gate. (The exact prompt is &lt;a href="#prompts"&gt;at the end of this post&lt;/a&gt;.)&lt;/p&gt;
&lt;h2&gt;The task table is memory the agent can't lose&lt;/h2&gt;
&lt;p&gt;The task table is the load-bearing idea, and it's worth being precise about why. It is not documentation. It is &lt;strong&gt;shared, durable memory&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Every worker can read it: what's assigned, what's done, what's left. And here's the part that matters most — if a worker runs out of context halfway through its block, it doesn't silently stop. It re-reads the table, sees which of its tasks are still &lt;code&gt;todo&lt;/code&gt;, and finishes them. Nothing falls through the cracks, because "did I do everything?" is answered by a file, not by a model's fading memory of a long conversation.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;| # | task                          | owner  | status |
|---|-------------------------------|--------|--------|
| 1 | reshape the onboarding flow   | chip A | done   |
| 2 | add the new permission role   | chip B | doing  |
| 3 | rework the results view       | chip C | todo   |
| 4 | migrate the export templates  | chip D | todo   |&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Anthropic hit exactly this in &lt;a href="https://www.anthropic.com/engineering/multi-agent-research-system" target="_blank" rel="noopener"&gt;their multi-agent research system&lt;/a&gt;: they persist the lead agent's plan to external memory "when the context window exceeds certain thresholds to prevent loss of information." The task table is the same instinct, made checkable — a plan that outlives the context window.&lt;/p&gt;
&lt;h2&gt;Where the brief comes from: cleaning the brain-dump&lt;/h2&gt;
&lt;p&gt;It's worth saying where the brief itself comes from, because that's half the craft. Mine usually starts as raw notes — a meeting, a call, a stream of "we need to change this, and this, oh and this." That's the dump. It is not a prompt yet.&lt;/p&gt;
&lt;p&gt;The prompt is what comes out after a pass of editing: grouping the ideas, resolving contradictions, turning vague wishes into concrete tasks with a clear objective, and marking the genuine open questions to hand to the agent. The tasks are defined by a human. This is the essential-complexity work from the &lt;a href="https://yous.dev/blog/ai-writes-the-code-who-steers-the-system"&gt;last post&lt;/a&gt; — deciding what to build — and it's exactly the part the machine can't do for you. A vague brief just produces a vague plan, faster.&lt;/p&gt;
&lt;h2&gt;Clean-context workers beat one long chat&lt;/h2&gt;
&lt;p&gt;For the actual building, the temptation is to keep going in the same conversation. Resist it. The more you pack into one context, the worse the output gets — and this is now well documented, not a hunch. Models attend well to the beginning and end of a long context and &lt;a href="https://arxiv.org/abs/2307.03172" target="_blank" rel="noopener"&gt;lose the middle&lt;/a&gt;, and every frontier model tested shows &lt;a href="https://redis.io/blog/context-rot/" target="_blank" rel="noopener"&gt;measurable degradation as the input grows&lt;/a&gt;. In practice, a long build chat is one where you keep having to say "you missed a few things, keep going," each nudge costs more tokens, and the quality drifts down.&lt;/p&gt;
&lt;p&gt;The alternative is to give each task, or group of related tasks, its &lt;strong&gt;own fresh context&lt;/strong&gt;. In Claude these get spun off as separate chats — I call them chips — each starting clean, each running in its own parallel worktree so they can't step on one another. The lead agent writes the prompt for each worker: the task, the constraints, the conventions, anything it learned while planning. The worker doesn't need the whole history; it needs its slice, stated well.&lt;/p&gt;
&lt;p&gt;This is precisely why Anthropic's research system gives "each subagent a self-contained task description, an output format, and a fresh context window," and why that architecture beat the single-agent version by more than 90% on their evaluations. Separate context windows aren't a workaround; they're where the extra reasoning capacity comes from. A clean room per task produces better work than one increasingly cluttered desk.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/agent-workflow.jpg" alt="Flow diagram: a messy brief becomes a numbered plan, fans out to parallel workers, then a review step, then a completed checklist" loading="lazy" width="1600" height="900"&gt;
&lt;figcaption&gt;The whole loop: brief → plan → parallel workers → review → done.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2&gt;Read every worker, then unify&lt;/h2&gt;
&lt;p&gt;When the workers finish — and this can run for hours — the job is not over. This is the step people skip. You read every worker's output, one by one. They will have found things: a bug the plan didn't anticipate, a better way to group something, a note that they left a decision for you, a suggestion to spin off a follow-up. Some of the most valuable output of the whole run lives in these reports.&lt;/p&gt;
&lt;p&gt;Then you unify. A final pass — feed the lead agent each worker's conclusions plus instructions to reconcile them — and it checks the table (is everything marked correctly? is anything still open?), integrates the branches, and produces the single coherent version you'll validate. And validation is now easy, which is the entire point: you have the task table in front of you and you walk it, task by task, confirming each is what you wanted. Small fixes go straight to the main chat. Then delete the spun-off chats and keep the sidebar clean.&lt;/p&gt;
&lt;h2 id="prompts"&gt;The prompts, so you can steal them&lt;/h2&gt;
&lt;p&gt;You only ever write three prompts. The prompts the workers receive? The lead agent writes those — it composes one for each worker out of the plan and everything it learned while planning. That delegation is the whole point, so there's no worker template for you to fill in. Your three are the &lt;strong&gt;kickoff&lt;/strong&gt;, the &lt;strong&gt;feedback&lt;/strong&gt;, and the &lt;strong&gt;unify&lt;/strong&gt;. Here they are, stripped of any project — the angle brackets are yours to fill.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. The kickoff.&lt;/strong&gt; A little context, the objective, then the whole ordered brief, then the instruction to &lt;em&gt;plan&lt;/em&gt; — not build.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;&amp;lt;What just happened — the meeting, the call — and where this input comes from.&amp;gt;
The objective is the usual: a solid plan with the tasks to do, from what I give
you below, each with a status indicator so every agent can coordinate — what's
done, what's left — and nothing gets forgotten. Then you'll spin off several
clean-context workers so each does its part well. We work on a new branch, &amp;lt;name&amp;gt;.
Here is the brief:
&amp;lt;THE ORDERED BRIEF. This is the human work: your raw meeting notes cleaned up
into concrete, well-scoped tasks — each change written as something to do, the
open questions you actually want the agent to weigh in on, the goal behind each
item, and pointers to any files or references it needs.&amp;gt;
With that: analyse it and write the usual sprint-kickoff plan (an md file),
defining absolutely everything. Ask me whatever you need first so it's well
defined, with nothing left in the drawer. When we've both confirmed the plan,
we put the agents to work.
Do not launch a single worker until I've reviewed and approved the plan.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;2. The feedback.&lt;/strong&gt; You read the plan it wrote, answer its questions, correct what's off — and only then is it cleared to run.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;I've reviewed the plan. Here's my feedback; fold it in and then we're ready.
- &amp;lt;Answer each question the plan asked.&amp;gt;
- &amp;lt;A correction or a new decision — e.g. "X should actually work like this;
check whether what we already have covers it, and change what's needed."&amp;gt;
- &amp;lt;A challenge to one of its choices — "I don't see why we'd keep Y;
wouldn't Z cover it?"&amp;gt;
The rest looks perfect.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;3. The unify.&lt;/strong&gt; The workers have finished; this folds every report into one version for you to validate.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;All the workers finished. Here are their reports:
&amp;lt;paste each worker's report&amp;gt;
Review every report and the work behind it, and evaluate it. Check the task
table: is everything marked correctly, and is anything still open or missed?
Pull out the bugs they found, the decisions they left to me, and any follow-ups
they suggest, and list them for me. Then, if it's all sound, unify it —
integrate the branches into one coherent version — so I can start validating
by hand. Invent no new scope; if something is unresolved, surface it rather
than paper over it.&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;The bigger shift: one mind per project&lt;/h2&gt;
&lt;p&gt;Step back and the arithmetic is startling. Work that would have taken a team several days — or that a role-segregated team could only do by passing specifications back and forth — comes together in a single day, with everything at least attempted and a checklist to verify it against. Replit's CEO put the ceiling bluntly: agents could give one person the leverage of a hundred. That overshoots, but the direction is right, and a good run genuinely feels like it.&lt;/p&gt;
&lt;p&gt;Which changes how I think small consulting and research projects should be staffed. For years the default was a segregated team — frontend, backend, QA, integration — each person deep in one lane. When each of those lanes is something an agent now does well, keeping the segregation mostly adds coordination cost: the frontend dev prompts their slice, waits on the backend dev to prompt theirs, they meet to reconcile an API contract, another prompt wires it together — when one well-framed prompt, with shared context and clear goals, could have produced the whole feature at once.&lt;/p&gt;
&lt;p&gt;I don't think this means four of every five people are redundant. I think it means that where you needed five people for one project, you can now run five projects of one person each. Each person can cover what used to be uncoverable, because expert front-end, back-end and QA knowledge is on tap — the nearest thing to having that whole team on call, people who know each craft better than you do, with you as the one who decides where they point. You become the profile that used to hand you tasks; the model becomes the one that carries them out. What the project needs is one person who holds the whole of it: captures the requirements, analyzes them, designs the approach, defines the tasks, sets the validation criteria, and steers. Call it the project's implementation lead, or just its mind. The industry seems to be drifting the same way — toward &lt;a href="https://www.pwc.com/us/en/tech-effect/ai-analytics/agentic-ai-workforce-redesign.html" target="_blank" rel="noopener"&gt;generalists who own a problem end to end&lt;/a&gt; rather than specialists isolated by role.&lt;/p&gt;
&lt;p&gt;You take on the real work of an engineer — deciding &lt;em&gt;how to solve the problem&lt;/em&gt; — and hand the typing to the machine. Coding used to be the longest phase, but it was never the most important one: no amount of fast typing rescues a bad design, and the design is the part that needs human judgment. You own the &lt;em&gt;what&lt;/em&gt; and the &lt;em&gt;why&lt;/em&gt;. The &lt;em&gt;how&lt;/em&gt; is the machine's.&lt;/p&gt;
&lt;h2&gt;The takeaway&lt;/h2&gt;
&lt;p&gt;The code was never the hard part. Making sure &lt;em&gt;all&lt;/em&gt; of it gets built, correctly, from a clear idea of what "all of it" even is — that's the job. A plan, a task table the agent can't lose, and a clean room per task are how you keep it honest to the whole of it. You steer; it builds; the table makes sure nothing is left in the drawer.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>AI writes the code. Who steers the system?</title>
    <link href="https://yous.dev/blog/ai-writes-the-code-who-steers-the-system"/>
    <id>https://yous.dev/blog/ai-writes-the-code-who-steers-the-system</id>
    <published>2026-07-04T00:00:00Z</published>
    <updated>2026-07-04T00:00:00Z</updated>
    <category term="agentic ai"/>
    <summary>When implementation is automated, an engineer's value moves from memorizing syntax to defining, guiding and judging the system. An essay on the compiler mindset, cybernetics and output-first learning.</summary>
    <content type="html">&lt;p class="post-sub"&gt;A tweet about how we learn sent me down a rabbit hole, and it landed on something I keep seeing in my own work: when the machine writes the code, the value moves to whoever can define, guide and judge the system.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/steering-the-system.jpg" alt="A hand reaching down to move a lone chess pawn" loading="lazy" width="1600" height="900"&gt;
&lt;figcaption&gt;The machine plays the pieces. You still choose the move.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;For most of software's history the job rewarded one skill above all: turn a specification into correct code, by hand, fast. We memorized syntax, internalized edge cases, and took quiet pride in being reliable human compilers. That skill is being commoditized in front of us. Coding agents now produce the boilerplate, the tests, the refactors — the manual execution — faster and more consistently than I can by hand.&lt;/p&gt;
&lt;p&gt;Which raises an uncomfortable question. When the execution is automated, what exactly are you &lt;em&gt;for&lt;/em&gt;?&lt;/p&gt;
&lt;p&gt;This started, of all places, with &lt;a href="https://letters.thedankoe.com/p/how-to-remember-everything-you-read" target="_blank" rel="noopener"&gt;a piece by Dan Koe&lt;/a&gt; about how we learn. I pulled the thread — his essay, then the older ideas underneath it — and it named something I'd felt for a while. The engineers pulling ahead right now aren't the ones who memorized the most. They're the ones who can frame a problem well, hold the whole system in their head, and steer.&lt;/p&gt;
&lt;h2&gt;The compiler mindset is now a liability&lt;/h2&gt;
&lt;p&gt;We were trained to treat knowledge as something to hoard. School rewarded it: study for hours, recall on demand, fill in the blank. In engineering it shows up as trying to keep every API, flag and library in your head so you feel competent in the room.&lt;/p&gt;
&lt;p&gt;That instinct is now working against you, for a reason Fred Brooks named forty years ago. In &lt;a href="https://worrydream.com/refs/Brooks_1986_-_No_Silver_Bullet.pdf" target="_blank" rel="noopener"&gt;&lt;em&gt;No Silver Bullet&lt;/em&gt;&lt;/a&gt; (1986) he split the difficulty of building software in two: the &lt;strong&gt;accidental&lt;/strong&gt; — fighting your tools, syntax, ceremony — and the &lt;strong&gt;essential&lt;/strong&gt; — working out what to build, the design, the constraints, the trade-offs. His argument was that we'd already wrung most of the accidental complexity out, so the real work was, and would stay, essential.&lt;/p&gt;
&lt;p&gt;AI has now driven the accidental complexity close to zero. Andrej Karpathy put the endpoint bluntly: &lt;a href="https://x.com/karpathy/status/1617979122625712128" target="_blank" rel="noopener"&gt;"the hottest new programming language is English."&lt;/a&gt; The limiting factor stopped being whether you can express an idea in a particular syntax; it became whether you can express the idea clearly at all.&lt;/p&gt;
&lt;p&gt;So memorizing syntax to feel sharp is polishing the part that just got cheap. Understanding — how the pieces relate, where a design will bend, what a trade-off actually costs — is the part that got more valuable. If you freeze up in an architecture discussion, it's usually not a gap in intelligence. It's that years of training taught your brain to compile instructions instead of building models.&lt;/p&gt;
&lt;h2&gt;Learning is a feedback loop, not a download&lt;/h2&gt;
&lt;p&gt;Here's the idea from Koe's essay that reframed it for me. Cybernetics — Norbert Wiener's 1948 science of control and communication — takes its name from the Greek &lt;em&gt;kybernētēs&lt;/em&gt;: the &lt;strong&gt;helmsman&lt;/strong&gt;. A helmsman doesn't memorize the ocean. He fixes a destination, reads the current state, measures the gap, and corrects — over and over.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/cybernetic-loop.jpg" alt="Block diagram of a cybernetic feedback loop: goal to comparator, an error signal to the action or rudder, into the system, and feedback back to the comparator" loading="lazy" width="1448" height="1086"&gt;
&lt;figcaption&gt;The helmsman's loop: set a goal, act, sense the gap, correct — repeat.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Learning works the same way, and the order matters: output first, input second. A concrete goal produces an &lt;em&gt;error signal&lt;/em&gt; — the gap between where you are and where you want to be — and that error signal is the only thing that tells your brain what's worth keeping. Read a system design with a real problem in hand and you retain the pattern that closes your gap; read it "to learn it" and it evaporates by morning.&lt;/p&gt;
&lt;p&gt;This isn't only a nice metaphor. It's the &lt;a href="https://www.sciencedirect.com/science/article/abs/pii/S0749596X09001156" target="_blank" rel="noopener"&gt;generation effect&lt;/a&gt; in cognitive science: information you produce to solve a problem is retained far better than the same information re-read (Slamecka &amp;amp; Graf, 1978), and its close cousin, the testing effect, shows that retrieving a fact strengthens the memory more than re-studying it. Flashcards &lt;em&gt;feel&lt;/em&gt; like studying; building something that forces you to recall &lt;em&gt;is&lt;/em&gt; studying. It's also why an hour of doomscrolling leaves you drained — endless input, no error signal, nothing that needed keeping.&lt;/p&gt;
&lt;p&gt;The practical residue is small: keep a buffer for the &lt;em&gt;why&lt;/em&gt;, not the &lt;em&gt;how&lt;/em&gt;. Not code snippets a model can hand back to you in context — the mental models, the trade-offs, the shape of a decision (&lt;code&gt;[System X] → latency vs. consistency&lt;/code&gt;). Store what a search can't return.&lt;/p&gt;
&lt;h2&gt;Where the human leverage actually is&lt;/h2&gt;
&lt;p&gt;Steve Jobs liked to close launches on an image of two street signs meeting: "technology alone is not enough — it's technology married with liberal arts, married with humanities, that yields the results that make our hearts sing" (iPad, 2011). In the agentic era that stops being a nice sentiment and turns into a job description. When implementation costs fall toward zero, leverage moves to taste, framing and point of view — the things that decide &lt;em&gt;what&lt;/em&gt; to build and &lt;em&gt;whether&lt;/em&gt; it's any good.&lt;/p&gt;
&lt;p&gt;Concretely, my loop now has three moves, and only the middle one belongs to the machine:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Frame it.&lt;/strong&gt; Before opening an editor, define the problem, the constraints, the attack surface, the real trade-off between speed, security and scale. A vague frame just produces vague code, faster.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Delegate the execution.&lt;/strong&gt; Hand the framed spec to the agent and let it do the mud — boilerplate, tests, the syntax.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Judge and steer.&lt;/strong&gt; Be the comparator. Inspect the output, find the structural flaw, run the edge case that breaks it, correct the direction, repeat. It's the same discipline as &lt;a href="https://yous.dev/blog/prompts-became-loops"&gt;turning a prompt into a loop&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;That last move is the whole game, and it's exactly where AI stops being enough. Addy Osmani, who runs Chrome's developer experience at Google, calls it &lt;a href="https://addyo.substack.com/p/the-70-problem-hard-truths-about" target="_blank" rel="noopener"&gt;the 70% problem&lt;/a&gt;: agents get you 70% of the way fast, but the last 30% — edge cases, security, integration, the parts that decide whether it survives contact with production — still belongs to an engineer who understands how the system actually works. His warning is worth sitting with: developers who lean on AI without that understanding ship "house of cards" code, and review is quietly becoming the new bottleneck. The tools amplify judgment. They don't supply it.&lt;/p&gt;
&lt;h2&gt;The takeaway&lt;/h2&gt;
&lt;p&gt;Stop grading yourself on how much you can recall. Memory was never the moat, and now it's the cheapest thing in the stack. The moat is being able to frame a problem, hold the whole system at once, and tell — with judgment — the difference between output that's correct and output that's merely plausible.&lt;/p&gt;
&lt;p&gt;AI can write the code. Someone still has to steer the system. Increasingly, that's the job.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>Talk to your agents from Telegram</title>
    <link href="https://yous.dev/blog/talk-to-your-agents-from-telegram"/>
    <id>https://yous.dev/blog/talk-to-your-agents-from-telegram</id>
    <published>2026-06-25T00:00:00Z</published>
    <updated>2026-06-25T00:00:00Z</updated>
    <category term="agentic ai"/>
    <summary>A tiny skill that lets a coding agent message you while it works — progress, blocking questions, and a remote control from your phone. Idle is free, and several agents share one chat.</summary>
    <content type="html">&lt;p class="post-sub"&gt;Turning loops into something you can steer: progress pings, blocking questions, and a remote control — all from a Telegram chat in your pocket.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/telegram-bridge.jpg" alt="An AI agent checking in from a Telegram chat" loading="lazy" width="1600" height="900"&gt;
&lt;figcaption&gt;The loop keeps looping; you steer it from wherever you are.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In &lt;a href="https://yous.dev/blog/prompts-became-loops"&gt;&lt;em&gt;Prompts became loops&lt;/em&gt;&lt;/a&gt; I argued that the unit of work had shifted. A prompt used to be a one-shot: you ask, it answers, done. Now the unit of work is a &lt;strong&gt;loop&lt;/strong&gt; — the agent keeps going, takes steps, checks results, and continues until the goal is met. Coding agents like Claude Code already live this way.&lt;/p&gt;
&lt;p&gt;But loops have an awkward gap. Once an agent is off running for minutes or hours, &lt;strong&gt;you're no longer at the keyboard&lt;/strong&gt;. And sooner or later the loop needs &lt;em&gt;you&lt;/em&gt;: a decision, a missing credential, a "yes, ship it." If you're away, the loop stalls — or worse, the agent guesses. The loop runs without you, but it can't &lt;em&gt;reach&lt;/em&gt; you.&lt;/p&gt;
&lt;p&gt;So I built a tiny skill to close that gap: &lt;strong&gt;&lt;a href="https://github.com/jcordon5/claude-telegram-bridge" target="_blank" rel="noopener"&gt;telegram-bridge&lt;/a&gt;&lt;/strong&gt;. It lets an agent talk to me on Telegram while it works — send progress, &lt;strong&gt;ask blocking questions and wait for my answer&lt;/strong&gt;, and — the part I like most — &lt;strong&gt;let me drive it from my phone&lt;/strong&gt;. The loop keeps looping; I steer it from wherever I am.&lt;/p&gt;
&lt;h2&gt;What it does, in one screen&lt;/h2&gt;
&lt;p&gt;Three primitives, all stdlib Python (no installs, no SDK):&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;# progress — fire and forget
telegram.py notify "Tests green. Deploying to staging."
# blocking question — posts to Telegram, WAITS for your reply, prints it back
telegram.py ask "Migration will drop the legacy column. Proceed? (yes/no)"
# listen — wait for your next message and hand it to the agent as a new prompt
telegram.py listen&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That's the whole surface. The agent calls &lt;code&gt;notify&lt;/code&gt;/&lt;code&gt;ask&lt;/code&gt; when it has something to say; it runs &lt;code&gt;listen&lt;/code&gt; in the background to receive instructions from you.&lt;/p&gt;
&lt;h2&gt;The trick that keeps it cheap&lt;/h2&gt;
&lt;p&gt;Here's the part that matters if you've ever watched an agent burn tokens idling.&lt;/p&gt;
&lt;p&gt;A naive "let me check Telegram every few seconds" loop is &lt;strong&gt;the model polling&lt;/strong&gt; — and every tick re-bills the entire conversation context. That's expensive and pointless.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;listen&lt;/code&gt; flips it. The polling is done by a &lt;strong&gt;background shell process&lt;/strong&gt; doing a long-poll against Telegram's API — pure network, &lt;strong&gt;zero model tokens&lt;/strong&gt; while it waits. The model is asleep. The instant a message arrives, the process exits and hands the text to the agent, which wakes up &lt;strong&gt;once&lt;/strong&gt;, does the work, relaunches &lt;code&gt;listen&lt;/code&gt;, and goes back to sleep.&lt;/p&gt;
&lt;blockquote&gt;Token cost scales with the number of messages you send, not with how long you're away. Idle is free.&lt;/blockquote&gt;
&lt;p&gt;This is the loops idea taken one step further: the loop doesn't just run on its own — it &lt;strong&gt;parks for free and resumes on a human signal&lt;/strong&gt;, and that signal can come from your pocket.&lt;/p&gt;
&lt;h2&gt;A dispatcher you didn't have to build&lt;/h2&gt;
&lt;p&gt;The first version handled one agent. Then the obvious question: &lt;em&gt;what if I'm running several agents at once?&lt;/em&gt; One refactoring a backend, one writing docs, one babysitting CI. They'd all share one bot, and Telegram's update queue is consume-once — so they'd &lt;strong&gt;race&lt;/strong&gt; for your messages. Chaos.&lt;/p&gt;
&lt;p&gt;The fix is a &lt;strong&gt;broker&lt;/strong&gt;: a single background process that owns the Telegram connection and &lt;strong&gt;routes each message to the right agent&lt;/strong&gt;. It's a dispatcher — except you don't write it, configure it, or even think about it. It's &lt;strong&gt;plug &amp;amp; play and on by default&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The first command &lt;strong&gt;auto-starts&lt;/strong&gt; the broker (detached, survives the session).&lt;/li&gt;
&lt;li&gt;Each agent &lt;strong&gt;auto-names&lt;/strong&gt; itself after its project directory.&lt;/li&gt;
&lt;li&gt;Every outbound message is &lt;strong&gt;auto-tagged&lt;/strong&gt; &lt;code&gt;[backend] …&lt;/code&gt;, &lt;code&gt;[docs] …&lt;/code&gt; so you always know who's talking.&lt;/li&gt;
&lt;li&gt;The broker &lt;strong&gt;shuts itself down&lt;/strong&gt; when nobody's around.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You never run a setup step. It just works.&lt;/p&gt;
&lt;h3&gt;How you target a session from your phone&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;[backend] ✅ migration applied, running tests
[docs]    ❓ generate the API reference now? reply to confirm&lt;/code&gt;&lt;/pre&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Reply&lt;/strong&gt; to a message → the answer goes to that agent. Zero typing of names.&lt;/li&gt;
&lt;li&gt;Prefix &lt;code&gt;docs: regenerate the changelog&lt;/code&gt; → routes by name.&lt;/li&gt;
&lt;li&gt;Send &lt;code&gt;/sessions&lt;/code&gt; → a tap-to-pick &lt;strong&gt;inline keyboard&lt;/strong&gt; of live agents.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;stop&lt;/code&gt; stops the one you're talking to; &lt;code&gt;stop all&lt;/code&gt; stops everybody.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And because messages are tagged and routed, when a second agent shows up the broker drops you a one-time heads-up explaining how to switch — so there's nothing to memorize.&lt;/p&gt;
&lt;blockquote&gt;Side note for non-Claude setups: the whole thing is ~400 lines of stdlib Python driven by exit codes (&lt;code&gt;0&lt;/code&gt; = message on stdout, &lt;code&gt;4&lt;/code&gt; = stop, &lt;code&gt;3&lt;/code&gt; = timeout). Any headless agent or script that can shell out can use it as a Telegram control plane — a ready-made dispatcher for agents that don't have one.&lt;/blockquote&gt;
&lt;p&gt;It also persists its read offset, so if the broker restarts, messages you sent while it was briefly down are delivered when it comes back instead of vanishing.&lt;/p&gt;
&lt;h2&gt;Where this is genuinely useful&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Long unattended runs.&lt;/strong&gt; Kick off a multi-hour refactor, walk away, get pinged only when it finishes or needs a call. Answer from a café.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Human-in-the-loop without babysitting.&lt;/strong&gt; Dangerous step? The agent asks and &lt;em&gt;blocks&lt;/em&gt; until you say go — no guessing, no sitting at the terminal.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Steering from your phone.&lt;/strong&gt; New idea at dinner? Message the bot; the agent picks it up as its next prompt, exactly as if you'd typed it in the app.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fleet control.&lt;/strong&gt; Several agents in parallel, one chat, each message routed to the right one — a dispatcher you didn't build.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Security by default.&lt;/strong&gt; An allowlist means a stranger who finds the bot sees nothing and can't drive anything; only your chat id can.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Getting started&lt;/h2&gt;
&lt;pre&gt;&lt;code&gt;# install into Claude Code (or any skills-compatible agent)
npx skills add https://github.com/jcordon5/claude-telegram-bridge.git&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Create a bot with @BotFather, drop the token + your chat id into the skill's &lt;code&gt;.env&lt;/code&gt;, and tell your agent: &lt;em&gt;"work on this and ping me on Telegram if you get stuck — and let me drive you from there."&lt;/em&gt; That's it.&lt;/p&gt;
&lt;p&gt;It pairs naturally with the loop mindset I wrote about in &lt;a href="https://yous.dev/blog/prompts-became-loops"&gt;&lt;em&gt;Prompts became loops&lt;/em&gt;&lt;/a&gt;: if your prompts have already become loops, this is how you put a remote control in your hand. The agent runs the loop; you steer it from your pocket.&lt;/p&gt;
&lt;p&gt;Repo: &lt;strong&gt;&lt;a href="https://github.com/jcordon5/claude-telegram-bridge" target="_blank" rel="noopener"&gt;github.com/jcordon5/claude-telegram-bridge&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>How to deploy a HuggingFace model with Ollama</title>
    <link href="https://yous.dev/blog/deploy-huggingface-models-with-ollama"/>
    <id>https://yous.dev/blog/deploy-huggingface-models-with-ollama</id>
    <published>2026-06-16T00:00:00Z</published>
    <updated>2026-06-16T00:00:00Z</updated>
    <category term="tutorials"/>
    <summary>An end-to-end guide: GGUF and quantization, the Modelfile, prompt templates per model family, the local API, and the errors that bite — from a HuggingFace repo to a running endpoint.</summary>
    <content type="html">&lt;p class="post-sub"&gt;A practical guide to taking an open-weight model from a HuggingFace repository to a local Ollama endpoint, ready to consume over API or CLI.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/ollama.jpg" alt="Ollama, a local inference runtime" loading="lazy" width="1600" height="755"&gt;
&lt;figcaption&gt;Ollama runs open-weight models locally.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2&gt;1. Background concepts&lt;/h2&gt;
&lt;h3&gt;1.1 What is Ollama?&lt;/h3&gt;
&lt;p&gt;Ollama is a local inference runtime built on top of &lt;code&gt;llama.cpp&lt;/code&gt;. It handles:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Loading quantized models in &lt;strong&gt;GGUF&lt;/strong&gt; format.&lt;/li&gt;
&lt;li&gt;Exposing a compatible HTTP API (port &lt;code&gt;11434&lt;/code&gt; by default) with endpoints like &lt;code&gt;/api/generate&lt;/code&gt;, &lt;code&gt;/api/chat&lt;/code&gt; and &lt;code&gt;/api/embeddings&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Managing prompt templates, sampling parameters, and model download/storage.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It installs as a service (&lt;code&gt;systemd&lt;/code&gt; on Linux, a native app on macOS/Windows) and ships a CLI (&lt;code&gt;ollama&lt;/code&gt;) that wraps the API.&lt;/p&gt;
&lt;h3&gt;1.2 What is GGUF?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;GGUF&lt;/strong&gt; (GPT-Generated Unified Format) is the binary format &lt;code&gt;llama.cpp&lt;/code&gt; uses to store quantized weights, the tokenizer and model metadata in a single file. It's the format Ollama consumes directly.&lt;/p&gt;
&lt;h3&gt;1.3 Common quantizations&lt;/h3&gt;
&lt;div class="table-wrap"&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Label&lt;/th&gt;&lt;th&gt;Bits/weight approx.&lt;/th&gt;&lt;th&gt;Quality&lt;/th&gt;&lt;th&gt;Typical use&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;BF16&lt;/code&gt; / &lt;code&gt;F16&lt;/code&gt;&lt;/td&gt;&lt;td&gt;16&lt;/td&gt;&lt;td&gt;Maximum (reference)&lt;/td&gt;&lt;td&gt;GPU with comfortable VRAM&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;Q8_0&lt;/code&gt;&lt;/td&gt;&lt;td&gt;8&lt;/td&gt;&lt;td&gt;Almost identical to F16&lt;/td&gt;&lt;td&gt;Good quality/size balance&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;Q6_K&lt;/code&gt;&lt;/td&gt;&lt;td&gt;~6.5&lt;/td&gt;&lt;td&gt;Very high&lt;/td&gt;&lt;td&gt;Machines with less VRAM&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;Q5_K_M&lt;/code&gt;&lt;/td&gt;&lt;td&gt;~5.5&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;Fast inference&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;Q4_K_M&lt;/code&gt;&lt;/td&gt;&lt;td&gt;~4.8&lt;/td&gt;&lt;td&gt;Acceptable&lt;/td&gt;&lt;td&gt;Large models on modest hardware&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;Q3_K_*&lt;/code&gt; / &lt;code&gt;Q2_K&lt;/code&gt;&lt;/td&gt;&lt;td&gt;2–3&lt;/td&gt;&lt;td&gt;Degraded&lt;/td&gt;&lt;td&gt;Only when there's no alternative&lt;/td&gt;&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;/div&gt;
&lt;p&gt;Rule of thumb: if the quantized model fits in VRAM, &lt;strong&gt;&lt;code&gt;Q8_0&lt;/code&gt; or &lt;code&gt;Q6_K&lt;/code&gt; are the default choice&lt;/strong&gt;. Drop to &lt;code&gt;Q4&lt;/code&gt; only when the model's size justifies it.&lt;/p&gt;
&lt;h2&gt;2. Requirements&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;A server with Ollama installed and the service running (&lt;code&gt;systemctl status ollama&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Enough disk space (the GGUF + ~2× for Ollama's internal layers during &lt;code&gt;create&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Internet access to download from HuggingFace.&lt;/li&gt;
&lt;li&gt;Optional: an NVIDIA GPU with drivers + CUDA. Ollama detects and uses the GPU automatically.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Quick checks:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;ollama --version
ollama list                  # models already registered
curl http://localhost:11434  # should respond "Ollama is running"
df -h                        # available space&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;3. Download the GGUF from HuggingFace&lt;/h2&gt;
&lt;h3&gt;3.1 Locate the file&lt;/h3&gt;
&lt;p&gt;On the model page, in the &lt;strong&gt;Files and versions&lt;/strong&gt; tab, look for the &lt;code&gt;.gguf&lt;/code&gt; files. You'll see something like:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;my-model-7B-BF16.gguf       14 GB
my-model-7B-Q8_0.gguf        7.5 GB   ← recommended
my-model-7B-Q6_K.gguf        5.8 GB
my-model-7B-Q4_K_M.gguf      4.4 GB&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Each file exposes a &lt;code&gt;resolve&lt;/code&gt; URL that serves the binary directly:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;https://huggingface.co/&amp;lt;user&amp;gt;/&amp;lt;repo&amp;gt;/resolve/main/&amp;lt;file&amp;gt;.gguf&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;3.2 Direct download with &lt;code&gt;wget&lt;/code&gt;&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;mkdir -p ~/models/my-model-7B
cd ~/models/my-model-7B
wget -c \
-O my-model-7B-Q8_0.gguf \
"https://huggingface.co/user/repo/resolve/main/my-model-7B-Q8_0.gguf?download=true"&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;-c&lt;/code&gt; lets you resume if the download breaks. For private repos add &lt;code&gt;--header="Authorization: Bearer hf_xxxxxxxx"&lt;/code&gt;.&lt;/p&gt;
&lt;h3&gt;3.3 Alternative: &lt;code&gt;huggingface-cli&lt;/code&gt;&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;pip install -U huggingface_hub
huggingface-cli download user/repo my-model-7B-Q8_0.gguf \
--local-dir ~/models/my-model-7B --local-dir-use-symlinks False&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;3.4 What if the repo has no GGUF?&lt;/h3&gt;
&lt;p&gt;Some repos only publish the original weights (HuggingFace Transformers &lt;code&gt;.safetensors&lt;/code&gt;). You have to convert and quantize to GGUF. The cleanest way is to use the &lt;code&gt;llama.cpp&lt;/code&gt; Docker image so you don't pollute the host:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;# Convert HF -&amp;gt; GGUF F16
docker run --rm -v ~/models:/models ghcr.io/ggml-org/llama.cpp:full \
--convert /models/my-model-hf-src \
--outfile /models/my-model-7B/my-model-7B-F16.gguf \
--outtype f16
# Quantize F16 -&amp;gt; Q8_0
docker run --rm -v ~/models:/models ghcr.io/ggml-org/llama.cpp:full \
--quantize /models/my-model-7B/my-model-7B-F16.gguf \
/models/my-model-7B/my-model-7B-Q8_0.gguf Q8_0&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Because the image runs with &lt;code&gt;--rm&lt;/code&gt;, the container is cleaned up automatically and nothing residual is left outside the mounted directory.&lt;/p&gt;
&lt;h2&gt;4. The Modelfile&lt;/h2&gt;
&lt;p&gt;A &lt;strong&gt;Modelfile&lt;/strong&gt; is the "recipe" Ollama uses to register a model. It defines where to get the weights, which chat template to apply, which default parameters to set, and which system prompt to ship pre-loaded. It's the conceptual equivalent of a &lt;code&gt;Dockerfile&lt;/code&gt;, but for models.&lt;/p&gt;
&lt;h3&gt;4.1 Structure&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;# (1) Weights source
FROM /absolute/path/to/model.gguf
# (2) Prompt template
TEMPLATE """..."""
# (3) Sampling and context parameters
PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER num_ctx 8192
PARAMETER stop "&amp;lt;|im_end|&amp;gt;"
# (4) Default system prompt (optional)
SYSTEM """You are a helpful, concise assistant."""
# (5) LoRA adapters (optional)
# ADAPTER /path/to/adapter.gguf
# (6) Embedded license (optional)
# LICENSE """..."""&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;4.2 Relevant directives&lt;/h3&gt;
&lt;div class="table-wrap"&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Directive&lt;/th&gt;&lt;th&gt;Function&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;FROM&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Path to the GGUF, or the name of another model already in Ollama (&lt;code&gt;FROM llama3:8b&lt;/code&gt;).&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;TEMPLATE&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Go-template that wraps the &lt;code&gt;prompt&lt;/code&gt; with special tokens.&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;RENDERER&lt;/code&gt; / &lt;code&gt;PARSER&lt;/code&gt;&lt;/td&gt;&lt;td&gt;(Modern Ollama) Native renderers for known architectures — avoids writing the template by hand.&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;PARAMETER&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Tunes sampling (&lt;code&gt;temperature&lt;/code&gt;, &lt;code&gt;top_p&lt;/code&gt;, &lt;code&gt;top_k&lt;/code&gt;, &lt;code&gt;min_p&lt;/code&gt;, &lt;code&gt;repeat_penalty&lt;/code&gt;), context (&lt;code&gt;num_ctx&lt;/code&gt;), and stop tokens (&lt;code&gt;stop&lt;/code&gt;).&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;SYSTEM&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Default system prompt if the request doesn't provide one.&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;ADAPTER&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Applies a LoRA on top of the base model.&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;MESSAGE&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Adds few-shot examples the model sees at the start.&lt;/td&gt;&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;/div&gt;
&lt;h2&gt;5. Prompt templates: why they matter&lt;/h2&gt;
&lt;p&gt;Each model family expects its messages wrapped in specific &lt;strong&gt;special tokens&lt;/strong&gt;. If the template doesn't match:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The model doesn't know where each turn starts/ends → erratic responses.&lt;/li&gt;
&lt;li&gt;It generates the wrong end token → it never stops, or stops mid-sentence.&lt;/li&gt;
&lt;li&gt;It ignores the &lt;code&gt;system prompt&lt;/code&gt; → it doesn't respect the role.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the most widespread formats.&lt;/p&gt;
&lt;h3&gt;5.1 ChatML (Qwen, OpenAI style)&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;&amp;lt;|im_start|&amp;gt;system
You are a helpful assistant.&amp;lt;|im_end|&amp;gt;
&amp;lt;|im_start|&amp;gt;user
Hi&amp;lt;|im_end|&amp;gt;
&amp;lt;|im_start|&amp;gt;assistant&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Typical Modelfile:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;FROM /models/qwen-7b-q8_0.gguf
TEMPLATE """{{- if .System }}&amp;lt;|im_start|&amp;gt;system
{{ .System }}&amp;lt;|im_end|&amp;gt;
{{ end }}&amp;lt;|im_start|&amp;gt;user
{{ .Prompt }}&amp;lt;|im_end|&amp;gt;
&amp;lt;|im_start|&amp;gt;assistant
"""
PARAMETER stop "&amp;lt;|im_start|&amp;gt;"
PARAMETER stop "&amp;lt;|im_end|&amp;gt;"&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;5.2 Llama 3 / Llama 4&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;&amp;lt;|begin_of_text|&amp;gt;&amp;lt;|start_header_id|&amp;gt;system&amp;lt;|end_header_id|&amp;gt;
You are helpful.&amp;lt;|eot_id|&amp;gt;&amp;lt;|start_header_id|&amp;gt;user&amp;lt;|end_header_id|&amp;gt;
Hi&amp;lt;|eot_id|&amp;gt;&amp;lt;|start_header_id|&amp;gt;assistant&amp;lt;|end_header_id|&amp;gt;&lt;/code&gt;&lt;/pre&gt;
&lt;pre&gt;&lt;code&gt;TEMPLATE """&amp;lt;|begin_of_text|&amp;gt;{{- if .System }}&amp;lt;|start_header_id|&amp;gt;system&amp;lt;|end_header_id|&amp;gt;
{{ .System }}&amp;lt;|eot_id|&amp;gt;{{ end }}&amp;lt;|start_header_id|&amp;gt;user&amp;lt;|end_header_id|&amp;gt;
{{ .Prompt }}&amp;lt;|eot_id|&amp;gt;&amp;lt;|start_header_id|&amp;gt;assistant&amp;lt;|end_header_id|&amp;gt;
"""
PARAMETER stop "&amp;lt;|eot_id|&amp;gt;"&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;5.3 Mistral / Mixtral (&lt;code&gt;[INST]&lt;/code&gt; format)&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;&amp;lt;s&amp;gt;[INST] You are helpful.
Hi [/INST]&lt;/code&gt;&lt;/pre&gt;
&lt;pre&gt;&lt;code&gt;TEMPLATE """[INST] {{ if .System }}{{ .System }}
{{ end }}{{ .Prompt }} [/INST]
"""
PARAMETER stop "[INST]"
PARAMETER stop "[/INST]"&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;5.4 Gemma&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;&amp;lt;start_of_turn&amp;gt;user
Hi&amp;lt;end_of_turn&amp;gt;
&amp;lt;start_of_turn&amp;gt;model&lt;/code&gt;&lt;/pre&gt;
&lt;pre&gt;&lt;code&gt;TEMPLATE """&amp;lt;start_of_turn&amp;gt;user
{{ if .System }}{{ .System }}
{{ end }}{{ .Prompt }}&amp;lt;end_of_turn&amp;gt;
&amp;lt;start_of_turn&amp;gt;model
"""
PARAMETER stop "&amp;lt;start_of_turn&amp;gt;"
PARAMETER stop "&amp;lt;end_of_turn&amp;gt;"&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;5.5 Phi-3&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;&amp;lt;|system|&amp;gt;
You are helpful.&amp;lt;|end|&amp;gt;
&amp;lt;|user|&amp;gt;
Hi&amp;lt;|end|&amp;gt;
&amp;lt;|assistant|&amp;gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;5.6 Native RENDERERs (a shortcut in modern Ollama)&lt;/h3&gt;
&lt;p&gt;Recent Ollama versions ship native renderers and parsers for specific architectures. Instead of writing the template, you declare:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;FROM /models/model-x.gguf
TEMPLATE {{ .Prompt }}
RENDERER architecture-name
PARSER   architecture-name
PARAMETER temperature 0.7
PARAMETER top_p 0.8&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The renderer injects the correct tokens, and the parser separates &lt;strong&gt;thinking&lt;/strong&gt; from &lt;strong&gt;content&lt;/strong&gt; in reasoning models. When an official renderer exists for the model's family, this is the most robust option and the one I recommend.&lt;/p&gt;
&lt;p&gt;To discover which template any already-registered model uses:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;ollama show --modelfile &amp;lt;model&amp;gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It's the best source of truth: copy and adapt from a model of the same family that already works.&lt;/p&gt;
&lt;h2&gt;6. Create the model in Ollama&lt;/h2&gt;
&lt;p&gt;With the GGUF downloaded and the &lt;code&gt;Modelfile_mymodel&lt;/code&gt; ready:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;ollama create mymodel:7b-q8_0 -f Modelfile_mymodel&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;What happens under the hood:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Ollama &lt;strong&gt;copies the GGUF into its blob store&lt;/strong&gt; (&lt;code&gt;~/.ollama/models/blobs/&lt;/code&gt; or &lt;code&gt;/usr/share/ollama/.ollama/models/blobs/&lt;/code&gt; depending on the install), named by its SHA-256.&lt;/li&gt;
&lt;li&gt;Generates additional &lt;strong&gt;layers&lt;/strong&gt; with the template, parameters and system prompt.&lt;/li&gt;
&lt;li&gt;Writes a &lt;strong&gt;manifest&lt;/strong&gt; tying everything under the tag &lt;code&gt;mymodel:7b-q8_0&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Important: the copy temporarily doubles disk usage. Once created, &lt;strong&gt;you can delete the original GGUF&lt;/strong&gt; (Ollama already has its copy in the blob store).&lt;/p&gt;
&lt;p&gt;Verify:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;ollama list
ollama show mymodel:7b-q8_0&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;7. Test the model&lt;/h2&gt;
&lt;h3&gt;7.1 Interactive CLI&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;ollama run mymodel:7b-q8_0
&amp;gt;&amp;gt;&amp;gt; Hi, how are you?&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;7.2 REST API (&lt;code&gt;/api/chat&lt;/code&gt;)&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;curl http://localhost:11434/api/chat -d '{
"model": "mymodel:7b-q8_0",
"messages": [
{"role": "system", "content": "Always answer in English."},
{"role": "user",   "content": "Summarize what GGUF is in one sentence."}
],
"stream": false
}'&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;7.3 Raw generation API (&lt;code&gt;/api/generate&lt;/code&gt;)&lt;/h3&gt;
&lt;p&gt;Useful for debugging the template — pass &lt;code&gt;"raw": true&lt;/code&gt; and send the special tokens yourself:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;curl http://localhost:11434/api/generate -d '{
"model": "mymodel:7b-q8_0",
"prompt": "&amp;lt;|im_start|&amp;gt;user\nHi&amp;lt;|im_end|&amp;gt;\n&amp;lt;|im_start|&amp;gt;assistant\n",
"raw": true,
"stream": false
}'&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;7.4 Reasoning models (thinking)&lt;/h3&gt;
&lt;p&gt;Modern reasoning families (Qwen3.6, DeepSeek-R1, etc.) emit a thinking trace separate from the final content. The API exposes it in the &lt;code&gt;message.thinking&lt;/code&gt; field:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;{
"message": {
"role": "assistant",
"thinking": "Let's work through it step by step...",
"content": "The result is 42."
}
}&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;To disable the trace on requests where you only want the answer:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;curl http://localhost:11434/api/chat -d '{
"model": "mymodel:7b-q8_0",
"messages": [{"role": "user", "content": "2+2"}],
"think": false
}'&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;8. Expose the service to the network&lt;/h2&gt;
&lt;p&gt;By default Ollama listens only on &lt;code&gt;127.0.0.1&lt;/code&gt;. To expose it on the LAN, edit the systemd override:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;sudo systemctl edit ollama&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And add:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Reload:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;sudo systemctl daemon-reload
sudo systemctl restart ollama&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;⚠️ Ollama ships no authentication. If you expose it beyond a trusted network, put a reverse proxy (nginx/Caddy/Traefik) in front with TLS and auth (Basic, an OAuth proxy, mTLS — whatever fits).&lt;/blockquote&gt;
&lt;p&gt;Useful variables:&lt;/p&gt;
&lt;div class="table-wrap"&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Variable&lt;/th&gt;&lt;th&gt;What for&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;OLLAMA_HOST&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Listen address/port.&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;OLLAMA_MODELS&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Blob store directory (handy to move it to another disk).&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;OLLAMA_KEEP_ALIVE&lt;/code&gt;&lt;/td&gt;&lt;td&gt;How long to keep the model loaded in VRAM (&lt;code&gt;5m&lt;/code&gt;, &lt;code&gt;1h&lt;/code&gt;, &lt;code&gt;-1&lt;/code&gt; forever).&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;OLLAMA_NUM_PARALLEL&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Concurrent requests per model.&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;OLLAMA_MAX_LOADED_MODELS&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Simultaneous models held in memory.&lt;/td&gt;&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;/div&gt;
&lt;h2&gt;9. Maintenance&lt;/h2&gt;
&lt;pre&gt;&lt;code&gt;ollama list                       # what's registered
ollama ps                         # what's loaded in VRAM right now
ollama show --modelfile &amp;lt;m&amp;gt;        # see a model's recipe
ollama cp old:tag new:tag         # duplicate/rename
ollama rm &amp;lt;model&amp;gt;                  # delete a model and free disk
journalctl -u ollama -f           # live logs&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;To iterate on the Modelfile (change the template, tweak &lt;code&gt;temperature&lt;/code&gt;, etc.) just edit it and run &lt;code&gt;ollama create&lt;/code&gt; again with the same tag — it overwrites.&lt;/p&gt;
&lt;h2&gt;10. Common errors and how to diagnose them&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The model won't stop generating.&lt;/strong&gt; A &lt;code&gt;PARAMETER stop&lt;/code&gt; with the correct end token is missing, or the template doesn't close the turn with the token the model learned. Check the repo's &lt;code&gt;tokenizer_config.json&lt;/code&gt; to confirm the &lt;code&gt;eos_token&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empty or garbled responses.&lt;/strong&gt; The template isn't the one for the model's family. Pull the Modelfile from another model in the same family (&lt;code&gt;ollama show --modelfile&lt;/code&gt;) and use it as a reference.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;Error: invalid file magic&lt;/code&gt;.&lt;/strong&gt; The GGUF is corrupt or incomplete. Resume the download with &lt;code&gt;wget -c&lt;/code&gt; and compare the size against what HuggingFace shows.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;Error: model requires more system memory than is available&lt;/code&gt;.&lt;/strong&gt; It doesn't fit in VRAM. Lower the quantization, reduce &lt;code&gt;num_ctx&lt;/code&gt;, or let Ollama offload layers to the CPU (slower, but it works).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The model loads but the GPU sits at 0%.&lt;/strong&gt; NVIDIA drivers misaligned with the CUDA version. &lt;code&gt;nvidia-smi&lt;/code&gt; should run clean; if it reports "Driver/library version mismatch", reboot or reinstall the drivers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hangs on the first request.&lt;/strong&gt; Ollama is loading the model into memory. The initial latency can be several seconds; later loads are instant while the model stays in VRAM.&lt;/p&gt;
&lt;h2&gt;11. Flow summary&lt;/h2&gt;
&lt;pre&gt;&lt;code&gt;┌─────────────────────┐
│  HuggingFace repo   │
│  (.gguf quantized)  │
└──────────┬──────────┘
│  wget / huggingface-cli
▼
┌─────────────────────┐
│  ~/models/&amp;lt;model&amp;gt;/   │
│  file.gguf          │
└──────────┬──────────┘
│
│  + Modelfile (TEMPLATE, PARAMETER, SYSTEM)
▼
┌─────────────────────┐
│  ollama create      │ ──► blob store + manifest
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│  ollama run / API   │   :11434/api/chat
└─────────────────────┘&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Three files, three commands: download, write the Modelfile, &lt;code&gt;ollama create&lt;/code&gt;. The rest is knowing which template your model wants and tuning the sampling parameters to the use case.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>The task file is the prompt</title>
    <link href="https://yous.dev/blog/the-task-file-is-the-prompt"/>
    <id>https://yous.dev/blog/the-task-file-is-the-prompt</id>
    <published>2026-06-15T00:00:00Z</published>
    <updated>2026-06-15T00:00:00Z</updated>
    <category term="agentic ai"/>
    <summary>How a CLAUDE.md and a folder of markdown tasks let a team share context and hand work to AI agents — no prompts, no third party.</summary>
    <content type="html">&lt;p class="post-sub"&gt;How a &lt;code&gt;CLAUDE.md&lt;/code&gt; and a folder of markdown tasks let a team share context and hand work to AI agents — version-controlled, vendor-independent, and with nobody writing prompts.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/tasks.jpg" alt="Diagram: a CLAUDE.md and a tasks folder that an agent reads, works from, and commits back" loading="lazy" width="1600" height="900"&gt;
&lt;figcaption&gt;The whole loop on one page.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Working with AI on a real project, day to day, quietly turns into a coordination problem. Someone spots a thing that needs doing. They open a chat, re-explain the project, think up the right prompt, phrase it, paste in the relevant files, and wait. Tomorrow a teammate does the same dance from scratch — same context re-typed, a slightly different prompt, a slightly different result.&lt;/p&gt;
&lt;p&gt;The work isn't the bottleneck. The &lt;em&gt;re-explaining&lt;/em&gt; is. So I stopped explaining to the AI and started writing it down once, in the repo, where both the humans and the agents already look.&lt;/p&gt;
&lt;h2&gt;Two files do most of the work&lt;/h2&gt;
&lt;p&gt;The whole system is two things that live in the codebase and travel with it through git.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A &lt;code&gt;CLAUDE.md&lt;/code&gt; at the root&lt;/strong&gt; — the standing context. What the project is, the non-negotiable rules, how we build, where the tasks are, a quick map of the repo. An agent reads it first and inherits the same constraints every teammate works under. It's the onboarding doc you never have to give twice.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;# CLAUDE.md — how we work in this repo
## What this is
One paragraph an agent can absorb in five seconds.
## Non-negotiable rules
- Never touch X without Y.
- Migrations are append-only.
## How we build
Stack, conventions, what "done" means, how to run the tests.
## Where the tasks live
docs/tasks/ — read STATE.md first, then pick one.
## Repo map
src/… · migrations/… · docs/…&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;A folder of tasks&lt;/strong&gt; — the living backlog, also in markdown. Start with one file; split it as it grows.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;docs/tasks/
STATE.md     # the big picture: what's done, what's next, by area
QUICK.md     # small things, feedback, one-liners
billing.md   # grouped by type once a single file gets noisy&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;The loop, minus the prompting&lt;/h2&gt;
&lt;p&gt;Once the context and the backlog are in the repo, the day-to-day collapses into something almost boring:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;You notice work that needs doing.&lt;/li&gt;
&lt;li&gt;You write it as a task and commit it. That's the whole "assignment."&lt;/li&gt;
&lt;li&gt;A teammate opens the repo and tells their agent: &lt;em&gt;pick up the next task&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;The agent reads &lt;code&gt;CLAUDE.md&lt;/code&gt; and the task files, does the work, ticks the box, and commits.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Nobody wrote a prompt. Nobody re-explained the project. The person who spotted the work didn't have to be the person who did it, and didn't have to be available when it got done. The task &lt;em&gt;was&lt;/em&gt; the instruction — which is exactly why it has to be written like one.&lt;/p&gt;
&lt;blockquote&gt;If a teammate can read the task and know what "done" looks like, so can an agent. If they can't, neither can the agent.&lt;/blockquote&gt;
&lt;p&gt;A task that works looks less like a wish and more like a small contract — goal, why, and a checkable definition of done:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;## Add CSV export to the reports page
Status: todo
Why: people keep asking for it.
Done when:
- there's a "Download CSV" button on the reports page
- the columns match the on-screen table exactly
- it's covered by a test
Context: the existing PDF export is the closest pattern to copy.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This is the same idea as &lt;a href="https://yous.dev/blog/prompts-became-loops"&gt;turning a prompt into a loop&lt;/a&gt;: the value isn't in the cleverness of the phrasing, it's in stating the goal, the constraints, and how you'll know it's finished. Write the task that well and the prompt writes itself.&lt;/p&gt;
&lt;h2&gt;You don't even have to write the task&lt;/h2&gt;
&lt;p&gt;Put that format — the task "contract" — into &lt;code&gt;CLAUDE.md&lt;/code&gt; itself, and the human stops writing tasks at all. You just tell your agent &lt;em&gt;"add a task to let users export reports as CSV,"&lt;/em&gt; and the agent, which already knows where the tasks live and what shape a task should take, writes it out for you: well-phrased, in the right format, in the right file. If something's missing — a why, a definition of done — it asks you before committing it.&lt;/p&gt;
&lt;p&gt;So the floor drops one more level. You describe the work in a sentence; the agent turns it into a proper task; the next agent picks it up and does it. You can still hand-write a task when you want the wording exactly so — but you no longer have to, and the format stays consistent no matter who (or what) wrote it.&lt;/p&gt;
&lt;h2&gt;Why markdown and git, not a board&lt;/h2&gt;
&lt;p&gt;You could keep all of this in a project-management tool, and if your company lives in Jira there are agent skills that read tasks straight from it. But putting it in the repo buys things a board can't:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;It's where the agent already is.&lt;/strong&gt; No integration, no API token, no third party between the work and the code.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It versions with the code.&lt;/strong&gt; The task, the commit that did it, and the change all sit in the same history. &lt;code&gt;git blame&lt;/code&gt; on a decision actually works.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It's diffable and reviewable.&lt;/strong&gt; Adding work is a commit. Changing scope is a diff. Disagreeing is a comment on a line.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It's vendor-independent.&lt;/strong&gt; No seat, no migration, no outage that takes your backlog offline.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Letting it grow&lt;/h2&gt;
&lt;p&gt;It starts as a single &lt;code&gt;TASKS.md&lt;/code&gt; and that's enough for a while. When one file gets noisy, split by intent rather than by person: a &lt;code&gt;STATE.md&lt;/code&gt; for the big moving picture, a &lt;code&gt;QUICK.md&lt;/code&gt; for small feedback, a file per area when an area earns one. The structure should mirror how you actually think about the work, not an org chart.&lt;/p&gt;
&lt;p&gt;And keep it from bloating. Once a task is done, either delete it — the git history already remembers it ever existed — or move it to a &lt;code&gt;DONE.md&lt;/code&gt; if you want a visible record of what shipped. Both are fine; the only real mistake is letting finished work pile up until the file an agent has to read is mostly noise.&lt;/p&gt;
&lt;p&gt;The point isn't the folder layout. It's the shift underneath it: the shared context stops living in people's heads and DMs, and starts living in the repo as plain text that humans write and agents execute. Humans decide &lt;em&gt;what&lt;/em&gt; and &lt;em&gt;why&lt;/em&gt;; the agents handle a lot of the &lt;em&gt;how&lt;/em&gt;; and the handoff between them is just a commit.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>You don't need to "stop prompting"</title>
    <link href="https://yous.dev/blog/prompts-became-loops"/>
    <id>https://yous.dev/blog/prompts-became-loops</id>
    <published>2026-06-15T00:00:00Z</published>
    <updated>2026-06-15T00:00:00Z</updated>
    <category term="agentic ai"/>
    <summary>When a task stops being a request and becomes a process, the prompt grows a loop. The full guide — goal, verification, guardrails, memory — and when not to bother.</summary>
    <content type="html">&lt;p class="post-sub"&gt;For years we talked about prompt engineering as if the main challenge were finding the right sentence. With modern agents a practical difference appears: many tasks should no longer be solved with a single answer, but with a cycle of work.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/loop.jpg" alt="An agentic loop running, iteration by iteration" loading="lazy" width="1600" height="893"&gt;
&lt;figcaption&gt;A loop, mid-iteration.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Better context, better role, better format, better examples. All of that still matters. But that cycle now tends to be called a &lt;em&gt;loop&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;The term can sound newer than it really is. In its simplest form, a loop is not magic and not a new profession separate from prompting. It's a well-designed prompt that includes a goal, verification, repetition, correction and a stop. The difference is not using a new word. It's that you stop asking for an output and start specifying a controlled process.&lt;/p&gt;
&lt;p&gt;A prompt asks for an answer.&lt;br&gt;A loop defines a way of working until a condition is met.&lt;/p&gt;
&lt;p&gt;That's the useful distinction.&lt;/p&gt;
&lt;h2&gt;What a prompt is&lt;/h2&gt;
&lt;p&gt;A prompt is an instruction given to a model to produce an output or take an action. It can be simple:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Summarise this article in five points.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It can be more specific:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Summarise this article for a technical audience, separating the main thesis, arguments, limitations and conclusions. Don't add external information.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And it can be very well designed:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Analyse this article for a technical audience.
Goal: extract the thesis, assess the strength of the arguments and detect possible exaggerations.
Format: executive summary, critical analysis and practical recommendations.
Quality criteria: don't invent information, separate facts from opinion, and flag uncertainty when the text provides no evidence.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That last one is no longer "write something." It's a work specification. But it's still a normal prompt if the model runs once, delivers the result and stops.&lt;/p&gt;
&lt;h2&gt;What a loop is&lt;/h2&gt;
&lt;p&gt;A loop is a cycle of work that repeats until an exit condition is met.&lt;/p&gt;
&lt;p&gt;Its basic structure is:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;goal → action → verification → correction → repetition → stop&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;An example in development:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Goal: get CI green.
Action: read logs, reproduce the failure, fix the root cause.
Verification: check the latest CI run.
Correction: if it fails, diagnose again.
Stop: finish when the latest run is success, or when the max iterations are reached.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This is no longer a simple request for an output. It's a task with feedback.&lt;/p&gt;
&lt;p&gt;The critical part is verification. Without verification there is no reliable loop. There's a long conversation.&lt;/p&gt;
&lt;p&gt;A serious loop needs to know when it's done. "Do it better" doesn't work. "The tests pass, the lint is clean and the API contract doesn't change" does.&lt;/p&gt;
&lt;h2&gt;Not every iterative prompt is a loop&lt;/h2&gt;
&lt;p&gt;Many prompts look like loops but aren't.&lt;/p&gt;
&lt;p&gt;This is not a loop:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Improve the code and make sure everything is fine.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Neither is this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Review your answer before replying.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;They're good intentions, not verifiable cycles.&lt;/p&gt;
&lt;p&gt;This one starts to be a loop:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Fix the authentication module.
After each attempt, run:
npm test -- tests/auth
npm run lint
If a command fails, read the error, diagnose the root cause and fix it.
Don't finish until both commands pass or you reach 6 iterations.
Don't disable tests, don't soften lint rules, and don't change the public API.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Here there's already a goal, check commands, behaviour on failure, constraints and a stop.&lt;/p&gt;
&lt;h2&gt;How to tell if a prompt creates a loop&lt;/h2&gt;
&lt;p&gt;A prompt creates a loop if it contains these elements:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;A verifiable goal.&lt;/li&gt;
&lt;li&gt;An initial action.&lt;/li&gt;
&lt;li&gt;An objective check.&lt;/li&gt;
&lt;li&gt;A rule to decide whether to continue.&lt;/li&gt;
&lt;li&gt;A rule to correct if it fails.&lt;/li&gt;
&lt;li&gt;A limit on iterations, time or cost.&lt;/li&gt;
&lt;li&gt;A stop condition.&lt;/li&gt;
&lt;li&gt;Guardrails against fake shortcuts.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;If verification is missing, it's not a loop.&lt;/p&gt;
&lt;p&gt;If repetition is missing, it's not a loop.&lt;/p&gt;
&lt;p&gt;If the stop condition is missing, it's not a controlled loop; it's drift.&lt;/p&gt;
&lt;p&gt;If the model can "approve itself" with no external evidence, the loop is weak.&lt;/p&gt;
&lt;h2&gt;Normal prompt vs prompt-loop&lt;/h2&gt;
&lt;p&gt;The difference isn't length. It's the nature of the task.&lt;/p&gt;
&lt;p&gt;Use a normal prompt when the task is direct, bounded and easy to review:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Summarise this text.
Translate this paragraph.
Give me three alternative titles.
Explain this concept.
Rewrite this email in a professional tone.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Use a prompt-loop when the task depends on iteration, tools or verification:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Fix a bug until the tests pass.
Research a topic until the claims are validated against sources.
Improve coverage until it reaches 85%.
Refactor without breaking the public contract.
Generate a proposal and review it against a rubric.
Watch CI until it's green.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The simple rule:&lt;/p&gt;
&lt;p&gt;If you can evaluate the result in one pass, use a normal prompt.&lt;br&gt;If the work needs to try, measure, correct and repeat, use a loop.&lt;/p&gt;
&lt;h2&gt;The common mistake: thinking the loop is the command&lt;/h2&gt;
&lt;p&gt;Some tools have commands like &lt;code&gt;/goal&lt;/code&gt;, &lt;code&gt;/loop&lt;/code&gt; or similar mechanisms. They're useful because they keep the cycle alive without the human having to relaunch instructions by hand.&lt;/p&gt;
&lt;p&gt;But the command is not the loop.&lt;/p&gt;
&lt;p&gt;The command only executes the repetition.&lt;/p&gt;
&lt;p&gt;The real loop lives in the specification:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;what you want to achieve
how it's checked
what's allowed to be touched
what's forbidden
what to do on failure
when to stop
what to report at the end&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Without that specification, a &lt;code&gt;/goal&lt;/code&gt; is just a faster way to repeat a bad prompt.&lt;/p&gt;
&lt;h2&gt;The foundation is still a good prompt&lt;/h2&gt;
&lt;p&gt;There's an exaggerated narrative that says: "stop prompting, start designing loops." It's a catchy line, but imprecise.&lt;/p&gt;
&lt;p&gt;The correct version would be:&lt;/p&gt;
&lt;p&gt;"Stop writing vague prompts. Design prompts with a goal, context, constraints, acceptance criteria and, when needed, verifiable repetition."&lt;/p&gt;
&lt;p&gt;A loop contains prompts. It doesn't replace them.&lt;/p&gt;
&lt;p&gt;The prompt defines intent.&lt;br&gt;Verification defines truth.&lt;br&gt;The loop defines repetition.&lt;br&gt;Constraints protect the system.&lt;br&gt;The stop condition keeps the agent from drifting.&lt;/p&gt;
&lt;h2&gt;Anatomy of a good prompt&lt;/h2&gt;
&lt;p&gt;A good prompt isn't necessarily long. It's precise. It should answer seven questions.&lt;/p&gt;
&lt;h3&gt;1. What you want to achieve&lt;/h3&gt;
&lt;p&gt;Bad:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Improve this code.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Better:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Reduce the complexity of the authentication module without changing its public API.&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;2. What "done" means&lt;/h3&gt;
&lt;p&gt;Bad:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Make sure it works.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Better:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Done means:
- the auth tests pass;
- the lint passes;
- no public endpoint changes;
- no new dependencies are added.&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;3. What context it should use&lt;/h3&gt;
&lt;p&gt;Bad:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Keep the project in mind.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Better:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Before touching code, read:
- ARCHITECTURE.md
- RULES.md
- tests/auth/*
- the README of the affected module&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;4. What constraints it cannot break&lt;/h3&gt;
&lt;p&gt;Bad:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Don't break anything.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Better:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Don't change migrations, public contracts, event names, environment variables or CI configuration.&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;5. How it should verify&lt;/h3&gt;
&lt;p&gt;Bad:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Run some tests.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Better:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Run:
npm test -- tests/auth
npm run lint
npm run typecheck&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;6. What to do on failure&lt;/h3&gt;
&lt;p&gt;Bad:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Fix errors.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Better:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;If a check fails, identify the first causal failure, reproduce it locally when possible, and fix the root cause. Don't make cosmetic changes until the main failure is resolved.&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;7. When it should stop&lt;/h3&gt;
&lt;p&gt;Bad:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Keep going until it's done.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Better:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Stop when all checks pass or when you reach 8 iterations. If it isn't resolved, deliver the blockers with evidence.&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;Anatomy of a prompt-loop&lt;/h2&gt;
&lt;p&gt;A good loop template might look like this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Start the "[loop name]" loop.
Goal:
[verifiable result]
Max iterations:
[N]
Context:
[files, documentation, business rules, sources or constraints to read]
Between iterations run:
[command, test, evaluation, rubric or check]
Exit when:
[objective success condition]
Procedure:
1. Discover the current state.
2. Identify the root cause or next reasonable action.
3. Apply the minimal necessary change.
4. Verify.
5. If it fails, use the feedback to correct.
6. If it passes, stop.
Guardrails:
- Don't change the success criterion.
- Don't disable checks.
- Don't hide errors.
- Don't replace a fix with a bypass.
- Don't make changes outside the scope.
- If there's a real blocker, stop and report it.
Final report:
- what was done;
- what was verified;
- what evidence proves success;
- what risks remain.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This structure is reusable because it turns an open-ended request into a controlled unit of work.&lt;/p&gt;
&lt;h2&gt;Example: normal prompt&lt;/h2&gt;
&lt;pre&gt;&lt;code&gt;Fix the CI on this branch.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The model can do many things. Some good, some dangerous. It can look at one error, touch a config, change a test, or assume it's already fixed. The problem isn't that the prompt is short. The problem is that it doesn't define what counts as success or which shortcuts are forbidden.&lt;/p&gt;
&lt;h2&gt;Example: prompt-loop&lt;/h2&gt;
&lt;pre&gt;&lt;code&gt;Start the "Fix CI Until Green" loop.
Goal:
The latest CI run on the current branch passes.
Max iterations:
8
Context:
Use the current branch. Inspect the latest failed GitHub Actions run before editing code.
Between iterations run:
gh run list --branch $(git branch --show-current) --limit 1 --json conclusion -q '.[0].conclusion'
Exit when:
The latest run conclusion is success.
Procedure:
1. Find the latest failed CI run.
2. Read the failing job logs.
3. Identify the root cause.
4. Reproduce locally when possible.
5. Apply the smallest safe fix.
6. Run the relevant local check.
7. Push.
8. Re-check CI.
Guardrails:
- Do not modify the check command or exit criteria.
- Do not skip, disable or weaken checks.
- Do not mark tests as skipped.
- Do not hide failures behind broad catch blocks.
- Do not change unrelated files.
- If blocked by credentials, unavailable services or missing secrets, stop and report evidence.
Final report:
Summarize root cause, files changed, checks run, CI status and remaining risks.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This isn't magic. It's good prompting in the shape of a loop.&lt;/p&gt;
&lt;h2&gt;What separates a good loop from a bad one&lt;/h2&gt;
&lt;p&gt;A bad loop repeats without criteria.&lt;br&gt;A good loop learns from each failure.&lt;/p&gt;
&lt;p&gt;A bad loop optimises the appearance of success.&lt;br&gt;A good loop protects the acceptance criterion.&lt;/p&gt;
&lt;p&gt;A bad loop lets the model decide whether its work is fine.&lt;br&gt;A good loop uses tests, commands, rubrics, sources or external review.&lt;/p&gt;
&lt;p&gt;A bad loop doesn't know when to stop.&lt;br&gt;A good loop has limits.&lt;/p&gt;
&lt;h2&gt;Guardrails: the least flashy and most important part&lt;/h2&gt;
&lt;p&gt;Guardrails stop the agent from winning by cheating.&lt;/p&gt;
&lt;p&gt;In code, an agent can pass tests by disabling them.&lt;br&gt;In research, it can pick only convenient sources.&lt;br&gt;In content, it can satisfy a shallow rubric by producing generic text.&lt;br&gt;In security, it can hide symptoms without fixing the root cause.&lt;/p&gt;
&lt;p&gt;That's why a loop must explicitly say what it cannot sacrifice.&lt;/p&gt;
&lt;p&gt;Example:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Don't reduce coverage.
Don't delete tests.
Don't change public contracts.
Don't ignore errors.
Don't change the success metric.
Don't replace validation with claims.
Don't invent sources.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Guardrails are part of the prompt. They're not decoration.&lt;/p&gt;
&lt;h2&gt;Verification: the centre of the loop&lt;/h2&gt;
&lt;p&gt;Verification can take several forms.&lt;/p&gt;
&lt;p&gt;In software:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;npm test
npm run lint
npm run typecheck
pytest
cargo test
go test ./...&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;In research:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Every critical claim must be backed by a source.
Separate facts, inferences and opinion.
Contrast contradictions between sources.
Flag uncertainty.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;In content:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Evaluate against a rubric:
- clarity;
- accuracy;
- usefulness;
- tone;
- evidence;
- absence of filler.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;In business:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Validate that the proposal:
- respects legal constraints;
- fits the ICP;
- includes next steps;
- doesn't promise capabilities that don't exist.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The important thing is that verification doesn't rely only on the model saying "this is fine."&lt;/p&gt;
&lt;h2&gt;When not to use loops&lt;/h2&gt;
&lt;p&gt;Don't use loops for everything.&lt;/p&gt;
&lt;p&gt;A loop adds cost, time and complexity. It's absurd to use one for simple tasks.&lt;/p&gt;
&lt;p&gt;You don't need a loop to:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;summarise an email;
translate a paragraph;
generate a list of ideas;
explain a concept;
turn notes into an outline;
rephrase a sentence.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Loops are also a bad fit when verification is purely subjective and you haven't defined criteria. "Make it prettier until it's perfect" is not a reliable loop. It's a factory of vague iterations.&lt;/p&gt;
&lt;p&gt;Before using a loop, ask yourself:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Is there an objective success condition?
Is there a way to measure progress?
Does failure produce useful feedback?
Does the task justify several iterations?
Are there clear limits?&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;If the answer is no, use a normal prompt.&lt;/p&gt;
&lt;h2&gt;When to use loops&lt;/h2&gt;
&lt;p&gt;Use loops when the work is verifiable and feedback improves the result.&lt;/p&gt;
&lt;p&gt;Clear cases:&lt;/p&gt;
&lt;h3&gt;Development&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;Fix CI.
Raise coverage.
Fix bugs.
Refactor with tests.
Migrate code while keeping compatibility.&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;Security&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;Fix SAST findings.
Review vulnerable dependencies.
Harden configuration.
Validate controls against a checklist.&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;Research&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;Compare providers.
Analyse incidents.
Review literature.
Build a report with verified sources.&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;Documentation&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;Update docs until they match the real API.
Generate a changelog from commits and PRs.
Detect inconsistencies between README, code and examples.&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;Product&lt;/h3&gt;
&lt;pre&gt;&lt;code&gt;Review a spec against requirements.
Validate user stories.
Detect ambiguities before development starts.&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;Memory turns loops into learning&lt;/h2&gt;
&lt;p&gt;A single-session loop can solve a task. A loop with memory can improve over time.&lt;/p&gt;
&lt;p&gt;Memory shouldn't be a junk drawer of notes. It should separate:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;TRIED:
What was attempted and what happened.
VERIFIED:
What facts were confirmed with evidence.
OPEN:
What's still pending or couldn't be resolved.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Example:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;# MEMORY.md
## TRIED
- Tried to update eslint-config. Failed due to a conflict with TypeScript 5.4.
## VERIFIED
- The "frontend-lint" job uses Node 20.
- The CI failure reproduces with npm run lint.
- The broken rule comes from packages/web/.eslintrc.
## OPEN
- Check whether to migrate the shared config in a separate PR.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The rule is simple: don't store hunches as facts. Bad memory amplifies mistakes. Good memory reduces repeated work.&lt;/p&gt;
&lt;h2&gt;A good prompt isn't longer: it's more contractual&lt;/h2&gt;
&lt;p&gt;A common mistake is believing that a quality prompt has to be enormous. Not necessarily.&lt;/p&gt;
&lt;p&gt;A good prompt is contractual: it defines obligations, limits and evidence.&lt;/p&gt;
&lt;p&gt;Bad:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Act as a senior expert, analyse deeply, think step by step and give me the best possible solution.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Better:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Goal: identify the root cause of the CI failure.
Use evidence from the logs.
Don't edit code until you isolate the first causal error.
Deliver: root cause, affected file, proposed fix and verification command.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The second is better because it reduces ambiguity. It doesn't need grandiosity. It needs control.&lt;/p&gt;
&lt;h2&gt;Checklist for a quality prompt&lt;/h2&gt;
&lt;p&gt;Before launching an important prompt, review this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Is the goal clear?
Is the definition of done verifiable?
Does the model have enough context?
Are there explicit constraints?
Is the output format defined?
Are there examples if the task is ambiguous?
Are facts separated from inferences?
Are the allowed tools or sources stated?
Are dangerous shortcuts forbidden?
Is there a stop criterion?&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;For a normal prompt, you don't need every point.&lt;br&gt;For a prompt-loop, almost all of them should be present.&lt;/p&gt;
&lt;h2&gt;Specific checklist for loops&lt;/h2&gt;
&lt;pre&gt;&lt;code&gt;Is there a verifiable goal?
Is there a command, test, metric, source or rubric?
Is there a max number of iterations?
Is it clear what to do on failure?
Is it clear when to stop?
Is changing the metric forbidden?
Is skipping checks forbidden?
Is there a final report with evidence?
Is there escalation if a real blocker appears?
Is there memory if the work spans several sessions?&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;If you can't complete this checklist, you probably don't need a loop — or you haven't defined the work well yet.&lt;/p&gt;
&lt;h2&gt;Prompt engineer vs loop engineer&lt;/h2&gt;
&lt;p&gt;The useful distinction isn't a job title. It's a focus.&lt;/p&gt;
&lt;p&gt;A prompt engineer optimises instructions.&lt;br&gt;A loop engineer optimises feedback.&lt;/p&gt;
&lt;p&gt;But they're not separate worlds. A well-designed loop is a mature form of prompting.&lt;/p&gt;
&lt;p&gt;The practical difference is this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Prompt:
"Write a function that does X."
Prompt-loop:
"Write a function that does X, run tests Y, fix until they pass, don't change API Z, and stop after N attempts if there's a blocker."&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The second isn't more sophisticated for fashion's sake. It's more useful because it defines how the work is checked.&lt;/p&gt;
&lt;h2&gt;The conclusion, without smoke&lt;/h2&gt;
&lt;p&gt;Loops don't kill the prompt. They make it more important.&lt;/p&gt;
&lt;p&gt;The more autonomy you give an agent, the better defined the initial prompt has to be. If the model is going to run commands, edit files, query systems or repeat actions, a vague instruction stops being an inconvenience and becomes a risk.&lt;/p&gt;
&lt;p&gt;The skill is not saying "use a loop."&lt;br&gt;The skill is designing a specification that says:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;what has to be achieved;
how it's verified;
what can't break;
how failure is corrected;
when it ends;
when it stops;
what evidence remains.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;For simple tasks, use simple prompts.&lt;/p&gt;
&lt;p&gt;For iterative tasks, use loops.&lt;/p&gt;
&lt;p&gt;For critical tasks, don't trust the model's goodwill: use external verifiers, limits, logs, memory and human review where appropriate.&lt;/p&gt;
&lt;p&gt;The prompt isn't dead. It leveled up. It's no longer just a sentence to get an answer. In complex work, it's the contract that governs an execution system.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>C1b3rwall: LLMs as a new criminal surface</title>
    <link href="https://yous.dev/blog/c1b3rwall-llm-security"/>
    <id>https://yous.dev/blog/c1b3rwall-llm-security</id>
    <published>2026-06-03T00:00:00Z</published>
    <updated>2026-06-03T00:00:00Z</updated>
    <category term="talks"/>
    <summary>Notes from presenting "Security in LLMs: the new criminal surface" — what the audience already feared, and what surprised them.</summary>
    <content type="html">&lt;p class="post-sub"&gt;I gave a talk at C1b3rwall 2026 — "Security in LLMs: the new criminal surface." Some notes on what the room already feared, and what surprised it.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/c1b3rwall.jpg" alt="Presenting Security in LLMs at C1b3rwall 2026" loading="lazy" width="1500" height="1029"&gt;
&lt;figcaption&gt;On stage at C1b3rwall 2026.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The premise was simple: every time we hand a language model more reach — tools, memory, the ability to act — we also hand attackers a surface that didn't exist a year ago. Not a metaphorical one. A real, reproducible, exploitable surface, with its own bestiary of techniques.&lt;/p&gt;
&lt;p&gt;What a security audience already expects to hear is "prompt injection." They've read the headlines. What tends to land harder is the second step: an injected instruction isn't interesting because it makes the model say something rude — it's interesting because the model can &lt;em&gt;do&lt;/em&gt; things. Read a file. Call an API. Trust the wrong document. The danger scales exactly with the autonomy.&lt;/p&gt;
&lt;blockquote&gt;The model isn't the vulnerability. The vulnerability is everything you let the model touch on your behalf.&lt;/blockquote&gt;
&lt;h2&gt;Three things that surprised the room&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;How little it takes. A poisoned web page or a crafted document is often enough — no exotic exploit, just text in the right place at the right time.&lt;/li&gt;
&lt;li&gt;How familiar the defenses feel. Least privilege, input validation, separating data from instructions, never trusting the client — the classics, wearing new clothes.&lt;/li&gt;
&lt;li&gt;How quickly "the AI did it" becomes an accountability gap. If a system can take consequential actions, "the model decided" is not an answer an investigator accepts.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The part I cared most about getting across is that this isn't a reason to stop. It's a reason to design. The same care that makes a SOC trustworthy makes an agent trustworthy: explicit goals, verified actions, hard limits on what can be touched, and evidence left behind. The attack surface is new. The discipline isn't.&lt;/p&gt;
&lt;p&gt;Thanks to everyone who came, asked sharp questions, and stayed afterwards to argue. That's the part you can't get from a slide deck.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.linkedin.com/posts/jose-antonio-cordon-munoz_c1b3rwall2026-aisecurity-ugcPost-7467837436451008512-Jkke" target="_blank" rel="noopener"&gt;The talk on LinkedIn ↗&lt;/a&gt;&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>Two small tools: RenderMD and Share</title>
    <link href="https://yous.dev/blog/rendermd-and-share"/>
    <id>https://yous.dev/blog/rendermd-and-share</id>
    <published>2026-05-20T00:00:00Z</published>
    <updated>2026-05-20T00:00:00Z</updated>
    <category term="tools"/>
    <summary>Software that needs no account and barely a server — a Markdown report generator and a zero-knowledge sharing service, and the moment they clicked together.</summary>
    <content type="html">&lt;p class="post-sub"&gt;The best personal tools ask for nothing — no account, no subscription, no server if a file will do. Two recent ones follow that rule closely.&lt;/p&gt;
&lt;p&gt;Most of what I build for myself has a narrow mouth and a deep throat: trivial to start, but still useful as the problem grows. It reads plain formats and writes plain formats. It survives being ignored for a year and still works when you come back. These two are good examples.&lt;/p&gt;
&lt;h2&gt;RenderMD — Markdown into real reports&lt;/h2&gt;
&lt;p&gt;I write almost everything in Markdown, but Markdown looks like a draft. &lt;a href="https://rendermd.yous.dev/" target="_blank" rel="noopener"&gt;RenderMD&lt;/a&gt; turns a &lt;code&gt;.md&lt;/code&gt; file into a self-contained HTML technical report: cover page, metadata pulled from the frontmatter, custom colours and code highlighting, ready to print to PDF.&lt;/p&gt;
&lt;p&gt;The important part is what it &lt;em&gt;doesn't&lt;/em&gt; need. You open &lt;code&gt;app.html&lt;/code&gt; in a browser — no server, no install — drag your file in, adjust the metadata it detected, and download. The frontmatter does the talking:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;---
title: "Report title"
subtitle: "Short description"
author: "Name Surname"
date: "2026-05-20"
classification: "Internal use"
---&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;For automation there's also a small &lt;code&gt;pandoc&lt;/code&gt; pipeline for batch generation in CI. But the point of the project is the browser flow: a tool that needs the cloud to render a Markdown file has confused its own importance.&lt;/p&gt;
&lt;h2&gt;Share — give something away without giving it to a server&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://share.yous.dev/" target="_blank" rel="noopener"&gt;Share&lt;/a&gt; is a zero-knowledge sharing service running entirely on Cloudflare Workers and KV. You encrypt in the browser with AES-GCM; the key never touches the server — it travels in the URL hash, the part the browser never sends:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;share.yous.dev/v/aZ8c3K1q#k=key-never-touches-the-server&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;On top of that: an optional password layer (PBKDF2, 250k iterations), a configurable TTL — one hour, a day, a week, or burn-after-read — and a small &lt;strong&gt;public API&lt;/strong&gt; so other tools can use it. The server stores an opaque blob and knows nothing about what's inside. That's the whole idea: the simplest design that makes "I can't read your data" a property of the architecture, not a promise.&lt;/p&gt;
&lt;blockquote&gt;If it can't run offline on a slow machine, it isn't a personal tool — it's someone else's product wearing your data.&lt;/blockquote&gt;
&lt;h2&gt;Where the two meet&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://rendermd.yous.dev/" target="_blank" rel="noopener"&gt;RenderMD&lt;/a&gt; integrates &lt;a href="https://share.yous.dev/" target="_blank" rel="noopener"&gt;Share&lt;/a&gt;'s public API directly — partly as a proof that the API actually works for outside tools, and partly because it solves a real annoyance. A self-contained &lt;code&gt;.html&lt;/code&gt; file is great until you try to &lt;em&gt;send&lt;/em&gt; it: handing someone a raw HTML file, especially on a phone, often doesn't render the way you'd hope. So RenderMD can push the report through Share and hand back a link instead.&lt;/p&gt;
&lt;p&gt;Open that link on anything — a phone with no laptop in sight — and the report renders perfectly, because the rendering happens in the browser and every browser is compatible. No "download this file and hope it opens." Just a URL that works everywhere, with Share's encryption underneath.&lt;/p&gt;
&lt;p&gt;Neither of these ships as a polished SaaS, because the interesting part is the system, not a funnel around it. Small surface, no telemetry, no server unless the problem genuinely needs one — and when two small tools click together like this, that's the whole reward. Field notes over case studies.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>47CON: reversing a WiFi drone</title>
    <link href="https://yous.dev/blog/47con-drones"/>
    <id>https://yous.dev/blog/47con-drones</id>
    <published>2026-04-19T00:00:00Z</published>
    <updated>2026-04-19T00:00:00Z</updated>
    <category term="talks"/>
    <summary>Capturing the traffic, deciphering the protocol, and writing our own control app — until a cheap WiFi drone flew itself with YOLO.</summary>
    <content type="html">&lt;p class="post-sub"&gt;A field note from 47CON, where we took a cheap WiFi-controlled drone apart — not the radio, the protocol — and ended up flying it autonomously with a computer.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/47con.jpg" alt="Reversing a WiFi drone at 47CON" loading="lazy" width="960" height="872"&gt;
&lt;figcaption&gt;47CON — hardware on the table.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This drone didn't talk over some exotic RF link. It did something much more familiar and much more interesting: it exposed its own WiFi network, and the phone app controlled it over that. Which means the entire control channel was just network traffic — and network traffic can be captured, read, and replayed.&lt;/p&gt;
&lt;h2&gt;Capture, decipher, rebuild&lt;/h2&gt;
&lt;p&gt;The work was a clean reversing loop. First we captured the traffic between the official app and the drone and stared at the packets until the structure gave way. Then we deciphered the control protocol — which bytes meant throttle, yaw, pitch, takeoff, land. Once the protocol was legible, we stopped needing the app at all.&lt;/p&gt;
&lt;p&gt;So we wrote our own. A small control application that ran on a laptop and exposed an &lt;strong&gt;open API&lt;/strong&gt;: each endpoint translated into the corresponding command and was sent straight to the drone. Suddenly the aircraft answered to HTTP instead of a phone.&lt;/p&gt;
&lt;blockquote&gt;Once you can speak a device's protocol, the app it shipped with is just one client among many — and not the privileged one.&lt;/blockquote&gt;
&lt;h2&gt;Then it flew itself&lt;/h2&gt;
&lt;p&gt;An open API is an invitation. With the drone reachable by code, we put a vision model in the loop: a &lt;strong&gt;YOLO&lt;/strong&gt; object detector reading the camera feed, deciding where to go, and driving the API directly. The same cheap toy that came with a clumsy phone app was now doing autonomous, AI-driven flight the manufacturer never designed for — tracking and reacting on its own.&lt;/p&gt;
&lt;h2&gt;The point isn't the drone&lt;/h2&gt;
&lt;p&gt;The lesson generalises to most of the cheap connected world. A WiFi control channel is a public, often unauthenticated surface; the moment its protocol is understood, control belongs to whoever decodes it. The cost pressure that makes a fun gadget is the same pressure that ships an open channel and trusts the bundled app to be the only thing talking to it.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Treat the WiFi channel as public. Anyone in range can listen.&lt;/li&gt;
&lt;li&gt;If commands aren't authenticated, the "owner" is just the first client to speak the protocol.&lt;/li&gt;
&lt;li&gt;An open, undocumented API is an open API. Someone will write the second client.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Thanks to 47CON for the stage, and to everyone who came up afterwards with their own drones and their own captures. The best conversations always happen with hardware on the table.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.linkedin.com/posts/jose-antonio-cordon-munoz_cybersecurity-reversing-drones-activity-7451622469607706624-zCKI" target="_blank" rel="noopener"&gt;The talk on LinkedIn ↗&lt;/a&gt;&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>The bicycle as the ultimate machine</title>
    <link href="https://yous.dev/blog/the-ultimate-machine"/>
    <id>https://yous.dev/blog/the-ultimate-machine</id>
    <published>2026-03-01T00:00:00Z</published>
    <updated>2026-03-01T00:00:00Z</updated>
    <category term="misc"/>
    <summary>A manifesto on efficiency, minimalism and freedom — and why the cleanest vehicle ever built is also the best design lesson I know.</summary>
    <content type="html">&lt;p class="post-sub"&gt;A manifesto on efficiency, minimalism, and freedom.&lt;/p&gt;
&lt;figure&gt;
&lt;img src="https://yous.dev/assets/bike.jpg" alt="A stripped-down bicycle" loading="lazy" width="1299" height="975"&gt;
&lt;figcaption&gt;The cleanest machine I own.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;There is a famous quote by Steve Jobs where he describes the computer as &lt;em&gt;"a bicycle for our minds."&lt;/em&gt; He was fascinated by a study showing that when a human climbs onto a bicycle, they instantly outpace every other species on Earth in terms of locomotive efficiency.&lt;/p&gt;
&lt;p&gt;We became the most efficient machines on the planet just by adding two wheels, a chain, and a pair of pedals.&lt;/p&gt;
&lt;p&gt;For me, the bicycle is not just a sport or a weekend hobby. It is a philosophy. It is design thinking brought to life in the physical world. In a society obsessed with adding more noise, heavier engines, and endless digital clutter, turning to the bicycle is the ultimate act of modern hacking.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;1. Zero Emissions, Maximum Autonomy&lt;/h2&gt;
&lt;p&gt;We live in a world where we are constantly told that comfort means sitting in a two-ton metal box, burning fuel, and waiting in line behind fifty other identical metal boxes just to move a couple of kilometres across town.&lt;/p&gt;
&lt;p&gt;The bicycle shatters that illusion. It is the cleanest vehicle ever created:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;0% Emissions:&lt;/strong&gt; Powered entirely by human energy — your morning coffee converted into instant, clean momentum.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;100% Efficiency:&lt;/strong&gt; No traffic jams, no parking hunts, no hidden taxes, and no unexpected mechanic bills that drain your wallet.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When you commute by bike, you realise that urban speed limits are a joke. A simple single-speed or fixed-gear bicycle can maintain a more consistent average speed through city streets than a sports car trapped in rush hour. I got tired of arguing about it, so I built &lt;a href="https://routebattle.yous.dev/" target="_blank" rel="noopener"&gt;RouteBattle&lt;/a&gt; to settle it with data — race a bike against a car across a real city and watch the podium. The data doesn't lie: bikes don't change the system; they completely bypass it.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;2. The Joy of Analogue Mechanics&lt;/h2&gt;
&lt;p&gt;In my day-to-day work, I build digital products, solve complex logical problems, and interface with cutting-edge software. But there is a unique, grounding satisfaction in stepping away from the screens, heading down to the garage, and getting your hands dirty with real, physical mechanics.&lt;/p&gt;
&lt;p&gt;There is a beautiful simplicity in a stripped-down bicycle. One chain. One chainring. One cog.&lt;/p&gt;
&lt;p&gt;Taking ten minutes to clean the drivetrain, lubricate the links, and tune the tension is almost meditative. It is a reminder of how high-performance design doesn't need to be over-engineered. True elegance is not about adding more features; it's about having nothing left to strip away.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;3. "Take the Long Way Home"&lt;/h2&gt;
&lt;p&gt;Choosing the bicycle changes how you experience your surroundings. A commute stops being a mindless chore you just want to get through, and becomes an active part of your day.&lt;/p&gt;
&lt;p&gt;When you are on two wheels, you notice the city. You feel the evening breeze at the end of a long day, you discover hidden streets, you ride along the river paths, and you intentionally take the long way home just because the flow feels right. It leaves your mind completely clear, your focus sharp, and your energy levels charged.&lt;/p&gt;
&lt;p&gt;Whether it's a quick two-minute dash to your favourite local coffee shop, the daily commute, or a long-distance bikepacking trip exploring new horizons with nothing but a few bags strapped to your frame, the bike gives you absolute control over your journey.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Join the Ride&lt;/h2&gt;
&lt;p&gt;You don't need an expensive carbon machine or professional gear to start hacking your daily mobility. All it takes is an old frame, a solid lock, and a shift in mentality.&lt;/p&gt;
&lt;p&gt;The next time you need to move across town, leave the keys behind. Step outside, hop on a saddle, push the pedals, and feel the immediate superpower of moving freely, efficiently, and sustainably.&lt;/p&gt;
&lt;p&gt;Your mind, your city, and the planet will thank you for it. &lt;strong&gt;Welcome to the clean movement.&lt;/strong&gt;&lt;/p&gt;</content>
  </entry>
</feed>
