<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Agentic Engineering]]></title><description><![CDATA[Where the hype ends, and the build begins]]></description><link>https://newsletter.agentengineering.co</link><image><url>https://substackcdn.com/image/fetch/$s_!J5ef!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fecfd107e-a784-4492-be0c-1ade75f87224_256x256.png</url><title>Agentic Engineering</title><link>https://newsletter.agentengineering.co</link></image><generator>Substack</generator><lastBuildDate>Wed, 16 Sep 2026 15:16:59 GMT</lastBuildDate><atom:link href="https://newsletter.agentengineering.co/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Packt]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[agenticengineering@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[agenticengineering@substack.com]]></itunes:email><itunes:name><![CDATA[Packt]]></itunes:name></itunes:owner><itunes:author><![CDATA[Packt]]></itunes:author><googleplay:owner><![CDATA[agenticengineering@substack.com]]></googleplay:owner><googleplay:email><![CDATA[agenticengineering@substack.com]]></googleplay:email><googleplay:author><![CDATA[Packt]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[🎙️ Episode 4: Forward Deployed Engineer, the job everyone's hiring for and no one can quite explain]]></title><description><![CDATA[Google's Tanya Dixit breaks down what the role involves, once you get past the title]]></description><link>https://newsletter.agentengineering.co/p/episode-4-forward-deployed-engineer</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/episode-4-forward-deployed-engineer</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Tue, 15 Sep 2026 15:07:32 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/215777479/3d285e48b99cbcd61f018390ebe12694.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Every so often the industry invents a job title, throws it at LinkedIn, and leaves the rest of us to figure out what it means. This year&#8217;s contender: <em>forward deployed engineer</em>. Half the internet insists it&#8217;s consulting with better branding. The other half seems convinced it&#8217;s the last job standing once AI writes all our code for us. I had opinions. I also had questions. </p><p>So for episode 4, I got <strong><a href="https://www.linkedin.com/in/tanya-dixit-computer-vision/">Tanya Dixit</a></strong> on the podcast, a forward deployed engineer at Google, to sort out which half is right (spoiler: neither, entirely). Yes, <em>also</em> called Tanya. No, we did not plan that but the transcript was a nightmare to edit.</p><p>Tanya&#8217;s route here was not a tidy straight line. She started out building satellite hardware in embedded systems, back before deep learning was cool enough to have its own hoodie. Then the math from her undergrad started clicking, she did a nano-degree, and somehow that spiralled into a master&#8217;s from ANU (university medal, no big deal) and a career that&#8217;s touched everything from enterprise AI to bushfire risk modelling. By the time she landed on forward deployed engineering, she&#8217;d basically done a full lap of the ML lifecycle. Build it, prove it, deploy it, fix it, repeat.</p><p>This is why the conversation doesn&#8217;t stay in the shallow end. We get into what her actual week looks like, and the coding-to-customer-calls ratio is not what most people guess. If you assume &#8220;forward deployed&#8221; is code for &#8220;technical enough to be in the room but not enough to be at the keyboard&#8221;... you&#8217;re in for a surprise.</p><div class="callout-block" data-callout="true"><p>Funnily enough, if this episode leaves you wanting to try the job rather than just hear about it, Tanya&#8217;s got you covered there too. On 19th September, she&#8217;s co-instructing a live workshop called <em><strong>Forward Deployed Engineering: From AI Demo to Production</strong></em> alongside <strong><a href="https://www.linkedin.com/in/keithbourne/">Keith Bourne</a></strong> (FDE at Tribe AI, author of <a href="https://www.amazon.com/Unlocking-Data-Generative-RAG-fundamentals/dp/1806381656/ref=sr_1_1?crid=2XYUOZ3539JON&amp;dib=eyJ2IjoiMSJ9.8qFnMqSH0Gq-G7SJbgcTYtiTT7RyvP2J3D2iJWzJh8nd5-CumwjJIMbpktOb6EA1TIzPF4aHZx95x0XDDMLlxJKAvLUGQdccP4QPsLqQactNjf5a0gwM1EivykU_H6UDwowxOMwR0quMeNKjmchoytrLfi8mLsntMpHRlKf7UQChIaxdCEUtbWuamwkvk18zBSN6fXPr_epLN9k1g9Zxal9NcN49f9O8NKfM9-y337A.hMZrhCi7KBHXx7QGdhsczfMQD65X_jKQYph2BGQoLDU&amp;dib_tag=se&amp;keywords=Unlocking+Data+with+Generative+AI+and+RAG&amp;qid=1789455927&amp;sprefix=unlocking+data+with+generative+ai+and+rag+%2Caps%2C336&amp;sr=8-1">Unlocking Data with Generative AI and RAG</a>). </p><p>It&#8217;s not a webinar where you nod along and take notes. You&#8217;re handed a realistic 90-day AI-agent deployment for a regulated customer, and you have to scope it, define the success metrics, design the security controls, and then defend the whole plan in a CISO hot seat, which is exactly as fun as it sounds. </p><p>You&#8217;ll walk away with an actual skills gap analysis, a career action plan, and a much clearer sense of whether this job is for you or just looks good from the outside.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://luma.com/fde-workshop?utm_source=nl&quot;,&quot;text&quot;:&quot;Get Me Into the FDE Hot Seat &#8594;&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://luma.com/fde-workshop?utm_source=nl"><span>Get Me Into the FDE Hot Seat &#8594;</span></a></p></div><p>We also get properly into the thing everyone eyeing this career path wants to know: where does an FDE&#8217;s job stop and the product team&#8217;s job start? Tanya draws that line with far more precision than I expected going in, and her answer for why companies are suddenly hiring for this role like it&#8217;s going out of fashion made a few things click for me. Turns out standard SaaS playbooks and AI products are not playing the same game at all.</p><p>Then there&#8217;s the bit I haven&#8217;t stopped thinking about since we recorded. As models get better and coding gets, in her words, increasingly commoditised, does that mean the industry just needs fewer FDEs? She doesn&#8217;t reach for the easy optimist answer, and she doesn&#8217;t reach for the easy doom answer either. What she gives instead is a proper mental model for which parts of this job survive and which don&#8217;t... and yes, we did end up in recursive self-improvement territory, because apparently no 2026 tech conversation is complete without it.</p><p>If you're building AI products, working anywhere near enterprise customers, or just trying to work out where the humans still fit in this whole picture, do yourself a favour and watch the full thing rather than skimming for quotes.</p><p>The full episode is live now with audio and video. Go watch it, and if Tanya&#8217;s take here makes you want more, she&#8217;s finally starting to write on Substack too. </p><div class="callout-block" data-callout="true"><p>And if Tanya&#8217;s <em>Forward Deployed Engineering (FDE) Workshop: From AI Demo to Production</em> is still sitting in an open tab somewhere on your browser...</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://luma.com/fde-workshop?utm_source=nl&quot;,&quot;text&quot;:&quot;Save your seat immediately&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://luma.com/fde-workshop?utm_source=nl"><span>Save your seat immediately</span></a></p></div><p>That&#8217;s it for this one. We&#8217;ll pick up the conversation next week.</p><p>Until then, keep building.</p><blockquote><p>Tanya D&#8217;cruz<br><em>Editor-in-Chief</em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[Inside the harness (Part 3): Change the tools]]></title><description><![CDATA[How the same ReAct loop becomes a coding agent and what happens when its context and capabilities no longer fit]]></description><link>https://newsletter.agentengineering.co/p/inside-the-harness-part-3-change</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/inside-the-harness-part-3-change</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Fri, 11 Sep 2026 15:04:32 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/0f69415c-8f00-4971-8d96-1c8434c98dfa_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="callout-block" data-callout="true"><h4><code>Editor&#8217;s note</code></h4><p>Two Fridays ago, our agent checked the weather. Last week, we gave it a plan and taught the harness to tap it on the shoulder when it wandered. This week, the same loop gets dropped into a codebase.</p><p>In the final part of <em>Inside the Harness</em>, we see how a change of tools creates a coding agent, then deal with what happens when its context and capabilities no longer fit. The loop stays put. The harness grows around it.</p><p>This also concludes our Friday scheduling experiment. Normal service resumes next issue. But before we dive in&#8230;</p><div class="poll-embed" data-attrs="{&quot;id&quot;:1205577}" data-component-name="PollToDOM"></div></div><p>Leave the harness exactly where it is. Swap the tools, point it at a codebase, and the same loop becomes a coding agent.</p><h3><strong><span>Same harness, new tools</span></strong></h3><p><span>The experiment here is deliberately simple: keep the harness and its system prompt as they are, and change only the tools. If the harness skeleton is genuinely generic, that one swap should be enough to turn a research agent into a coding agent.</span></p><p><span>The system prompt, therefore, stays almost untouched:</span></p><div class="callout-block" data-callout="true"><p><code>&lt;agent name=&#8221;Tara&#8221;&gt;</code></p><p><code>You&#8217;re a helpful agent.</code></p><p><code>&lt;/agent&gt;</code></p><p><code>&lt;task&gt;</code></p><p><code>Your task is to use the given tools to solve the user&#8217;s problem.</code></p><p><code>&lt;/task&gt;</code></p><p><code>&lt;notes&gt;</code></p><p><code>- Parallel tool calling is supported.</code></p><p><code>&lt;/notes&gt;</code></p></div><p>Only the toolset changes. Out go the research tools; in come three tools that let the agent work directly with a codebase, defined here in JSON:</p><div class="callout-block" data-callout="true"><p><code>[</code></p><p><code>{</code></p><p><code>&#8220;name&#8221;: &#8220;read_file&#8221;,</code></p><p><code>&#8220;description&#8221;: &#8220;Read a file from disk.&#8221;,</code></p><p><code>&#8220;input_schema&#8221;: {</code></p><p><code>&#8220;type&#8221;: &#8220;object&#8221;,</code></p><p><code>&#8220;properties&#8221;: {</code></p><p><code>&#8220;path&#8221;: { &#8220;type&#8221;: &#8220;string&#8221; }</code></p><p><code>},</code></p><p><code>&#8220;required&#8221;: [&#8221;path&#8221;]</code></p><p><code>}</code></p><p><code>},</code></p><p><code>{</code></p><p><code>&#8220;name&#8221;: &#8220;write_file&#8221;,</code></p><p><code>&#8220;description&#8221;: &#8220;Write content to a file. Set is_appending=true to append instead of overwrite.&#8221;,</code></p><p><code>&#8220;input_schema&#8221;: {</code></p><p><code>&#8220;type&#8221;: &#8220;object&#8221;,</code></p><p><code>&#8220;properties&#8221;: {</code></p><p><code>&#8220;path&#8221;: { &#8220;type&#8221;: &#8220;string&#8221; },</code></p><p><code>&#8220;content&#8221;: { &#8220;type&#8221;: &#8220;string&#8221; },</code></p><p><code>&#8220;is_appending&#8221;: { &#8220;type&#8221;: &#8220;boolean&#8221;, &#8220;default&#8221;: false }</code></p><p><code>},</code></p><p><code>&#8220;required&#8221;: [&#8221;path&#8221;, &#8220;content&#8221;]</code></p><p><code>}</code></p><p><code>},</code></p><p><code>{</code></p><p><code>&#8220;name&#8221;: &#8220;bash&#8221;,</code></p><p><code>&#8220;description&#8221;: &#8220;Execute a shell command.&#8221;,</code></p><p><code>&#8220;input_schema&#8221;: {</code></p><p><code>&#8220;type&#8221;: &#8220;object&#8221;,</code></p><p><code>&#8220;properties&#8221;: {</code></p><p><code>&#8220;command&#8221;: { &#8220;type&#8221;: &#8220;string&#8221; }</code></p><p><code>},</code></p><p><code>&#8220;required&#8221;: [&#8221;command&#8221;]</code></p><p><code>}</code></p><p><code>}</code></p><p><code>]</code></p></div><p>That&#8217;s it! <code>read_file</code> inspects the codebase, <code>write_file</code> changes it, and <code>bash</code> runs commands. Nothing else in the loop moves.</p><h4><strong><span>How this fixes a bug</span></strong></h4><p>Say the user asks the agent to fix a failing test. The run looks roughly like this:</p><div class="callout-block" data-callout="true"><p><code>user: add() in src/adder.py has a bug, pytest is failing, can you fix it?</code></p><p><code>assistant &#8594; bash(&#8221;pytest tests/test_adder.py&#8221;)</code></p><p><code>tool_result:</code></p><p><code>FAILED tests/test_adder.py::test_negative</code></p><p><code>assert add(-1, 1) == 0 # got -2</code></p><p><code>assistant &#8594; read_file(&#8221;src/adder.py&#8221;)</code></p><p><code>tool_result:</code></p><p><code>def add(a, b):</code></p><p><code>return a - b # &#8592; obvious at a glance</code></p><p><code>assistant &#8594; write_file(</code></p><p><code>&#8220;src/adder.py&#8221;,</code></p><p><code>&#8220;def add(a, b):\n return a + b\n&#8221;</code></p><p><code>)</code></p><p><code>assistant &#8594; bash(&#8221;pytest tests/test_adder.py&#8221;)</code></p><p><code>tool_result:</code></p><p><code>2 passed in 0.03s</code></p><p><code>assistant: Fixed. add() was written as a - b; changed it back to a + b, all tests pass.</code></p></div><p>That&#8217;s a complete code-change loop: run the test, inspect the source, make the change, then verify it. Underneath, it is still ReAct&#8212;think &#8594; tool use &#8594; tool result &#8594; think. No new control flow required.</p><p><span>Drawn out:</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vP9h!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vP9h!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png 424w, https://substackcdn.com/image/fetch/$s_!vP9h!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png 848w, https://substackcdn.com/image/fetch/$s_!vP9h!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png 1272w, https://substackcdn.com/image/fetch/$s_!vP9h!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vP9h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png" width="1456" height="91" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/efa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:91,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:41506,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.agentengineering.co/i/215219617?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vP9h!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png 424w, https://substackcdn.com/image/fetch/$s_!vP9h!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png 848w, https://substackcdn.com/image/fetch/$s_!vP9h!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png 1272w, https://substackcdn.com/image/fetch/$s_!vP9h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefa179a2-5035-415b-9447-8ed234a04ea1_3204x200.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Three tools and one loop are enough for a straightforward bug fix. For something larger, such as a feature spanning several files, add the Plan-then-Act layer from Part 2. The job changes; the underlying structure does not.</p><p>That is the point: <em><strong>the harness skeleton is generic, while the tools determine the agent&#8217;s shape</strong></em><strong>.</strong> Replace <code>web_search</code> and <code>web_fetch</code> with <code>read_file</code>, <code>write_file</code>, and <code>bash</code>, and a research agent becomes a coding agent.</p><p><span>Speaking of context cost &#8212; some tools&#8217; return values are naturally huge. A bash build can dump thousands of lines of stdout. A </span><code>read_file</code><span> on a real file is hundreds of lines easy. A </span><code>web_fetch</code><span> against a real article is tens of thousands of tokens in one shot. Stuff all of that into context as-is and ten turns later the window pops.</span></p><div><hr></div><p style="text-align: center;"><em>For less hype and more engineering, pull up a chair.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.agentengineering.co/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h4><strong><span>Why does write_file support appending?</span></strong></h4><p><span>It is less about convenience than context. Without an append option, adding one line to a 1,000-line file could mean sending the entire file through the tool call again&#8212;and leaving all 1,000 lines sitting in the context window. Incremental writes avoid paying that cost for a tiny change.</span></p><p><span>The problem gets worse with tool results. A build can produce thousands of lines of output, while reading a real file can consume hundreds more. Keep feeding all of that back into the loop, and even a generous context window begins to look rather less generous.</span></p><p><span>The harness now needs somewhere else to put it. That brings us to context offloading.</span></p><h3><strong><span>Offloading: When context won&#8217;t fit</span></strong></h3><p><span>Every tool result goes back into the message history. A repository-wide search or noisy build can add thousands of tokens in a single call&#8212;useful in the moment, dead weight a few turns later.</span></p><p><span>Even a 200k-token context window begins to look rather less generous during a serious coding task. As it fills, the agent can lose track of earlier details, contradict itself, or simply run out of room.</span></p><p><span>The harness layer&#8217;s first answer is context offloading.</span></p><h4><strong><span>How offloading works</span></strong></h4><p>The rule is simple: when a tool result exceeds a set size, the harness writes the raw content to disk and leaves a path in the context. If the agent needs the details later, it retrieves them with <code>read_file</code>.</p><p>Suppose a <code>bash</code> call produces 1,000 tokens. The tool result the agent receives might look like this:</p><div class="callout-block" data-callout="true"><p><code>{</code></p><p><code>&#8220;role&#8221;: &#8220;tool&#8221;,</code></p><p><code>&#8220;tool_call_id&#8221;: &#8220;call_042&#8221;,</code></p><p><code>&#8220;content&#8221;: &#8220;[OFFLOADED] saved to abc123.log (1024 tokens). Use read_file to retrieve.&#8221;</code></p><p><code>}</code></p></div><p>The 1,000 tokens never enter the agent&#8217;s context. The harness intercepts the output and replaces it with a roughly 30-token pointer.</p><p>From the agent&#8217;s perspective, nothing else changes. It still calls <code>bash</code>, receives a tool result, and decides whether it needs to read the file. No new tool. No change to the ReAct loop. Just a smaller payload<strong>.</strong></p><h4><strong><span>Offloading is not compression</span></strong></h4><p><span>It is easy to confuse offloading with </span><strong><span>context compression</span></strong><span>. Both make room in the context window, but they do it differently:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bKBK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bKBK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png 424w, https://substackcdn.com/image/fetch/$s_!bKBK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png 848w, https://substackcdn.com/image/fetch/$s_!bKBK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png 1272w, https://substackcdn.com/image/fetch/$s_!bKBK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bKBK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png" width="578" height="1863" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1863,&quot;width&quot;:578,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:86047,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.agentengineering.co/i/215219617?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bKBK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png 424w, https://substackcdn.com/image/fetch/$s_!bKBK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png 848w, https://substackcdn.com/image/fetch/$s_!bKBK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png 1272w, https://substackcdn.com/image/fetch/$s_!bKBK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c170465-4b1e-4a2c-b8c7-295cde2bec7d_578x1863.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>They diverge on three axes:</span></p><ul><li><p><strong><span>Trigger timing:</span></strong><span> Offloading happens </span><strong><span>the moment the tool_result is born</span></strong><span> &#8212; the harness sees the return is over threshold and externalizes it on the spot. Compression happens </span><strong><span>the moment context is about to blow</span></strong><span> &#8212; only when occupancy is near the ceiling does the harness go back, scan history, and summarize.</span></p></li><li><p><strong><span>Lossy vs. lossless:</span></strong><span> Offloading is </span><strong><span>lossless and recoverable</span></strong><span> &#8212; not a byte of the original is lost; only the storage location changed, and the agent gets it back verbatim with read_file. Compression is </span><strong><span>lossy</span></strong><span> &#8212; summarize crushes 10 turns of conversation into a 200-word summary, and the dropped info is gone for good.</span></p></li><li><p><strong><span>Layer of abstraction:</span></strong><span> Offloading lives at the </span><strong><span>mechanism layer</span></strong><span>, handling hot/cold tiering of individual tool_results. Compression lives at the </span><strong><span>last-resort layer</span></strong><span>, packing whole stretches of history into archives. One is like the OS&#8217;s paging scheduler; the other is like tarring up old logs.</span></p></li></ul><p><span>The framing I like is this: </span><em><strong><span>offloading is hot/cold tiered storage; Compression is archival compression.</span></strong></em><span> The former goes first because it&#8217;s lossless and cheap. The latter is the fallback because it&#8217;s lossy but life-saving.</span></p><p>That is offloading&#8217;s place in the harness: when tool results are too large, move them instead of shrinking them.</p><div class="poll-embed" data-attrs="{&quot;id&quot;:1205618}" data-component-name="PollToDOM"></div><p>But fitting the agent&#8217;s working data solves only half the problem. Its capabilities may not fit either. That&#8217;s where Skill comes in.</p><h3><strong><span>Skill: Capability organization at the harness layer</span></strong></h3><p>So far, the harness has been helping one agent complete one task. Production introduces a different problem: the same agent may need dozens&#8212;or hundreds&#8212;of specialist capabilities. It might write Python one moment and diagnose a Kubernetes alert the next, each with a different set of instructions and pitfalls.</p><p>Loading all that guidance into the system prompt does not scale. If one Skill takes 3,000 tokens, fifty of them consume 150,000 tokens before the task has even begun. Even if they fit, the model must reason through a wall of mostly irrelevant instructions on every turn.</p><p>The harness&#8217;s answer to this capability explosion is Skill: make every capability available without loading all of them into context at once.</p><h4><strong><span>Load only the index at startup</span></strong></h4><p>Skill uses a simple form of progressive loading. Each Skill lives in a folder containing a <code>SKILL.md</code> file. At startup, the harness reads only the file&#8217;s frontmatter&#8212;a small metadata header containing its name and description&#8212;and adds that information to the system prompt&#8217;s <code>&lt;skills&gt;</code> index.</p><p>The full instructions remain on disk until the agent decides it needs them.</p><p>Here is the <code>&lt;skill-system&gt;</code> block from Claude Code:</p><div class="callout-block" data-callout="true"><p><code>&lt;skill-system&gt;</code></p><p><code>&lt;instruction&gt;</code></p><p><code>You have access to skills that provide optimized workflows for specific tasks. Each skill contains best practices, frameworks, and references to additional resources.</code></p><p><code>**Progressive Loading Pattern:**</code></p><p><code>1. When a user query matches a skill&#8217;s use case, immediately call `read_file` on the skill&#8217;s main file using the path attribute provided in the skill tag below</code></p><p><code>2. If an explicit requested skill is provided in the system context, load that skill first even if the user message is short</code></p><p><code>3. Read and understand the skill&#8217;s workflow and instructions</code></p><p><code>4. The skill file contains references to external resources under the same folder</code></p><p><code>5. Load referenced resources only when needed during execution</code></p><p><code>6. Follow the skill&#8217;s instructions precisely</code></p><p><code>&lt;/instruction&gt;</code></p><p><code>&lt;skills&gt;</code></p><p><code>&lt;skill name=&#8221;image-generation&#8221; path=&#8221;/Users/henry/.agents/skills/image-generation/SKILL.md&#8221;&gt;</code></p><p><code>Use this skill when the user requests to generate, create, imagine, or visualize images ...</code></p><p><code>&lt;/skill&gt;</code></p><p><code>&lt;skill name=&#8221;skill-creator&#8221; path=&#8221;/Users/henry/.agents/skills/skill-creator/SKILL.md&#8221;&gt;</code></p><p><code>Create new skills, modify and improve existing skills, and measure skill performance ...</code></p><p><code>&lt;/skill&gt;</code></p><p><code>&lt;/skills&gt;</code></p><p><code>&lt;/skill-system&gt;</code></p></div><h4><strong><span>What the prompt leaves out</span></strong></h4><p><span>Two absences matter:</span></p><ul><li><p><strong><span>No new tool.</span></strong><span> The harness does not need an invoke_skill() or load_skill() function. When the model selects a Skill, it uses the coding agent&#8217;s existing read_file tool to load the relevant SKILL.md. Skill simply reuses the filesystem.</span></p></li><li><p><strong><span>No Skill body.</span></strong><span> The index contains only a name, description, and path. The full instructions stay on disk until they are needed.</span></p></li></ul><p><span>That is progressive disclosure at the harness layer: keep the map in context, then load the instructions on demand.</span></p><h4><strong><span>Why not retrieve the right Skill?</span></strong></h4><p><span>Why not treat Skill selection as a search problem? Use BM25 or embeddings to match the user&#8217;s request against every Skill description, load the top results, and move on.</span></p><p><span>That can work&#8212;but if retrieval happens only once, the candidate set is fixed before the agent fully understands the task.</span></p><p><span>&#8220;Debug the alert that just fired&#8221; might begin with checking logs. Those logs may point towards a recent deployment, and resolving the incident may eventually require a company-specific postmortem. Each Skill becomes relevant at a different point in the reasoning chain.</span></p><p><span>Keeping the compact </span><code>&lt;skills&gt;</code><span> index in the system prompt lets the model reconsider its options on every ReAct turn. Retrieval could also be rerun each time, so the real choice is not search versus reasoning. It is one-shot selection versus selection that evolves with the task.</span></p><h4><strong><span>The full sequence</span></strong></h4><p>Startup, selection, and loading look like this:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Pvco!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Pvco!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png 424w, https://substackcdn.com/image/fetch/$s_!Pvco!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png 848w, https://substackcdn.com/image/fetch/$s_!Pvco!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png 1272w, https://substackcdn.com/image/fetch/$s_!Pvco!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Pvco!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png" width="1456" height="2137" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2137,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:285203,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.agentengineering.co/i/215219617?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Pvco!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png 424w, https://substackcdn.com/image/fetch/$s_!Pvco!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png 848w, https://substackcdn.com/image/fetch/$s_!Pvco!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png 1272w, https://substackcdn.com/image/fetch/$s_!Pvco!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd05123a4-46b9-4c1a-b86a-70813b2d693f_2110x3097.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The one line in this sequence diagram worth observing is the </span><code>read_file</code><span> call in the middle. It uses the same generic tool the Coding Agent already had. No new abstraction. Harness support for Skill costs zero net new tool surface.</span></p><h4><strong><span>Skill on the harness map</span></strong></h4><p>Skill does not change the ReAct loop or add another tool. It simply separates what the agent should always see from what it should load only when needed.</p><p>That small act of deferral is what lets an agent carry a hundred capabilities without dragging a hundred playbooks through every turn. With that, the harness map is complete.</p><h3><strong><span>Closing the loop on </span></strong><em><strong><span>Inside the Harness</span></strong></em></h3><p>Over the past three Fridays, we&#8217;ve taken the agent harness apart together, one mechanism at a time. If you&#8217;ve been with us since that first weather forecast, thank you for following the loop all the way through.</p><p>By now, a harness should look less like black-box middleware and more like a stack of small mechanisms:</p><ul><li><p><strong>ReAct makes the agent move.</strong> It supplies the basic rhythm: reason, act, observe, repeat.</p></li><li><p><strong>Plan-then-Act gives it direction.</strong> The agent works from an explicit plan instead of chasing whatever catches its attention next.</p></li><li><p><strong>Nudge keeps it on track.</strong> The harness brings the plan back into focus when it begins to fade.</p></li><li><p><strong>Offloading protects the context window.</strong> Oversized results move elsewhere without being permanently discarded.</p></li><li><p><strong>Skill makes capabilities manageable.</strong> The index stays visible; the full instructions appear only when needed.</p></li></ul><p>Each layer answers a failure exposed by the one beneath it. Each does one job, and you can tune each without rebuilding the entire system.</p><p>That is the larger point of harness engineering. The ReAct loop makes an agent run. The harness is what gives it a chance of surviving contact with reality.</p><div class="callout-block" data-callout="true"><p><strong>Keep up with Henry</strong></p><p>If three weeks inside the harness have left you wanting to poke at one yourself, Henry (the author responsible for your past three Fridays inside the harness) is building <a href="https://github.com/deer-flow/llm-space">LLM Space</a>, an open-source desktop app for prototyping agent ideas, inspecting harness runs, and replaying the bits that went wrong. Rather on theme, really.</p><p>&#8594; <a href="https://github.com/deer-flow/llm-space">Explore LLM Space on GitHub</a><br>&#8594; <a href="https://x.com/henry19840301">Follow Henry on X</a></p></div><p>That&#8217;s it for this one. We&#8217;ll pick up the conversation next week.</p><p>Until then, keep building.</p><blockquote><p>Tanya D&#8217;cruz<br><em>Editor-in-Chief</em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[Webinar: Your RAG pipeline doesn’t know when it’s wrong]]></title><description><![CDATA[Imran Ahmad on turning one-shot retrieval into a system that can evaluate its evidence, try again and admit when it doesn&#8217;t know]]></description><link>https://newsletter.agentengineering.co/p/free-webinar-your-rag-pipeline-doesnt</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/free-webinar-your-rag-pipeline-doesnt</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Tue, 08 Sep 2026 15:03:37 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/214542949/71ef2ad002e7068eb9a98afa3c22189d.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div class="callout-block" data-callout="true"><h4><code>Editor&#8217;s note</code></h4><p>The more conversations we have with people building these systems, the more useful material we find that cannot be contained within a single weekly issue. Rather than leave those sessions sitting in an archive or force them into our usual publishing cadence, we&#8217;ve decided to make a little more room in the week for them.</p><p>Between our regular issues, we&#8217;ll occasionally share technical sessions that deserve more than a forgotten recording link. A couple of weeks ago, we brought you <a href="https://open.substack.com/pub/agenticengg/p/free-webinar-building-ai-agent-harnesses?r=8770aj&amp;utm_campaign=post-expanded-share&amp;utm_medium=web">Nicole K&#246;nigstein&#8217;s webinar</a> on building AI agent harnesses for finance. This time, it&#8217;s <a href="https://www.linkedin.com/in/cloudanum/">Imran Ahmad</a>&#8217;s free webinar on the architecture of Agentic RAG.</p><p>Imran is a data scientist, AI architect, visiting professor, and the author of <em><a href="https://www.amazon.com/Agents-Every-Engineer-Must-Build/dp/1806109018/ref=sr_1_1?crid=357XEKWTU324E&amp;dib=eyJ2IjoiMSJ9.ElM92zcKd8pyXC5FH4xCW6-n6eV4tuLPkc6TfC4xzgVFRYF48gYjzNLRtQV8uBKLRcxWM_dmXd4X8crNRWB_QycD1I7HDEswhiAG6faLDVwTZKLIYzSIpvhnFvVPHYlnj55yUpmFhidqOsYmGxITTPGBO3x25GxFaFAkIQTGo9gPr04XVXnlEHGHcqwpy3I6m8_iXN32y0_28X5b7jocJAGzsqH5_uDIODXPc6OrmpQ.2jQzZVNakiuguitgnFF2c-_JyxnL-OstfY0ANHT2RmM&amp;dib_tag=se&amp;keywords=30+Agents+Every+AI+Engineer+Must+Build&amp;qid=1788788020&amp;sprefix=%2Caps%2C278&amp;sr=8-1">30 Agents Every AI Engineer Must Build</a></em>. His work looks beyond agent demos to the decisions that matter in production: when to add agency, where to place evaluation, and whether another loop is worth the extra cost and latency.</p><p>We&#8217;ve also put together a free take-home pack with Imran&#8217;s complete slide deck, a practical guide to 10 production agent patterns, and a one-page map of 30 agent architectures. You&#8217;ll find it later in the article.</p><p>Enough preamble. Let&#8217;s get into it!</p></div><p><span>Traditional RAG has an awkward habit: it keeps going.</span></p><p><span>Ask it about an internal company acronym, and it retrieves the closest available chunks. Misspell that acronym, and it still retrieves the closest available chunks. Those chunks are passed to the model, and the model produces an answer that may sound perfectly credible&#8212;even when the information needed to answer the question was never found.</span></p><p><span>Every component completed its assigned task. The system still failed.</span></p><p><span>That is the problem Imran explores in this free Agentic Engineering webinar on the architecture of agentic RAG.</span></p><p><span>His argument is not that conventional RAG is broken, nor that every retrieval pipeline now needs an agent. It is that traditional RAG is largely a one-way street. Once retrieval begins, there is often no meaningful point at which the system can inspect what it found, recognise that it is insufficient, and change course.</span></p><p><span>Agentic RAG introduces that missing loop.</span></p><h3><strong><span>RAG is really automated context generation</span></strong></h3><p><span>Large language models know a great deal, but they do not know everything about your organisation.</span></p><p><span>They may not know your internal terminology, current policies, proprietary research, or the particular meaning your company has assigned to an acronym. RAG addresses this by finding relevant information in an external knowledge source and adding it to the model&#8217;s prompt.</span></p><p><span>In other words, RAG automates context generation.</span></p><p><span>The familiar pipeline looks something like this: documents are divided into chunks, converted into embeddings, and stored in a vector database. When someone asks a question, the system searches for similar chunks, places them in the prompt, and asks the model to generate an answer from that context.</span></p><p><span>It is a useful architecture. There is a reason it became one of the first patterns companies reached for when building applications with large language models.</span></p><p><span>But its simplicity relies on several assumptions: that the user&#8217;s question is also a good search query, that the first retrieval will find the right evidence, that similarity means relevance, and that adding more context will improve the answer.</span></p><p><span>None of those assumptions is reliably true.</span></p><h3><strong><span>The problem with a one-way pipeline</span></strong></h3><p><span>Consider a user asking about an internal tool but misspelling its name.</span></p><p><span>The vector search does not necessarily respond with &#8220;I can&#8217;t find that.&#8221; It returns whichever chunks happen to be mathematically closest. A reranker may rearrange those chunks, but it is still ranking a weak set of results. The model then receives that material and does what it was designed to do: generate a plausible response.</span></p><p><span>The failure began before generation. The model was handed evidence that never supported the question.</span></p><p><span>Traditional RAG has no natural recovery mechanism here. It retrieves once and moves forward. It does not necessarily ask whether the query was misunderstood, whether a different source should be searched, or whether the available evidence is strong enough to justify an answer.</span></p><p><span>Adding an agent changes the shape of that pipeline. Retrieval can become a loop:</span></p><p style="text-align: center;"><strong>Plan</strong></p><p style="text-align: center;">&#8595;</p><p style="text-align: center;"><strong>Retrieve </strong></p><p style="text-align: center;">&#8595;</p><p style="text-align: center;"><strong>Evaluate </strong></p><p style="text-align: center;">&#8595;</p><p style="text-align: center;"><strong>Adapt</strong></p><p style="text-align: center;">&#8595;</p><p style="text-align: center;"><strong>Try again</strong></p><p><span>That sounds like a small architectural change. In practice, however, it changes where the system is allowed to exercise judgement.</span></p><div class="callout-block" data-callout="true"><p style="text-align: center;"><strong>A little something for being here</strong></p><p style="text-align: center;">If you&#8217;d rather build the patterns than just read about them, Imran is teaching a live, six-hour workshop on September 12.</p><p style="text-align: center;">In <em>10 Essential AI Agents Every Engineer Must Build</em>, you&#8217;ll build 10 working agents spanning RAG, tool use, fact-checking, multi-agent orchestration, multimodal AI, and autonomous planning. You&#8217;ll leave with the complete code repository, reusable architectures, and a complimentary copy of <em><a href="https://www.amazon.com/Agents-Every-Engineer-Must-Build/dp/1806109018/ref=sr_1_1?crid=357XEKWTU324E&amp;dib=eyJ2IjoiMSJ9.ElM92zcKd8pyXC5FH4xCW6-n6eV4tuLPkc6TfC4xzgVFRYF48gYjzNLRtQV8uBKLRcxWM_dmXd4X8crNRWB_QycD1I7HDEswhiAG6faLDVwTZKLIYzSIpvhnFvVPHYlnj55yUpmFhidqOsYmGxITTPGBO3x25GxFaFAkIQTGo9gPr04XVXnlEHGHcqwpy3I6m8_iXN32y0_28X5b7jocJAGzsqH5_uDIODXPc6OrmpQ.2jQzZVNakiuguitgnFF2c-_JyxnL-OstfY0ANHT2RmM&amp;dib_tag=se&amp;keywords=30+Agents+Every+AI+Engineer+Must+Build&amp;qid=1788788020&amp;sprefix=%2Caps%2C278&amp;sr=8-1">30 Agents Every AI Engineer Must Build</a></em>.</p><p style="text-align: center;">Tickets are normally $169.99. Agentic Engineering subscribers get 20% off with the private code <strong>AE20</strong>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://luma.com/10aiagents?coupon=AE20&amp;utm_source=agenticengnl&quot;,&quot;text&quot;:&quot;Get your ticket here&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://luma.com/10aiagents?coupon=AE20&amp;utm_source=agenticengnl"><span>Get your ticket here</span></a></p></div><h3><strong><span>Add agency where a specific failure needs it</span></strong></h3><p><span>&#8220;Agentic RAG&#8221; can make it sound as though the entire system must suddenly become autonomous. That is not the useful way to think about it.</span></p><p><span>Agency can be introduced at different points in the pipeline, depending on the failure you are trying to correct.</span></p><p><span>At the retrieval stage, an agent can interpret the user&#8217;s request, break a complicated question into smaller searches, rewrite a weak query, or choose between available retrieval tools.</span></p><p><span>After retrieval, it can examine the evidence and ask whether the results are truly relevant. If important information is missing, it can search again rather than blindly passing weak context downstream.</span></p><p><span>At the generation stage, another check can compare the final answer with the retrieved evidence, add citations, or reject claims that the sources do not support.</span></p><p><span>You may not need all three.</span></p><p><span>If retrieval is the main weakness, putting another agent around generation could add cost without fixing the actual problem. If the answer is poorly grounded despite strong retrieval, changing the search strategy may not help.</span></p><p><span>The important question is not, &#8220;How do we make this RAG pipeline agentic?&#8221;</span></p><p><span>It is, &#8220;Where does our existing pipeline fail, and would adding judgement at that point help it recover?&#8221;</span></p><p><span>Imran&#8217;s own course materials put the principle neatly: &#8220;Only add agency when you can name the failure it fixes.&#8221; The accompanying </span><a href="https://www.neurals.ca/agents/agentic-rag/"><span>Agentic RAG course hub</span></a><span> maps different failure modes to patterns including adaptive retrieval, corrective RAG, GraphRAG and groundedness checks.</span></p><h3><strong><span>A useful system needs the courage to say &#8220;I don&#8217;t know&#8221;</span></strong></h3><p><span>One of the most valuable behaviours in a knowledge system may also be one of the least impressive in a demo: refusing to answer.</span></p><p><span>If the retrieved passages do not contain the requested information, a reliable system should be able to say so. It might ask the user to clarify an acronym, try another source or explain that it could not find enough evidence.</span></p><p><span>That is much more useful than producing a polished answer from irrelevant material.</span></p><p><span>An evaluator inside the retrieval loop can provide that gate. It can score the evidence, check whether it covers the question, and decide whether the system should proceed, search again, or stop.</span></p><p><span>Of course, adding an evaluator does not make the system correct. The evaluator can make a poor judgement too. Its decisions need test cases, thresholds, and traces of their own.</span></p><p><span>The improvement is architectural rather than magical: the system now has a place where failure can be noticed.</span></p><h3><strong><span>More context is not always better context</span></strong></h3><p><span>Another assumption behind basic RAG is that retrieving more information gives the model a better chance of answering accurately.</span></p><p><span>Initially, it often does. But the relationship does not continue indefinitely.</span></p><p><span>As more chunks enter the prompt, relevant evidence competes with irrelevant material. Different sources may contradict each other. A precise but lengthy passage can pull the model away from the user&#8217;s original question. Imran describes this as a signal-and-noise problem: the goal is not to maximise context but to preserve the right signal.</span></p><p><span>Large context windows do not remove this problem. Being able to place the equivalent of several books into a prompt does not mean the model will identify the right paragraph, weigh every source correctly, or remain focused on the original task.</span></p><p><span>This is where an agent can help curate context rather than merely accumulate it. It can search for missing evidence, reject weak results, and recognise contradictions before generation.</span></p><p><span>But again, that judgement has a price.</span></p><h3><strong><span>Retrieval should reflect the question being asked</span></strong></h3><p><span>A short factual question and a request to explain a complete block of code do not require the same kind of context.</span></p><p><span>With a fixed RAG pipeline, they may still be treated in much the same way. The same chunk size, overlap, retrieval algorithm, and number of results are used regardless of the question.</span></p><p><span>An agentic retrieval layer can make the process more adaptive. Depending on the tools and indexes available, it can choose a retrieval strategy, change the query, search multiple sources, or gather a wider section of material when the answer depends on information spread across several passages.</span></p><p><span>For a compound question, it might separate the request into several searches and confirm that each part has been answered. For a vague query, it might rewrite the search or ask the user for clarification. For a technical question, it may retrieve a larger, more coherent unit rather than an isolated fragment.</span></p><p><span>The underlying trade-offs do not disappear. Larger chunks carry more context but reduce precision. Smaller chunks may retrieve precisely while cutting away something essential. Semantic chunking can preserve meaning, but it introduces more computation and complexity.</span></p><p><span>Agency allows the system to make choices within those trade-offs. It does not abolish them.</span></p><h3><strong><span>If the system can change course, you need to see where it went</span></strong></h3><p><span>A fixed pipeline is relatively easy to follow. The query went in, a set of chunks came back, and the model generated an answer.</span></p><p><span>An agentic system may rewrite the query, call more than one retriever, reject the first result, alter its plan, and try again. That adaptability is the point&#8212;but it also makes the system harder to inspect.</span></p><p><span>Imran recommends a provenance layer that records the path the system took. What did it search for? Which evidence did it retrieve? Why was one result accepted and another rejected? Which tool did it choose? Where did the final claim come from?</span></p><p><span>Without those traces, it becomes difficult to distinguish a retrieval failure from an evaluation failure or a generation failure. The system may be more capable, but the team responsible for it will know less about why it behaved as it did.</span></p><p><span>Observability is therefore not an optional dashboard added after the architecture is complete. It is part of what makes an adaptive retrieval loop governable.</span></p><h3><strong><span>Agency sends a bill</span></strong></h3><p><span>Each planning step, evaluation, and retry may require another model call. That adds latency, token usage, and another possible point of failure.</span></p><p><span>A system that runs three searches and evaluates each one may retrieve better evidence than a one-shot pipeline. It may also take longer and cost considerably more.</span></p><p><span>Imran&#8217;s advice is refreshingly blunt: &#8220;Don&#8217;t hire a PhD for the job of a receptionist.&#8221;</span></p><p><span>A small, focused model may be enough to classify a query, choose a retriever, or judge whether two passages are relevant. Some of these decisions can be handled by local or quantised models. Others may not require an LLM at all.</span></p><p><span>The most sophisticated model should not automatically be placed inside every part of the loop. The model&#8212;and the loop itself&#8212;should be proportionate to the decision being made.</span></p><p><span>This is where agentic RAG becomes an engineering problem rather than a naming exercise. The team has to decide where another round of reasoning is worth its cost.</span></p><div class="callout-block" data-callout="true"><p style="text-align: center;"><strong><span data-color="#fd2e1b" style="color: rgb(253, 46, 27);">Still thinking about it?</span></strong></p><p style="text-align: center;"><span>You could be building these systems with Imran on September 12. </span></p><p style="text-align: center;"><span>Agentic Engineering subscribers can take 20% off the workshop with code </span><strong><span>AE20</span></strong><span> if ten working agents sound more useful than ten more bookmarked tutorials.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://luma.com/10aiagents?coupon=AE20&amp;utm_source=agenticengnl&quot;,&quot;text&quot;:&quot;SAVE YOUR SEAT&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://luma.com/10aiagents?coupon=AE20&amp;utm_source=agenticengnl"><span>SAVE YOUR SEAT</span></a></p></div><h3><span>The goal is not maximum agency</span></h3><p><span>The choice is not between an outdated RAG pipeline and a fully autonomous knowledge agent.</span></p><p><span>There is a spectrum between them.</span></p><p><span>A conventional pipeline may be entirely sufficient for a narrow, predictable knowledge base. A retrieval evaluator may solve the most serious failure without changing anything else. A more complex investigation might justify query decomposition, multiple retrieval tools, and several rounds of evidence gathering.</span></p><p><span>The right architecture is the smallest one that can recognise and recover from the failures that matter.</span></p><p><span>That is the real case for agentic RAG. It does not simply retrieve more information or attach an agent to a familiar acronym. It gives the system a chance to notice that the evidence is weak before fluent language turns that weakness into an answer.</span></p><h3>Your webinar companion pack</h3><ul><li><p><span>Explore the </span><a href="https://github.com/cloudanum/webinar-agentic-rag-labs"><span>hands-on labs</span></a> and supporting material on GitHub.</p></li><li><p>Download <em><strong><a href="https://agentic-rag-toolkit.packt-4364.chatgpt.site/">The Agent Engineer&#8217;s Field Guide</a></strong></em> for 10 production agent patterns, including when to use them, what tends to break first, and how to evaluate them.</p></li><li><p>Download <em><strong><a href="https://agentic-rag-toolkit.packt-4364.chatgpt.site/">30 Agents at a Glance</a></strong></em> for a one-page map of all 30 architectures, organised from foundational patterns to specialised and domain-specific systems.</p></li><li><p>Download Imran&#8217;s <a href="https://agentic-rag-toolkit.packt-4364.chatgpt.site/">slide deck</a> to revisit the architectures, retrieval loops, and implementation patterns later.</p></li><li><p><span>Continue through the patterns in Imran&#8217;s book, </span><em><strong><a href="https://www.amazon.com/Agents-Every-Engineer-Must-Build/dp/1806109018/ref=sr_1_1?crid=357XEKWTU324E&amp;dib=eyJ2IjoiMSJ9.ElM92zcKd8pyXC5FH4xCW6-n6eV4tuLPkc6TfC4xzgVFRYF48gYjzNLRtQV8uBKLRcxWM_dmXd4X8crNRWB_QycD1I7HDEswhiAG6faLDVwTZKLIYzSIpvhnFvVPHYlnj55yUpmFhidqOsYmGxITTPGBO3x25GxFaFAkIQTGo9gPr04XVXnlEHGHcqwpy3I6m8_iXN32y0_28X5b7jocJAGzsqH5_uDIODXPc6OrmpQ.2jQzZVNakiuguitgnFF2c-_JyxnL-OstfY0ANHT2RmM&amp;dib_tag=se&amp;keywords=30+Agents+Every+AI+Engineer+Must+Build&amp;qid=1788788020&amp;sprefix=%2Caps%2C278&amp;sr=8-1"><span>30 Agents Every AI Engineer Must Build</span></a></strong></em><strong>.</strong></p></li></ul><div><hr></div><p style="text-align: center;"><em>For less hype and more engineering, pull up a chair.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.agentengineering.co/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p>That&#8217;s it for this one. We&#8217;ll pick up the conversation next week.</p><p>Until then, keep building.</p><blockquote><p>Tanya D&#8217;cruz<br><em>Editor-in-Chief</em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[Inside the harness (Part 2): Give it a plan]]></title><description><![CDATA[How planning and nudges keep a ReAct loop from losing the plot]]></description><link>https://newsletter.agentengineering.co/p/inside-the-harness-part-2-give-it</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/inside-the-harness-part-2-give-it</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Fri, 04 Sep 2026 15:33:29 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/9fea4ab5-0384-496c-857c-e442eaa15441_1478x1064.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="callout-block" data-callout="true"><h4><code>Editor&#8217;s note</code> </h4><p>Last week in <em>Part 1</em>, we stripped the agent harness back to its ReAct loop and left it walking in rather expensive circles.</p><p>This week, we give it a plan. <em>Part 2</em> looks at how <code>write_todos</code>, a quick briefing and well-timed nudges add strategy without changing the loop underneath.</p><p>Turns out, even sophisticated agents benefit from writing things down <em>and</em> being reminded to look at what they wrote.</p></div><p><span>Throw a complex question at pure multi-turn ReAct, and a structural flaw shows up &#8212; the model walks step by step with no global view. Hit it with a research question that needs five or six angles assembled together, and it&#8217;ll grab the most obvious two or three in the first few steps and call it done. It doesn&#8217;t even know which angles it skipped.</span></p><p><span>The fix is crude but effective: </span><em><span>make it write the plan down first, then execute</span></em><span>.</span></p><h3><strong><span>A new tool: write_todos</span></strong></h3><p><span>In the Deep Research harness, we add one new tool: </span><code>write_todos</code><span>. It does nothing! It literally just writes a TODO list into the agent&#8217;s state. The flip switch is one line in the system prompt:</span></p><div class="callout-block" data-callout="true"><p><code>&lt;agent name=&#8221;Tara&#8221;&gt;</code></p></div><p><span>You&#8217;re a helpful agent.</span></p><div class="callout-block" data-callout="true"><p><code>&lt;/agent&gt;</code></p><p><code>&lt;task&gt;</code></p><p><code>Your task is to use the given tools to solve the user&#8217;s problem.</code></p><p><code>Use `write_todos` to make a plan and break down tasks.</code></p><p><code>&lt;/task&gt;</code></p><p><code>&lt;notes&gt;</code></p><p><code>- Parallel tool calling is supported.</code></p><p><code>- If you are unfamiliar with a topic, you can use a one-time `web_search` to get a general understanding of it before planning.</code></p><p><code>- Don&#8217;t forget to update todos using `write_todos` after each task is done.</code></p><p><code>&lt;/notes&gt;</code></p></div><p><span>Look at the line inside </span><code>&lt;task&gt;: Use \write_todos `to make a plan and break down tasks.`</code><span> That single sentence flips the model from &#8220;see one step, take one step&#8221; to &#8220;plan first, then act.&#8221;</span></p><p><span>The tool&#8217;s schema is dirt simple:</span></p><div class="callout-block" data-callout="true"><p><code>{</code></p><p><code>&#8220;name&#8221;: &#8220;write_todos&#8221;,</code></p><p><code>&#8220;description&#8221;: &#8220;Create or update the current plan.&#8221;,</code></p><p><code>&#8220;input_schema&#8221;: {</code></p><p><code>&#8220;type&#8221;: &#8220;object&#8221;,</code></p><p><code>&#8220;properties&#8221;: {</code></p><p><code>&#8220;todo_list&#8221;: {</code></p><p><code>&#8220;type&#8221;: &#8220;array&#8221;,</code></p><p><code>&#8220;items&#8221;: {</code></p><p><code>&#8220;type&#8221;: &#8220;object&#8221;,</code></p><p><code>&#8220;properties&#8221;: {</code></p><p><code>&#8220;content&#8221;: { &#8220;type&#8221;: &#8220;string&#8221; },</code></p><p><code>&#8220;status&#8221;: {</code></p><p><code>&#8220;type&#8221;: &#8220;string&#8221;,</code></p><p><code>&#8220;enum&#8221;: [&#8221;pending&#8221;, &#8220;in_progress&#8221;, &#8220;completed&#8221;]</code></p><p><code>}</code></p><p><code>},</code></p><p><code>&#8220;required&#8221;: [&#8221;content&#8221;, &#8220;status&#8221;]</code></p><p><code>}</code></p><p><code>}</code></p><p><code>},</code></p><p><code>&#8220;required&#8221;: [&#8221;todo_list&#8221;]</code></p><p><code>}</code></p><p><code>}</code></p></div><h3><strong><span>The new behavior</span></strong></h3><p><span>Add this one tool and the loop visibly changes shape. Take a sample question: &#8220;research the current state of Chinese EV exports to Europe in 2026.&#8221; The transcript now looks like:</span></p><div class="callout-block" data-callout="true"><p><code>user: Research the state of Chinese EV exports to Europe in 2026</code></p><p><code>assistant &#8594; tool_use: write_todos([</code></p><p><code>{ content: &#8220;Understand the overall scale of Chinese EV exports to Europe in 2026&#8221;, status: &#8220;in_progress&#8221; },</code></p><p><code>{ content: &#8220;Map the main players and their European market share&#8221;, status: &#8220;pending&#8221; },</code></p><p><code>{ content: &#8220;Track EU tariff policy and anti-subsidy investigations&#8221;, status: &#8220;pending&#8221; },</code></p><p><code>{ content: &#8220;Summarize local manufacturing and channel strategy&#8221;, status: &#8220;pending&#8221; }</code></p><p><code>])</code></p><p><code>assistant &#8594; tool_use: web_search(&#8221;2026 Chinese EV Europe export data&#8221;)</code></p><p><code>... (several web_search / web_fetch turns) ...</code></p><p><code>assistant &#8594; tool_use: write_todos([</code></p><p><code>{ content: &#8220;Understand the overall scale of Chinese EV exports to Europe in 2026&#8221;, status: &#8220;completed&#8221; },</code></p><p><code>{ content: &#8220;Map the main players and their European market share&#8221;, status: &#8220;in_progress&#8221; },</code></p><p><code>...</code></p><p><code>])</code></p><p><code>... (work through the remaining items) ...</code></p><p><code>assistant: &lt;final answer&gt;</code></p></div><p><span>The biggest difference vs. pure multi-turn ReAct is that </span><code>write_todos</code><span> call right at the top &#8212; it pins down the exam outline upfront. If execution skips a dimension, there&#8217;s an unchecked pending item sitting on the plan. The model comes back and picks it up.</span></p><div><hr></div><p style="text-align: center;"><em>Building agents that need to work beyond the demo? Join us for the engineering that holds up in production.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.agentengineering.co/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p><span>Side-by-side helps. Pure multi-turn ReAct first:</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!91Uj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!91Uj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png 424w, https://substackcdn.com/image/fetch/$s_!91Uj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png 848w, https://substackcdn.com/image/fetch/$s_!91Uj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png 1272w, https://substackcdn.com/image/fetch/$s_!91Uj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!91Uj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png" width="1456" height="164" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:164,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:22744,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.agentengineering.co/i/214003496?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!91Uj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png 424w, https://substackcdn.com/image/fetch/$s_!91Uj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png 848w, https://substackcdn.com/image/fetch/$s_!91Uj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png 1272w, https://substackcdn.com/image/fetch/$s_!91Uj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bc54a89-0de1-4098-8ce6-4c6b0479bd91_1780x200.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><span>Each step only looks at the current </span><code>tool_result</code><span>. No looking back, no looking ahead. Now Plan-then-Act:</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!N2iI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!N2iI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png 424w, https://substackcdn.com/image/fetch/$s_!N2iI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png 848w, https://substackcdn.com/image/fetch/$s_!N2iI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png 1272w, https://substackcdn.com/image/fetch/$s_!N2iI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!N2iI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png" width="1456" height="88" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:88,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:41428,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.agentengineering.co/i/214003496?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!N2iI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png 424w, https://substackcdn.com/image/fetch/$s_!N2iI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png 848w, https://substackcdn.com/image/fetch/$s_!N2iI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png 1272w, https://substackcdn.com/image/fetch/$s_!N2iI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b8deff7-c8ef-4d5f-a165-b35d73807dd4_3319x200.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><span>What&#8217;s new: a plan at the start, and a state-update after each step. The loop itself didn&#8217;t change. We just added one tool and one line of system prompt.</span></p><p><span>That&#8217;s the Plan-then-Act skeleton. In practice, you&#8217;ll run into two small pitfalls:</span></p><ul><li><p><span>First, on unfamiliar topics, the plan goes off the rails. The model has only a vague grasp of the field. The four sub-tasks it lists from that half-baked understanding might miss the most important dimension from the start.</span></p></li><li><p><span>Second, mid-run, the model forgets to update todos. It finishes task two but doesn&#8217;t call </span><code>write_todos</code><span> to mark it completed, and the new sub-task it discovered along the way never makes it onto the list. The plan quietly decays.</span></p></li></ul><p><span>Two small patches, one for each.</span></p><h3><strong><span>Patch A: Brief yourself before planning</span></strong></h3><p><span>The first pitfall is easy to grok &#8212; </span><em><span>the model can&#8217;t plan something it doesn&#8217;t recognize</span></em><span>.</span></p><p><span>Say the user throws a niche term at it, like &#8220;research how pp-ocr performs on long-tail Chinese classical text.&#8221; The model has only a fuzzy notion of pp-ocr (&#8221;Baidu&#8217;s open-source OCR engine, I think?&#8221;). If it goes straight to </span><code>write_todos</code><span>, it&#8217;ll probably list four lukewarm items: figure out what it is, check performance, list pros and cons, summarize. The dimensions that actually matter (its multilingual branches, version history, benchmark against a specific competitor) never get touched.</span></p><h4><strong><span>One-time briefing</span></strong></h4><p><span>The patch is light. In the system prompt, explicitly allow the agent one web_search </span><em><strong><span>before</span></strong></em><span> planning, purely to build basic familiarity with the topic:</span></p><div class="callout-block" data-callout="true"><p><code>&lt;notes&gt;</code></p><p><code>- Parallel tool calling is supported.</code></p><p><code>- If you are unfamiliar with a topic, you can use a one-time `web_search`</code></p><p><code>to get a general understanding of it before planning.</code></p><p><code>- Don&#8217;t forget to update todos using `write_todos` after each task is done.</code></p><p><code>&lt;/notes&gt;</code></p></div><p><span>Two key phrases here: one-time and before planning. This isn&#8217;t a blanket &#8220;search whenever you want&#8221; &#8212; it&#8217;s a hard-coded </span><em><strong><span>protocol</span></strong></em><span> in the system prompt: before you plan, you can do exactly one web_search to brief yourself.</span></p><p><span>With that briefing, the model searches &#8220;pp-ocr&#8221; once, skims the first few snippets, learns it&#8217;s a PaddlePaddle-family OCR model with v1 through v5, multilingual branches for Chinese / English / Japanese / Korean, server-side and lightweight variants&#8230; </span><em><strong><span>then</span></strong></em><span> writes the todos. The plan turns from &#8220;look up what it is&#8221; (</span><em><span>useless</span></em><span>) into &#8220;compare v4 vs v5 recognition accuracy on handwritten classical text&#8221; (</span><em><span>actionable</span></em><span>).</span></p><p><span>That patches the input side of planning. The next pitfall is on the output side &#8212; the model forgets to update its own todos.</span></p><h3><strong><span>Patch B: Nudge after every step</span></strong></h3><p><span>The second pitfall &#8212; </span><em><strong><span>the model forgets to maintain its todos</span></strong></em><strong><span>.</span></strong></p><p><span>You&#8217;ve seen this. The opening plan lists four neat items. Item one gets done; the model jumps to item two without updating anything. Halfway through item three, it realizes it needs a new sub-task, but instead of appending it to the list, it just searches. By the end of the run, the plan state looks the same as it did after step one. Updated once, never again.</span></p><p><span>Putting &#8220;remember to update todos&#8221; in the system prompt isn&#8217;t enough to fix this. The prompt is a one-shot backdrop. After a few turns, the model&#8217;s attention weights have shifted to the more recent </span><code>tool_results</code><span>. The backdrop&#8217;s reminders fade.</span></p><h4><strong><span>Mechanism: nudge</span></strong></h4><p><span>The harness&#8217;s solution is called a </span><em><strong><span>nudge</span></strong><span> &#8212; </span><strong><span>after every tool_result</span></strong></em><span>, when the todo list is non-empty, the harness hard-codes an extra reminder onto the context for the model:</span></p><div class="callout-block" data-callout="true"><p><code>Don&#8217;t forget to update todos using `write_todos` after each task is done.</code></p></div><p><span>This isn&#8217;t a user message, and it isn&#8217;t an assistant message the model generated. It&#8217;s a system-reminder spliced in by the harness right after the </span><code>tool_result</code><span>. From the model&#8217;s perspective: every time it finishes a step and is about to plan its next move, this line is in its face.</span></p><p><span>In sequence diagram form:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SpWB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SpWB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png 424w, https://substackcdn.com/image/fetch/$s_!SpWB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png 848w, https://substackcdn.com/image/fetch/$s_!SpWB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png 1272w, https://substackcdn.com/image/fetch/$s_!SpWB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SpWB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png" width="1456" height="1427" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1427,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:119994,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.agentengineering.co/i/214003496?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SpWB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png 424w, https://substackcdn.com/image/fetch/$s_!SpWB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png 848w, https://substackcdn.com/image/fetch/$s_!SpWB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png 1272w, https://substackcdn.com/image/fetch/$s_!SpWB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F465b2827-7429-463c-989f-efbecdde23b7_1484x1454.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Two harness actions worth noticing: one, pass the tool_result back to the LLM as-is; two, when &#8220;current todo list is non-empty,&#8221; </span><em><strong><span>also</span></strong></em><span> append a system-reminder. Both arrive in the same message. From the model&#8217;s view, it looks as if the tool return came bundled with a built-in nudge.</span></p><p><span>And the nudge </span><em><strong><span>fires on every tool_result.</span></strong></em><span> Not just at the start, not every few rounds. As long as the plan isn&#8217;t done, it keeps firing. The model&#8217;s attention gets pulled back to &#8220;what&#8217;s my plan progress?&#8221; again and again, and the probability of forgetting to update drops sharply.</span></p><p><span>Patch A covers the input side of planning. Patch B covers the output side. Together with the </span><code>write_todos</code><span> tool itself, the harness for Deep Research is done &#8212; from single-turn ReAct, to multi-turn loop, to Plan-then-Act with two patches. That&#8217;s all of it.</span></p><p><span>But what does the same skeleton look like with a different toolset? </span>That&#8217;s the crucial question we&#8217;ll answer in Part 3, where we cut to coding&#8230; then follow the loop until its tool results and capabilities no longer fit.</p>]]></content:encoded></item><item><title><![CDATA[Inside the harness (Part 1): Start with the loop]]></title><description><![CDATA[How a weather lookup becomes deep research without changing the ReAct machinery underneath]]></description><link>https://newsletter.agentengineering.co/p/inside-the-harness-part-1-start-with</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/inside-the-harness-part-1-start-with</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Fri, 28 Aug 2026 15:34:31 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/5786e877-6860-4e14-b8b3-d9b5c46684a0_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="callout-block" data-callout="true"><h4><code>Editor&#8217;s note</code></h4><p>Agent harnesses have acquired a remarkable amount of jargon for something often built around a simple ReAct loop: think, call a tool, inspect the result, repeat. </p><p>Over the next three issues, we&#8217;re taking that loop apart and watching the harness grow around it as the problems get harder. <em>Part 1</em> strips the ReAct loop back to its basics. <em>Part 2</em> adds planning and nudges to help keep it on course. <em>Part 3</em> turns the same loop into a coding agent, then tackles what happens when its context and capabilities no longer fit.</p><p>We&#8217;re publishing all three parts on Fridays to see whether a technical deep dive fares better once the midweek inbox stampede has passed. After <em>Part 3</em>, we&#8217;ll return to our usual schedule.</p></div><p><span>Crack open any agent thread these days, and the buzzwords fly: Multi-Agent, Deep Research, Skill, Sub-Agent, Orchestrator&#8230; each one more abstract and more intimidating than the last. But peel the fancy wrapping off, layer by layer, and what&#8217;s actually running underneath is the same loop: the model thinks, calls a tool, looks at the result, thinks again</span><strong><span>.</span></strong><span> A bare-bones ReAct loop.</span></p><p><span>The staircase looks roughly like this:</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZXQs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZXQs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png 424w, https://substackcdn.com/image/fetch/$s_!ZXQs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png 848w, https://substackcdn.com/image/fetch/$s_!ZXQs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png 1272w, https://substackcdn.com/image/fetch/$s_!ZXQs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZXQs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png" width="1456" height="165" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:165,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:69403,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.agentengineering.co/i/213143990?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ZXQs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png 424w, https://substackcdn.com/image/fetch/$s_!ZXQs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png 848w, https://substackcdn.com/image/fetch/$s_!ZXQs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png 1272w, https://substackcdn.com/image/fetch/$s_!ZXQs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c535808-dc20-4a39-952a-c3a195f62ad5_2482x281.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><span>Each step adds almost nothing to the last. Multi-turn just lets the loop spin more.</span></p><p>Plan gives the loop a persistent TODO list<span>. Coding Agent swaps </span><code>web_search</code><span> for read_file and bash. Offloading and Skill are two different compromises for when stuff stops fitting in the context window. The point of internalising a harness isn&#8217;t memorising each level&#8217;s name but rather seeing clearly what problem each level solves that the level below it didn&#8217;t.</span></p><p><span>That&#8217;s the path this series takes. We start with the weather forecast.</span></p><h3><strong><span>Single-turn ReAct loop: Weather forecast</span></strong></h3><p><span>Picture the smallest possible scenario: the user asks, &#8220;is it cold in Beijing today?&#8221; We give the model a single tool: </span><code>get_weather(city).</code></p><p><span>The model takes one look at the question and makes one decision: &#8220;I can&#8217;t answer this on my own, I need to look it up.&#8221; So it emits a </span><code>tool_use</code><span>: </span><code>get_weather(city=&#8221;Beijing&#8221;)</code><span>. The harness picks up that </span><code>tool_use</code><span>, actually runs the function, takes the result (say </span><code>{&#8221;temp&#8221;: 3, &#8220;condition&#8221;: &#8220;sunny&#8221;}</code><span>), and shoves it back into the message queue as a </span><code>tool_result</code><span>. The model takes another look and decides whether to call another tool or answer the user. If it answers, what it emits is plain text (no </span><code>tool_use</code><span>) and the loop ends.</span></p><p><span>That&#8217;s one full ReAct round-trip: </span><strong><span>Reason (decide what to call) &#8594; Act (emit tool_call) &#8594; Observe (look at tool_result)</span></strong><span>, then either go again or hand back the final answer.</span></p><h4><strong><span>What the harness actually is</span></strong></h4><p>Strip it down, and the harness is almost embarrassingly simple<span>:</span></p><div class="callout-block" data-callout="true"><p><code>messages = [{&#8221;role&#8221;: &#8220;user&#8221;, &#8220;content&#8221;: user_input}]</code></p><p><code>while True:</code></p><p><code>response = llm.call(messages, tools=TOOLS)</code></p><p><code>messages.append(response)</code></p><p><code>if not response.tool_calls:</code></p><p><code># Model didn&#8217;t call any tools &#8212; loop terminates</code></p><p><code>break</code></p><p><code># Run all tool_calls in parallel, pack results back in</code></p><p><code>results = run_tools_parallel(response.tool_calls)</code></p><p><code>messages.extend(results)</code></p><p><code>return response.content</code></p></div><p><span>The harness doesn&#8217;t ask &#8220;have we searched enough?&#8221; or &#8220;should we try a different keyword?&#8221; The only thing it checks is: </span><em><span>does this assistant turn contain a tool_call? If yes, execute. If no, exit.</span></em><span> All the thinking is the model&#8217;s; the harness is just a courier.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XzRJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XzRJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png 424w, https://substackcdn.com/image/fetch/$s_!XzRJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png 848w, https://substackcdn.com/image/fetch/$s_!XzRJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png 1272w, https://substackcdn.com/image/fetch/$s_!XzRJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XzRJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png" width="1456" height="1069" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1069,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:161615,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.agentengineering.co/i/213143990?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XzRJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png 424w, https://substackcdn.com/image/fetch/$s_!XzRJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png 848w, https://substackcdn.com/image/fetch/$s_!XzRJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png 1272w, https://substackcdn.com/image/fetch/$s_!XzRJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa449d49-58e3-4d9c-82af-d04eec8a5561_2084x1530.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><strong><span>Emitting multiple tool_calls in one turn</span></strong></h4><p><span>The interesting bit comes with the next question: &#8220;which is colder, Beijing or Shanghai?&#8221;</span></p><p><span>A competent model isn&#8217;t going to look up Beijing first and then Shanghai. It emits two </span><code>tool_use</code><span> calls in the same assistant turn:</span></p><div class="callout-block" data-callout="true"><p><code>{</code></p><p><code>&#8220;role&#8221;: &#8220;assistant&#8221;,</code></p><p><code>&#8220;content&#8221;: [</code></p><p><code>{&#8221;type&#8221;: &#8220;tool_use&#8221;, &#8220;id&#8221;: &#8220;t1&#8221;, &#8220;name&#8221;: &#8220;get_weather&#8221;, &#8220;input&#8221;: {&#8221;city&#8221;: &#8220;Beijing&#8221;}},</code></p><p><code>{&#8221;type&#8221;: &#8220;tool_use&#8221;, &#8220;id&#8221;: &#8220;t2&#8221;, &#8220;name&#8221;: &#8220;get_weather&#8221;, &#8220;input&#8221;: {&#8221;city&#8221;: &#8220;Shanghai&#8221;}}</code></p><p><code>]</code></p><p><code>}</code></p></div><p><span>On the harness side, as long as you wrote run_tools to execute in parallel (the </span><code>run_tools_parallel</code><span> in the pseudocode above), both queries fire at once. Two </span><code>tool_results</code><span> come back together. Next turn, the model sees them side by side and answers the comparison directly.</span></p><p><span>This is a property of ReAct that people consistently underrate: a single assistant turn can fan out to any number of </span><code>tool_calls</code><span>. When you write your harness, don&#8217;t default to &#8220;one at a time&#8221; &#8212; the moment you hit &#8220;compare three cities&#8221; or &#8220;read ten files in parallel,&#8221; it&#8217;ll be painfully slow.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YwAY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YwAY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png 424w, https://substackcdn.com/image/fetch/$s_!YwAY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png 848w, https://substackcdn.com/image/fetch/$s_!YwAY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png 1272w, https://substackcdn.com/image/fetch/$s_!YwAY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YwAY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png" width="1456" height="1829" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1829,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:145380,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.agentengineering.co/i/213143990?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!YwAY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png 424w, https://substackcdn.com/image/fetch/$s_!YwAY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png 848w, https://substackcdn.com/image/fetch/$s_!YwAY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png 1272w, https://substackcdn.com/image/fetch/$s_!YwAY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45149146-170c-40a8-a7c9-f41ab73a80f8_1656x2080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><strong><span>Is single-turn enough?</span></strong></h4><p>For a straightforward weather check, yes. But ask for Beijing&#8217;s chance of rain this Friday compared with the same week over the past five years, and one lookup no longer does the job.</p><p>The model has to produce the forecast, retrieve the historical data, compare the two, and perhaps go back to fill in anything it missed. It needs a <em><strong><mark data-color="#e6b8af" style="background-color: rgb(230, 184, 175); color: rgb(0, 0, 0);">multi-turn loop</mark></strong></em>, with each step shaped by what the previous one returned.</p><h3><strong><span>Multi-turn loop: Deep research</span></strong></h3><p><span>Swap </span><code>get_weather</code><span> for two more general tools (</span><code>web_search</code> and <code>web_fetch</code><span>) and you&#8217;ve got the skeleton of Deep Research.</span></p><p><span>Two tools, but the behaviour space they open up jumps by orders of magnitude. The model can search, pull back a pile of candidate links, pick the most relevant one or two and fetch the full text, realise it&#8217;s missing something and search again with a different keyword, fetch the next link. The whole trajectory is </span><em><strong><span>meandering, stop-and-go</span></strong></em><span>. How many steps, when to pivot &#8212; entirely the model&#8217;s call.</span></p><p>In a multi-turn run, every tool call depends on what the previous one returned. Search and retrieval interleave, and the trajectory is not known in advance: the agent might converge in three turns or still be missing pieces at fifteen.</p><h4><strong><span>The shape of a multi-turn loop</span></strong></h4><p><span>Drawn as a sequence diagram, the shape is clean:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eyia!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eyia!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png 424w, https://substackcdn.com/image/fetch/$s_!eyia!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png 848w, https://substackcdn.com/image/fetch/$s_!eyia!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png 1272w, https://substackcdn.com/image/fetch/$s_!eyia!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eyia!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png" width="1456" height="1286" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1286,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:152913,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.agentengineering.co/i/213143990?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eyia!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png 424w, https://substackcdn.com/image/fetch/$s_!eyia!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png 848w, https://substackcdn.com/image/fetch/$s_!eyia!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png 1272w, https://substackcdn.com/image/fetch/$s_!eyia!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0fc9b7d-938c-49cc-9aa8-b0ec4df829f5_2047x1808.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The harness code is unchanged. The loop simply spins more times<span>.</span></p><h4><strong><span>Who decides when the loop stops?</span></strong></h4><p><span>This is the most counterintuitive part of multi-turn ReAct: the LLM decides when to halt, not the harness.</span></p><p>Whether the loop continues<span> or not is the </span><em><strong><span>model&#8217;s call</span></strong></em><strong><span>.</span></strong><span> If it thinks it doesn&#8217;t have enough, it emits another </span><code>web_search</code><span>. If it thinks it does, it skips </span><code>tool_use</code><span> and just produces text. The harness doesn&#8217;t get a vote.</span></p><h3><strong><span>What&#8217;s still missing</span></strong></h3><p><span>The multi-turn loop solves &#8220;one step isn&#8217;t enough.&#8221; But it has an obvious flaw: the model is purely reactive &#8212; wherever its eyes land, that&#8217;s where it goes next. No global view. Easy to wander off. A complex question that really should be researched along five dimensions might trap it on the first dimension for seven or eight searches while the others get forgotten. Or it picks a bad keyword early and turns spinning in the wrong direction, eating tokens and producing nothing.</span></p><p>Pure ReAct can keep moving, but it cannot guarantee that it is going anywhere. It has tactics, but no strategy. And an agent without a strategy is just a rather expensive way to walk in circles.</p><p>In <em>Part 2</em> of this series, we give it a plan.</p><div><hr></div><p style="text-align: center;"><em>Agentic Engineering grew out of a desire to create something useful for people trying to make sense of agentic AI and beyond. If you&#8217;d like to support that direction, please tell your friends about us.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share Agentic Engineering&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.agentengineering.co/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share Agentic Engineering</span></a></p><div><hr></div><p>That&#8217;s it for this one. We&#8217;ll pick up the conversation next week.</p><p>Until then, keep building.</p><blockquote><p><span>Tanya D&#8217;cruz</span><br><em>Editor-in-Chief</em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[Webinar: Building AI agent harnesses for finance]]></title><description><![CDATA[Nicole K&#246;nigstein on why reliable financial agents are an engineering problem, not simply a model problem]]></description><link>https://newsletter.agentengineering.co/p/free-webinar-building-ai-agent-harnesses</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/free-webinar-building-ai-agent-harnesses</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Sat, 22 Aug 2026 16:02:59 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/212138401/5d8761732d5f9606dd5ca28c84c57010.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div class="callout-block" data-callout="true"><h4><code>Editor&#8217;s note</code></h4><p><span>A quick programming detour from us this week.</span></p><p><span>If you regularly read Agentic Engineering, you know our core issues tend to follow a fairly deliberate cadence. This isn&#8217;t one of those issues. Yesterday, we hosted a free live session with </span><strong><a href="https://www.linkedin.com/in/nicole-koenigstein/"><span>Nicole K&#246;nigstein</span></a><span> </span></strong><span>on building AI agent harnesses for finance, and rather than let a useful conversation disappear into the webinar graveyard, we wanted to make the full recording freely available here.</span></p><p><span>Nicole is an AI researcher working on agentic systems and transformers, an instructor, consultant, and O&#8217;Reilly author. She teaches large language models and agentic architectures for organizations including the University of Oxford and O&#8217;Reilly Media, and much of her work focuses on the considerably harder problem that begins after an AI prototype works: making the resulting system reliable enough for production.</span></p><p><span>Enough preamble. Let&#8217;s get into the harness.</span></p></div><p><span>The conversation around agents has moved quickly from prompts to context, loops, orchestration, and now harnesses. But finance is a particularly unforgiving place to discover that a capable model and a reliable system are two very different things.</span></p><p><a href="https://www.linkedin.com/in/nicole-koenigstein/"><span>Nicole</span></a><span> spent the session unpacking what has to exist around the model: coordination, verification, observability, control, model selection, evaluation, and the boundaries between agents. Her central argument is a useful one for anyone building agents, whether or not you work in finance: </span><em><span>an agent is the model plus the harness.</span></em></p><p><span>And once you start thinking about agents that way, reliability becomes less about finding a smarter model and much more about engineering the system around it.</span></p><p><span>The video above contains the complete session and Q&amp;A. Below, we&#8217;ve turned the central ideas from Nicole&#8217;s talk into a written companion for those who would rather read or return to the concepts later.</span></p><div class="callout-block" data-callout="true"><h4>Want to go from understanding the harness to engineering one?</h4><p>This free webinar introduced the architecture behind reliable AI agents for finance. We see how the model, harness, coordination layer, and reliability mechanisms fit together. </p><p>On August 29, Nicole will take that foundation into <strong>Agentic AI for Finance</strong>, our four-day intensive workshop where you&#8217;ll move from understanding those systems to building them hands-on using real financial workflows.</p><p>We&#8217;ve broken down what you&#8217;ll build, what you&#8217;ll take away, and the other benefits of joining the intensive at the end of this article, so keep reading. Or, if you&#8217;re already convinced and would rather skip the rest of our pitch, you can book your place right away.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://luma.com/agentic-ai-for-finance&quot;,&quot;text&quot;:&quot;Book your seat now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://luma.com/agentic-ai-for-finance"><span>Book your seat now</span></a></p></div><h3><strong><span>The AI engineering problem has moved</span></strong></h3><p><span>For a long time, the dominant pattern in machine learning was fairly straightforward: task &#8594; model &#8594; output. If performance was poor, you improved the model.</span></p><p><span>Agents change that equation. Once tools, retrieval, memory, multiple agents, and decision points are incorporated into the architecture, the question shifts from &#8220;Can the model solve this task?&#8221; to &#8220;Can the system execute a stable sequence of decisions?&#8221;</span></p><p><span>That shift also explains the move from prompt engineering to context engineering and, now, harness engineering. Prompt engineering optimizes the input. Context engineering optimizes what the model sees. Harness engineering deals with everything that governs how the system actually operates: orchestration, tools, memory, retrieval, verification, governance, security, feedback, and execution loops.</span></p><p><span>The result is a different operating model: execute &#8594; observe &#8594; verify &#8594; correct &#8594; repeat. The model still matters, but it is now one component inside a much larger engineering problem.</span></p><h3><strong><span>Strong agents can still make a weak system</span></strong></h3><p><span>This was perhaps the most important warning in the session. It is tempting to assume that if every individual agent in a multi-agent system performs well, the resulting system should perform well too. It doesn&#8217;t necessarily follow.</span></p><p><span>Nicole points to the compounding reliability problem: when independent components execute sequentially, overall system success depends on the success probabilities across those components. Errors can compound as the chain becomes longer.</span></p><p><span>Imagine an agentic financial workflow involving ten decisions. Even if the individual components are reasonably capable, failures at one stage can affect everything downstream. Add retries as the default recovery mechanism, and the system may eventually produce an answer, but now you have another problem: cost and latency.</span></p><p><span>Nicole describes three consequences of trying to brute-force reliability through repeated attempts: more compute, more money, and more latency.</span></p><p><span>And average latency can hide the real production experience. A system with an acceptable mean response time but poor p95 or p99 latency can still feel broken to a meaningful portion of its users.</span></p><p><span>Reliability, in other words, cannot simply be retried into existence.</span></p><h3><strong><span>The difficult part often lives between agents</span></strong></h3><p><span>One of the more useful distinctions Nicole makes is between communication and coordination.</span></p><p><span>Communication is the exchange of messages. Coordination controls what the agents </span><em><span>do</span></em><span>. That matters because an agent handoff isn&#8217;t equivalent to a conventional function call.</span></p><p><span>When one agent hands something to another, the receiving agent has to </span><em><strong><span>interpret</span></strong></em><span> the message. A schema can ensure that the message has the right structure, but it cannot guarantee that the meaning survived the handoff intact.</span></p><p><span>The same applies to delegation. Calling a sub-agent isn&#8217;t simply calling another tool. You are introducing another probabilistic inference process into the system. That means the boundaries between agents deserve their own engineering attention.</span></p><p><span>Nicole&#8217;s argument here is particularly important for evaluation: if you evaluate only the final output, you may completely miss a coordination problem occurring within the system.</span></p><p><span>You need signals around the handoffs themselves. You need telemetry. You need to understand whether agents are communicating correctly, whether coordination is progressing as intended, and where meaning begins to drift.</span></p><p><span>As Nicole puts it, you can&#8217;t improve coordination using output-only feedback.</span></p><h3><strong><span>Maybe the model doesn&#8217;t need to change</span></strong></h3><p><span>This leads to a more interesting possibility.</span></p><p><span>A great deal of AI development still revolves around the model: use a more capable model, fine-tune it, or wait for the next generation. But Nicole&#8217;s research explores another direction: keep the model fixed and improve the harness around it.</span></p><p><span>The adjustable state of the system can move away from model weights and into the surrounding harness. Execution traces and trajectories can be scored, fed back into the system, and used to improve how it behaves.</span></p><p><span>Nicole took this further in her own open-source work.</span></p><p><span>Rather than hard-coding every coordination choice before the system has even executed a task, she began experimenting with </span><em><strong><span>learned coordination</span></strong></em><span>: allowing aspects of the harness to become adaptive, auditable, and transferable.</span></p><p><span>In her experiments, the benefits were particularly visible on coordination-heavy tasks, where she reported improvements in reliability and quality alongside lower costs. Simpler procedural tasks saw less improvement.</span></p><p><span>There is an important engineering principle hiding in there: not every task needs the same agent architecture.</span></p><h3><strong><span>Stop giving every job to your strongest model</span></strong></h3><p><span>That principle also applies to model selection. There is an understandable instinct when building agentic systems to reach for the strongest reasoning model available, particularly for high-stakes domains. But stronger and more expensive everywhere does not mean </span><em><span>better</span></em><span>.</span></p><p><span>In Nicole&#8217;s coding experiments, harder, coordination-heavy tasks benefited from a stronger planner, while simpler stages could be handled by cheaper models. She found that distributing models according to the demands of each stage could reduce token consumption and improve speed. More surprisingly, in her experiments, the resulting code was also judged better than code produced using the top-tier model at maximum effort throughout.</span></p><p><span>This isn&#8217;t an argument for universally choosing smaller models. It is an argument for </span><em><strong><mark data-color="#fff1f2" style="background-color: rgb(255, 241, 242); color: rgb(0, 0, 0);"><span>matching capability to responsibility</span></mark></strong></em><span>.</span></p><p><span>Nicole offered a useful mental model during the Q&amp;A: think about building a multi-agent system the way you would build a team.</span></p><p><span>If you were staffing the equivalent workflow inside an organization, you probably wouldn&#8217;t hire the same person for every role. Some jobs require deep expertise. Others require analysis, verification, classification, retrieval, or straightforward execution.</span></p><p><span>The skills should be complementary. Agent architectures can be designed the same way.</span></p><h3><strong><span>Coding agents need harnesses too</span></strong></h3><p><span>There was also a useful detour into coding agents.</span></p><p><span>Nicole described an experience in which an AI coding assistant, used without sufficiently strict instructions and boundaries, altered working agent code and introduced serious bugs and security problems.</span></p><p><span>Her takeaway was not that coding agents shouldn&#8217;t be used. Quite the opposite.</span></p><p><span>It was that using them effectively requires the same engineering discipline we expect elsewhere in an agentic system: clear instructions, guardrails, skills, hooks, review, and people who understand the system well enough to recognize when generated code is wrong.</span></p><p><span>There is a broader lesson here. AI can accelerate implementation, but it does not remove the need to know what good implementation looks like.</span></p><h3><strong><span>What this means for financial agents</span></strong></h3><p><span>Finance makes all of these problems harder to ignore.</span></p><p><span>A research assistant, portfolio-analysis system, financial-document agent, or other agentic workflow may need to retrieve evidence, coordinate multiple reasoning steps, invoke tools, maintain state, validate intermediate results, and operate within clearly defined constraints.</span></p><p><span>And probabilistic systems bring their own hazards.</span></p><p><span>During the Q&amp;A, Nicole specifically warned about forward-looking bias in financial applications. Models are extremely good at pattern matching, which can also lead them to derive patterns and make suggestions that are not adequately supported by the underlying facts.</span></p><p><span>So the problem is not merely whether an LLM knows finance.</span></p><p><span>The problem is whether the system surrounding it can constrain what it does, verify what it produces, expose how it arrived there, and respond sensibly when something goes wrong.</span></p><p><span>That is the territory of the harness.</span></p><h3><strong><span>The agent is bigger than the model</span></strong></h3><p><span>Nicole closed the main presentation with a progression worth keeping: executable reasoning &#8594; inspectable states &#8594; feedback-driven control &#8594; learned harness decisions.</span></p><p><span>It captures a shift we&#8217;re increasingly interested in at Agentic Engineering.</span></p><p><span>The model is becoming one component inside a much larger engineering problem.</span></p><p><span>The interesting work is moving outward: into the runtime, the context, the tools, the coordination layer, the evaluation system, the boundaries, the feedback loops, and the machinery responsible for keeping the whole thing under control.</span></p><p><span>That is especially visible in finance, where the cost of a system being </span><em><span>mostly right</span></em><span> can become very real, very quickly.</span></p><p><span>And perhaps that&#8217;s the simplest way to think about harness engineering: </span><em><strong><span>Capability comes from the model. Dependability has to be engineered around it.</span></strong></em></p><div><hr></div><h3 style="text-align: center;">The foundation is set. Now, build the thing.</h3><p>This free webinar covered the architecture. You learnt why a financial agent is not simply a model with a clever prompt, but a system in which the harness controls what the model can access, do, remember, verify, and prove.</p><p><a href="https://luma.com/agentic-ai-for-finance">Agentic AI for Finance</a>, Nicole&#8217;s four-day intensive, is where you <em>build</em> that system.</p><p>You&#8217;ll connect models to financial data, tools, and memory; engineer verification and governance into the workflow; and see what happens when all those neat boxes on an architecture diagram actually have to work together. The aim is to leave knowing how to build an agent that is useful, observable, and safe to operate.</p><h4 style="text-align: center;">What you&#8217;ll leave with</h4><p>Alongside the hands-on build experience, you&#8217;ll get:</p><ul><li><p>Four days of live access to the workshop with <a href="https://www.linkedin.com/in/nicole-koenigstein/">Nicole</a></p></li><li><p>Downloadable code templates, datasets, and architecture blueprints</p></li><li><p>Lifetime access to the workshop recordings</p></li><li><p>A free copy of <a href="https://www.amazon.com/Building-Agents-Finance-Financial-Architectures-ebook/dp/B0GSVS1RWJ/ref=sr_1_1?crid=QEEFABEOPZZG&amp;dib=eyJ2IjoiMSJ9.0QzpGX-Yg29AvYaA2xR0Uoont3_Dr4LuPNhLdN-6Q2RKP-ny11icnhiFuFPXBfO_qp2gjiTyi9GM-vKvSjwGPPdJDv8JS6200yVst1UacaJaUNuGX1QUCeEMEgPNbjSkl7LENiL000hGdNTfh616MKO4NKdrO_cge99u6aIcuKQ9a7c_oEE5iDaVYUQxG_kr8yLAVSVWF0uUEftc5t6HurrvEpQQc4DSacS6MoF6qOY.cGr83Bf3rNIVQkXoLRFxOjra7D2-ZipikTmpNF1FHZE&amp;dib_tag=se&amp;keywords=Building+AI+Agents+for+Finance&amp;qid=1787412046&amp;sprefix=building+ai+agents+for+finance%2Caps%2C349&amp;sr=8-1">Building AI Agents for Finance</a> to go deeper into the engineering covered during the program</p></li><li><p>A certificate on completion, plus practical experience relevant to work across quantitative research, fintech, portfolio management, and financial engineering</p></li></ul><p><a href="https://luma.com/agentic-ai-for-finance">Agentic AI for Finance</a> runs across August 29&#8211;31 and September 12&#8211;13, giving you time between sessions to digest what you&#8217;ve built before moving into the next stage.</p><p>So if you enjoyed Nicole&#8217;s webinar and want to take the next step from understanding the architecture to engineering it yourself, book your spot for Agentic AI for Finance below.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://luma.com/agentic-ai-for-finance&quot;,&quot;text&quot;:&quot;Book your spot now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://luma.com/agentic-ai-for-finance"><span>Book your spot now</span></a></p><div><hr></div><p>That&#8217;s it for this one. We&#8217;ll pick up our regular programming next week.</p><p>Until then, don&#8217;t just engineer the agent. Engineer the harness.</p><blockquote><p><span>Tanya D&#8217;cruz</span><br><em>Editor-in-Chief</em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[🎙️ Episode 3: The agent autonomy trap]]></title><description><![CDATA[Podcast with Ben Auffarth on bounded autonomy, calibrated confidence, meaningful human oversight and why businesses should fund the dataset before the agent]]></description><link>https://newsletter.agentengineering.co/p/episode-3-the-case-for-less-autonomous</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/episode-3-the-case-for-less-autonomous</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 20 Aug 2026 15:03:37 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/211965147/54ef00c61004d185e6b97d6116b1d464.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><span>AI agents are usually placed on an imaginary ladder. At the bottom, they assist. A little higher, they complete tasks. Eventually, they plan, act, learn and operate without us. Every step towards autonomy is treated as progress.</span></p><p><span>But autonomy is not free. The further a system can go on its own, the more room it has to drift, hallucinate, misinterpret a goal or make a consequential decision before someone notices.</span></p><p><span>For this episode of the Agentic Engineering podcast, I spoke with Ben Auffarth, Director and Chief Data Officer at Chelsea AI, founder of Chelsea AI Ventures, and author of several bestselling AI and machine learning books, including </span><em><a href="https://www.amazon.com/Generative-LangChain-production-ready-applications-LangGraph/dp/1837022011/ref=sr_1_1?crid=1LSOJHBEXGLG1&amp;dib=eyJ2IjoiMSJ9.4VGrkVel_z4pXJ-zkVpgjJr1H50j_XBLpXMI-7XsshRg5z71pjF0os_PUg7AO1PI32LlREYG-WwBXJgUT1OZ_n4lFFTr449UXkFMuuMjIMUPaVw0i4XsXNGlOu8eWcu-usB0ge4B7vDXLnpQ9kRYysDYHxgbJpVZ5vrE0SHQnnI_LqmewlYTwwureX1EYJ2wFzqLDOGSc-sGA8t3FHllt2Q0v8kYTKIDmfSQf-p-LRk._aKH_mgissnIRFrT9a8a7xUHLBMNkIuyNcXSNLmDZwo&amp;dib_tag=se&amp;keywords=Generative+AI+with+LangChain&amp;qid=1787222316&amp;sprefix=generative+ai+with+langchain%2Caps%2C505&amp;sr=8-1"><span>Generative AI with LangChain</span></a></em><span>, </span><em><a href="https://www.amazon.com/Machine-Learning-Time-Python-state-ebook/dp/B09GS44ZP4/ref=sr_1_1?crid=1GHH5HGC3AF2O&amp;dib=eyJ2IjoiMSJ9.V6T27nBIc2d7O4OGkXEeOvpXIYX7XuY2NsR33KqWlzXBZ71e4Z_AEkXN8gVk9qWRw_vm6SekP_dVBBzPpmZKbeW_iMsVrIW10S_ab2KBtnPpYfFo8VzsYxH76OslEVwQp7BSAZisc1dWDsu7vrY3G-DafGqhh39L01HOGT1UzNtadMmJ0opvIVc-cR2vYjR2ya2QvqwNiN8SiwpdRpW5Eas_kHN6nn63Om5heZ3y_Rk.9gIXJsJ_yLFyT0JQFST61OHmuYW1ElDDd00hZDa7Dv0&amp;dib_tag=se&amp;keywords=Machine+Learning+for+Time-Series+with+Python&amp;qid=1787222345&amp;sprefix=generative+ai+with+langchain%2Caps%2C395&amp;sr=8-1"><span>Machine Learning for Time-Series with Python</span></a></em><span> and </span><em><a href="https://www.amazon.com/Artificial-Intelligence-Python-Cookbook-algorithms-ebook/dp/B08GQ8TN8C/ref=sr_1_1?crid=Q18DQHWXHMFV&amp;dib=eyJ2IjoiMSJ9.q5a4atn9A3e4CGblFiKJPakEsugGAtmYvNPiT9tiX_zplXdGvb6HDaD231exozzMQ5dE0EBIel0LiB9e3YMl0cnm-fsMDWhUoD3TYTsdGY0OEQpSF4qrYFyx7BQtSp93l1HTBgXTtGGmaTGO64RQH-pA1k-jHjY_fgkaLes8cHjTxELb6rlYQMsCEAt3mm1EjoR6rjITnmqMQV2O9qqyQbFB_Z4oIAK_lLAfdnYNaTs.QfmD4a-40rKv1y8IoZt7TT8gOc3WXmcbAAenc6xX9qM&amp;dib_tag=se&amp;keywords=Artificial+Intelligence+with+Python+Cookbook&amp;qid=1787222386&amp;sprefix=machine+learning+for+time-series+with+python%2Caps%2C538&amp;sr=8-1"><span>Artificial Intelligence with Python Cookbook</span></a></em><span>. Ben holds a PhD in computational neuroscience and has built systems across finance, insurance, travel and other regulated environments.</span></p><p><span>Our conversation centred on a digital-footprint intelligence platform his team is building for estate agents. It is the kind of problem that sounds made for an autonomous agent: search for a person, gather information, assess the evidence and produce a useful profile.</span></p><p><span>Ben&#8217;s team made a different choice. They use LLMs extensively during development (to build datasets, run experiments, and improve the underlying system) but limit their authority when a real user is being evaluated.</span></p><p><span>That decision led us to a more useful question than simply asking &#8220;How autonomous can this agent become?&#8221;, which is </span>&#8594;</p><p><span>How much autonomy does the system actually need?</span></p><h3><strong><span>What does Neighborhood do, and where does AI fit into the system?</span></strong></h3><p><em><span>Neighborhood helps estate agents prioritise prospective property buyers. An enquiry may contain little more than a name, email address, telephone number and sometimes an address. The platform enriches those details with publicly available information, tries to verify that it has found the correct person and returns a profile with match scores and source information.</span></em></p><p><em><span>The difficult part is not finding someone with the right name; it is knowing whether the information belongs to this particular person. We use LLMs to help build datasets, extract information and run experiments quickly. At inference time, however, the LLM does not make the final matching or prioritisation decision. Those decisions are handled by more deterministic models and statistical methods. That is what I mean by bounded autonomy.</span></em></p><h3><strong><span>What tells the system that it has found enough information?</span></strong></h3><p><em><span>A general search engine may return many possible matches and leave someone to work through them. An LLM-based search tool can return a neat answer with a confidence percentage, but that number is not necessarily calibrated. If a model says it is 70 per cent confident, it does not mean that comparable answers are correct 70 per cent of the time.</span></em></p><p><em><span>We wanted confidence to have a measurable meaning. Neighborhood uses statistical principles, information theory and record-linkage methods to estimate the probability of a match. A common name in a large city begins with more uncertainty, and each additional piece of matching evidence adjusts that probability. The system stops relying on whatever the LLM &#8220;feels&#8221; and works with confidence that can be compared against data.</span></em></p><h3><strong><span>Can missing information become a hidden penalty?</span></strong></h3><p><em><span>Some people have extensive digital footprints, while others share very little or use platforms the system cannot access. The absence of information should not automatically become evidence that someone is suspicious or unlikely to proceed.</span></em></p><p><em><span>We found that a comprehensive profile does not necessarily mean someone is a stronger buyer. We therefore try not to downgrade people heavily when little information is available. More information may help an estate agent ask better questions, but that is a product assumption to test&#8212;not a universal truth. Sometimes the most accurate answer is simply that we do not know enough.</span></em></p><h3><strong><span>Is a more autonomous agent necessarily more advanced?</span></strong></h3><p><em><span>Autonomy is often presented as a measure of technical maturity. I do not think it is something to pursue for its own sake. The more autonomy we add, the more control we give up and the more variance we introduce.</span></em></p><p><em><span>Agents can look impressive in short demonstrations, but benchmarks involving longer tasks still show very low completion rates. If reliability matters, moving down the autonomy scale can be the better engineering decision. Use the LLM for tasks it performs relatively well, such as extracting information, and place deterministic decisions and constraints around it.</span></em></p><div class="callout-block" data-callout="true"><p>If you&#8217;re enjoying Ben&#8217;s approach to building AI systems that can actually be measured, he&#8217;s leading a hands-on workshop on August 29.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/the-genai-build-lab-build-production-ready-rag-on-a-budget-tickets-1994016271345?keep_tld=true&quot;,&quot;text&quot;:&quot;Take a look&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.eventbrite.co.uk/e/the-genai-build-lab-build-production-ready-rag-on-a-budget-tickets-1994016271345?keep_tld=true"><span>Take a look</span></a></p></div><h3><strong><span>What does meaningful human oversight look like?</span></strong></h3><p><em><span>Putting a person at the end of a workflow does not guarantee meaningful oversight. If the system is usually correct and its output looks professional, the reviewer may gradually stop examining the evidence and simply approve each recommendation.</span></em></p><p><em><span>In Neighborhood, we avoid asking whether an entire profile &#8220;looks right.&#8221; Instead, we ask the reviewer to identify a believable anchor, such as a company or location that connects the information to the correct person. They can also say they are unsure. Human oversight should contribute evidence or judgement the system does not possess; otherwise, it is merely rubber-stamping.</span></em></p><h3><strong><span>How do you evaluate whether the system&#8217;s confidence is justified?</span></strong></h3><p><em><span>A system may retrieve accurate information and still attach it to the wrong person. That is why we expose the evidence behind a profile, our confidence in each piece and its provenance. The estate agent can inspect the original sources and provide feedback, while the final decision remains with them.</span></em></p><p><em><span>The business measure is not whether a profile looks comprehensive. It is whether the system helps estate agents spend their time with the right prospective buyers. Evaluation must cover identity matching, evidence quality, confidence calibration and the usefulness of the eventual recommendation&#8212;not only the final score.</span></em></p><h3><strong><span>Are we confusing a successful demo with a reliable system?</span></strong></h3><p><em><span>Agent demonstrations usually show examples in which the system finds the right information, follows the right steps and produces a convincing answer. Production introduces ambiguous identities, inaccessible data, regional differences and inputs that look nothing like the demonstration.</span></em></p><p><em><span>At Neighborhood, we used agents and repeated runs to help create a labelled dataset, then combined those outputs with human evaluation. We train and tune more deterministic models against that data and track the results of each experiment. That may be less exciting than showing an agent browsing autonomously, but it gives us something we can measure and reproduce.</span></em></p><h3><strong><span>What should a business leader establish before deploying an agent?</span></strong></h3><p><em><span>I would ask for four things: a labelled evaluation set, a baseline to improve upon, an understanding of the cost of a wrong answer and a measurement of response variance. Run the same input several times and see whether the system gives the same answer. If it does not, average performance may conceal an agent that cannot be trusted on an individual decision.</span></em></p><p><em><span>Businesses often want to fund the agent before they fund the dataset. That is backwards. Without representative data, you cannot reliably build, compare or govern the system. In many AI projects, the dataset, not the model, is the real bottleneck.</span></em></p><h3><strong><span>How much autonomy should engineers give an agent?</span></strong></h3><p><em><span>There is a perception that because a system uses AI, less engineering is required. In practice, the opposite is often true.</span></em></p><p><em><span>You need to define the outcome clearly, establish a benchmark early, and constrain what the system is allowed to do. Deterministic checks and filters can prevent entire categories of failure. An orchestration framework such as LangGraph can also make the available paths explicit: some decisions can be made by the model, while others remain ordinary conditional logic.</span></em></p><p><em><span>The question is not whether the agent could make another decision autonomously. It is whether that autonomy improves the system enough to justify the additional variance and loss of control.</span></em></p><div><hr></div><p><span>Ben&#8217;s case for bounded autonomy is not an argument against agents. His team uses LLMs to create data, accelerate experimentation, extract information, and improve development. The argument is about placing autonomy where it creates value </span><em><span>and</span></em><span> removing it where consistency matters more.</span></p><p><span>That distinction becomes especially important now as agents move from helping people produce content or write code towards influencing decisions about buyers, customers, patients and applicants.</span></p><p><span>The most advanced system may not be the one permitted to do everything. It may be the one whose builders know exactly where the model is useful, where deterministic software should take over and where a person must still make the call.</span></p><p><span>Before asking how autonomous an agent can become, businesses should decide what happens when it is confidently wrong.</span></p><div><hr></div><p style="text-align: center;"><em>Less hype, more engineering conversations with the experts. </em></p><p style="text-align: center;"><em>Pull up a chair.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.agentengineering.co/subscribe?"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[#8: You can't trust what you don't evaluate]]></title><description><![CDATA[NVIDIA's Balamurugan Balakreshnan maps the evaluation practices that separate reliable AI agents from convincing demos]]></description><link>https://newsletter.agentengineering.co/p/you-cant-trust-what-you-dont-evaluate</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/you-cant-trust-what-you-dont-evaluate</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 13 Aug 2026 15:01:06 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/3305a72b-2824-4e2c-bdd6-4570f2f553da_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="callout-block" data-callout="true"><h4><code>Editor&#8217;s note</code></h4><p>Over the past few weeks, we&#8217;ve been talking about context engineering, loops, GraphRAG, and all the machinery that sits around modern AI agents.</p><p>This week, we&#8217;re deliberately taking a step back.</p><p>A year ago, most discussions focused on evaluating models. Today, more teams are discovering that evaluating agents is a very different problem. Once an AI system starts planning, calling tools, retrieving memory, and coordinating multiple steps, checking whether the final answer looks right simply isn&#8217;t enough. The execution path matters just as much as the destination.</p><p>So before we continue exploring more advanced engineering patterns, it felt worth revisiting one of the foundations.</p><p>In this issue, <a href="https://www.linkedin.com/in/balamurugan-balakreshnan/">Balamurugan Balakreshnan</a>, who works with enterprise AI teams at NVIDIA, walks through how he thinks about evaluation in agentic systems, why trust has to be engineered from the beginning, and which parts of an agent deserve measuring before production.</p></div><p>One of the biggest challenges I see in getting agentic AI systems into production isn&#8217;t the technology itself. It&#8217;s trust.</p><p>We&#8217;ve made enormous progress in building models that can reason, write code, call tools, and orchestrate complex workflows. But once those systems start making decisions on their own, a different question seems to start taking up space: How do we know we can trust what they&#8217;re doing?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Agentic Engineering! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Of course, the simplest answer is to have humans validate every output. But that will never sustain once you move beyond prototypes. If an agent is handling thousands of requests a day, human review becomes the bottleneck. So we need another way to build confidence.</p><h3>Trust has to be designed</h3><p>Models are trained on data that reflects the world at a particular point in time. We can improve their responses by providing fresh context during inference, whether that&#8217;s enterprise knowledge, live business data, or current events. That helps, but better context doesn&#8217;t automatically make an output trustworthy. Trust still has to be earned through <em><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">evaluation</mark></em>.</p><p>In my experience, the best place to start is during development, not after deployment. Before a model ever reaches production, it should be evaluated for things like correctness, relevance, completeness, grounding, and safety. None of this is easy. Good evaluation depends on good benchmark datasets, domain expertise, and clear success criteria. A coding model won&#8217;t be evaluated the same way as a reasoning model, and every organisation will have its own priorities beyond the standard metrics.</p><p>And the work doesn&#8217;t stop once the model ships.</p><p>Production systems need continuous evaluation because real users rarely behave as we expect. The data collected in production often tells you more about your system than anything you learned during testing. As models continue through fine-tuning, post-training optimisation, or other updates, both the data and the evaluation process evolve alongside them.</p><h3>Why evaluating agents is different</h3><p>Evaluating a model is one thing. Evaluating an agent is something else entirely.</p><p>An agent doesn&#8217;t just generate text. It interprets the user&#8217;s intent, builds a plan, decides which tools to call, retrieves information, and coordinates multiple actions before producing a result. In many systems, there&#8217;s also a planning or supervisor agent deciding which specialist agents, MCP servers, or data sources should be involved. In multi-agent systems, every additional agent introduces another layer of decision-making.</p><p>That&#8217;s why evaluating the final answer is no longer enough.</p><p>We also need to understand <strong>how the agent arrived there</strong>.</p><p>When an agent runs in production, I think about questions like these:</p><ul><li><p>Was the execution plan appropriate for the user&#8217;s request?</p></li><li><p>Did it choose the right specialist agents?</p></li><li><p>Did it select the right tools or MCP servers?</p></li><li><p>Did it retrieve the right information from memory or external knowledge?</p></li><li><p>If a judge model was involved, has that model itself been properly evaluated?</p></li></ul><p>These questions matter because agents are making decisions continuously as they execute a workflow. Trust comes not <em>just</em> from the quality of the final answer, but from confidence that the system made good decisions <em><mark data-color="#fce5cd" style="background-color: rgb(252, 229, 205); color: rgb(0, 0, 0);">all the way through</mark></em>.</p><h3>So what should we actually evaluate?</h3><p>This is the question I get asked most often. My honest and perhaps bland answer is: <strong>it depends</strong>.</p><p>Every application is different. The metrics you care about will vary depending on your domain, the framework you&#8217;re using, and what success looks like for your users. There isn&#8217;t a universal scorecard for agentic AI.</p><p>That said, I&#8217;ve found there are a handful of areas that every team should think about.</p><h4>Task performance</h4><p>The first question is the simplest one. <strong><span>Did the agent accomplish what it was asked to do in the first place?</span></strong></p><p>That sounds obvious, but it&#8217;s surprisingly easy to focus on intermediate steps and forget the end goal. At a minimum, I&#8217;d want to understand things like:</p><ul><li><p>Goal achievement</p></li><li><p>Subtask completion rate</p></li><li><p>End-to-end success rate</p></li><li><p>First-attempt success rate</p></li><li><p>Human evaluation</p></li></ul><h4>Reasoning quality</h4><p>A correct answer isn&#8217;t always the result of good reasoning. Sometimes an agent simply gets lucky.</p><p>Because agents plan before they act, it&#8217;s important to evaluate the quality of that planning as well. Questions worth asking include: Was the plan coherent? Did it retain the right context? Could it recover when something went wrong? Did it follow instructions, or did it drift off course?</p><p>Some useful metrics here include:</p><ul><li><p>Plan coherence</p></li><li><p>Step accuracy</p></li><li><p>Self-correction ability</p></li><li><p>Hallucination rate</p></li><li><p>Context retention</p></li><li><p>Decision quality</p></li><li><p>Instruction adherence</p></li></ul><h4>Tool selection and orchestration</h4><p>Most production agents don&#8217;t work alone. They&#8217;re constantly interacting with APIs, databases, MCP servers, and external tools.</p><p>Choosing the wrong tool can be just as damaging as producing the wrong answer, so it&#8217;s important to evaluate how well the agent orchestrates the systems around it.</p><p>Some useful measures include:</p><ul><li><p>Tool call accuracy</p></li><li><p>Tool selection precision</p></li><li><p>Parameter correctness</p></li><li><p>API call success rate</p></li><li><p>Error recovery rate</p></li><li><p>Orchestration efficiency</p></li><li><p>Unnecessary tool invocation rate</p></li></ul><h4>Safety and reliability</h4><p>No production system is complete without guardrails.</p><p>Agents should stay within their intended scope, handle failures gracefully, and avoid actions that could compromise users or data. Depending on your application, you might look at:</p><ul><li><p>Refusal accuracy</p></li><li><p>Guardrail compliance</p></li><li><p>Harmful action rate</p></li><li><p>Scope creep</p></li><li><p>Failure mode coverage</p></li><li><p>Graceful degradation</p></li><li><p>Uptime and availability</p></li><li><p>Data leakage incidents</p></li></ul><h4>Cost and efficiency</h4><p>Finally, there&#8217;s the operational side of the equation.</p><p>An agent that produces excellent results but takes five minutes and thousands of tokens to complete a simple task may not be practical in production. That&#8217;s why it&#8217;s worth tracking metrics such as:</p><ul><li><p>Latency per task</p></li><li><p>Token usage</p></li><li><p>Cost per successful task</p></li><li><p>Number of execution steps</p></li><li><p>Retry overhead</p></li><li><p>Throughput</p></li><li><p>Memory efficiency</p></li></ul><p>These aren&#8217;t meant to be a checklist that every team must follow. Think of them as a menu rather than a mandate. Different frameworks expose different signals, and different applications will naturally prioritise different measures.</p><p>The important thing is to be deliberate about what you&#8217;re evaluating. If you only measure the final response, you&#8217;ll miss most of what determines whether an agent is actually reliable.</p><p>For me, that&#8217;s the biggest shift agentic AI introduces.</p><p>Trust doesn&#8217;t come from a good answer alone. It comes from understanding how that answer was produced. That means evaluating the planning process, the decisions made along the way, the tools that were selected, the knowledge that was retrieved, and the actions the agent ultimately took.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ai8G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ai8G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Ai8G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Ai8G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Ai8G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ai8G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2199721,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.agentengineering.co/i/210883386?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ai8G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Ai8G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Ai8G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Ai8G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eb033c2-4420-45d0-9e48-00d6e1433cd9_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That&#8217;s it for this one. We&#8217;ll pick up the conversation next week.</p><p>Until then, keep building.</p><blockquote><p><span>Tanya D&#8217;cruz</span><br><em>Editor-in-Chief</em></p></blockquote><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Agentic Engineering! Subscribe for free to receive new posts every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[🎙️ Episode 2: An agent loop needs a way out]]></title><description><![CDATA[Podcast with Rick Hightower on verification, stopping conditions, and the harnesses that keep autonomous agents honest]]></description><link>https://newsletter.agentengineering.co/p/9-an-agent-loop-needs-a-way-out</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/9-an-agent-loop-needs-a-way-out</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 06 Aug 2026 15:03:48 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/210053880/dd364856108051a51d6d7f09890a8896.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>An agent loop looks simple when it is drawn on a slide. Give the model a goal, let it take an action, check the result, and repeat until the work is done. The difficulty, however, hides inside the word &#8220;until.&#8221;</p><p>How does the system know that it is getting closer to the goal rather than repeating itself? Who decides whether the result is correct? What prevents the agent from changing the test until its own work passes? And if the model can continue taking actions, spending tokens, and modifying files, what finally makes it stop?</p><p>For this episode of the Agentic Engineering podcast, I spoke with <a href="https://www.linkedin.com/in/rickhigh/">Rick Hightower</a>, a senior AI engineer and certified Claude architect who has spent more than 20 years building software across companies including Apple, Capital One, and the NFL. Rick now focuses on the less-visible parts of agentic systems: harnesses, context, state, verification, and loops.</p><p>Rick began working with AI before today&#8217;s command-line agents and coding harnesses were available. The models could suggest what to do, but he still had to move the output into the development environment, run the tests, inspect the result, and decide what should happen next. As he puts it, he was the loop.</p><p>Today, harnesses can take on more of that responsibility. But before we remove the person who was checking the work and deciding when to continue, we need to understand what will replace them.</p><p>So we began with the most basic question: what actually turns a series of prompts into a loop?</p><div class="callout-block" data-callout="true"><h4><code>Editor&#8217;s note</code></h4><p>We&#8217;ve tucked a subscriber-only perk into this edition. If Rick&#8217;s ideas resonate with you, keep reading&#8230; the private code is waiting below.</p></div><h3>What turns a series of prompts into a loop?</h3><p><em>Before command-line agents and coding harnesses existed, I was the loop. I would ask ChatGPT or Claude to research an approach or help port code, move the result into the development environment, run it, inspect what happened and then decide what to try next.</em></p><p><em>A real agent loop automates that feedback process. It has an expected outcome, a way to verify the result, and a mechanism for returning useful feedback when the result fails. It continues until the work meets the stated criteria or the system determines that it cannot complete the task with the information and tools available.</em></p><p><em>I use Mermaid diagrams as a simple example. The agent can generate a diagram, lint the syntax, and correct any errors. Once it renders, another model can inspect the image and compare it with the original requirements. The first check tells us whether the diagram is valid. The second tells us whether it actually represents what we asked for.</em></p><p><em>Without verification and feedback, we do not really have a loop. We have a sequence of prompts.</em></p><h3>Who verifies the verifier?</h3><p><em>If the same agent interprets the requirement, writes the code and creates the test, it may define success in a way that agrees with its own implementation. It may even change the test to make its work pass.</em></p><p><em>The harness should control what the agent is allowed to modify. Permissions can prevent it from editing protected tests, and hooks can detect when it tries. There may be legitimate reasons to update a test, but that change should be explained and, in higher-risk situations, reviewed by a person.</em></p><p><em>Wherever possible, I prefer deterministic verification: does the code compile, does the test pass, does the output match the schema, or does the file pass a linter? When that is not possible, I use a separate judge or adversarial subagent.</em></p><p><em>The judge should receive the requirements and the final output, but not necessarily the full context of the agent that produced it. If both agents share the same reasoning history and assumptions, the judge may inherit the same bias. I think of it as the difference between the person doing the work and the teacher grading it.</em></p><h3>How does a loop know when it is stuck?</h3><p><em>A loop can continue producing new output without making meaningful progress. It may repeat the same action, revisit an approach that already failed or improve one part of the result while breaking several others.</em></p><p><em>I have encountered this while tuning systems that extract structured information from documents. The loop may be working towards a target accuracy across several categories. One attempt improves one category but causes the others to fall. The next attempt reverses the change. The system is doing work, but the overall result has stopped improving.</em></p><p><em>The harness needs an exit condition for that situation. It might stop after four or five attempts without progress, after a fixed number of turns or after reaching a token or cost limit. It should also recognise when it cannot succeed because data is missing, a downstream system is unavailable, or the available tools are not enough.</em></p><p><em>Ideally, the loop stops because the work is complete. But production systems also need a responsible way to stop when completion is not possible. At that point, the agent should change strategy, ask for help, or hand the problem to a person.</em></p><div class="callout-block" data-callout="true"><p style="text-align: center;">If this conversation has you wondering whether your agent is a reliable system or just a prompt with permissions, Rick is teaching a live, hands-on workshop on August 29.</p><p style="text-align: center;">In <em><a href="https://www.eventbrite.co.uk/e/engineering-reliable-agentic-ai-systems-tickets-1992373400474?aff=agenticeng&amp;discount=AE40">Engineering Reliable Agentic AI Systems</a></em>, you&#8217;ll spend four hours getting into the machinery behind production agents. And here&#8217;s something just for the Agentic Engineering circle: subscribers get <strong>40% off</strong> the ticket price.</p><p style="text-align: center;">Your private code is <strong>AE40</strong>. Use it at checkout.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/engineering-reliable-agentic-ai-systems-tickets-1992373400474?aff=agenticeng&amp;discount=AE40&quot;,&quot;text&quot;:&quot;Claim your place&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.eventbrite.co.uk/e/engineering-reliable-agentic-ai-systems-tickets-1992373400474?aff=agenticeng&amp;discount=AE40"><span>Claim your place</span></a></p></div><h3>Are we blaming models for failures caused by the harness?</h3><p><em>When an agent behaves badly, the model usually receives the blame. But many failures begin elsewhere in the system.</em></p><p><em>A team adds a feature and changes the prompt. That new feature works, but three older behaviours stop working. A model provider releases an update, the retrieval system starts loading different context, or someone changes the few-shot examples. The output becomes worse, and it looks as though the model has suddenly lost a capability.</em></p><p><em>The most common problem I have seen is drift. Agent systems have many moving parts, and changing one of them can create regressions somewhere else. That is why teams need a known set of inputs, expected outputs, and a way to grade them after every meaningful change.</em></p><p><em>I describe the principle as &#8220;slow down and go fast.&#8221; If you can detect regressions early, you can move quickly without repeatedly breaking the behaviours you already depend on. Otherwise, developing these systems can feel like nailing Jell-O to a wall.</em></p><p><em>The harness controls the context, tools, permissions, state, verification, and stopping conditions around the model. Before replacing the model or rewriting the prompt, inspect those layers.</em></p><h3>Where does memory live when the model has none?</h3><p><em>LLMs are stateless, but a useful agent needs to know what has already happened, what remains unresolved, and whether its latest action changed anything.</em></p><p><em>I separate memory and state into several categories. There is conversational memory, which may be compressed into a summary. There are important facts that should be stored explicitly because summarisation can lose them. There are instructions and procedures in files or agent skills. There is information retrieved from a database or RAG system when the task requires it. And there are plans, tickets, logs, and repository history that allow the work to continue beyond one session.</em></p><p><em>The goal is not to keep everything in the active context. If the context becomes too large, important information gets buried, and the model begins forgetting things it would normally handle correctly.</em></p><p><em>I prefer to keep the main agent focused and use subagents for narrow tasks. Once one task is complete, I can clear the context and begin the next one using the durable information stored in the specifications, plans, and tickets.</em></p><p><em>If you cannot clear the context because the agent would forget everything it has been doing, the important state probably needs to be written down somewhere outside the conversation.</em></p><div class="callout-block" data-callout="true"><h3>Keep up with Rick</h3><p>Rick publishes practical deep dives on harness engineering, loop engineering, agent memory, context, and production AI systems.</p><p>Find him on <a href="https://medium.com/@richardhightower">Medium as </a><strong><a href="https://medium.com/@richardhightower">@richardhightower</a></strong>, or subscribe to <em><strong><a href="https://rickhigh.substack.com/">Hightower&#8217;s AI Harness Engineering</a></strong></em> on Substack.</p></div><h3>Is a specification more than a very long prompt?</h3><p><em>A prompt usually tells the agent what to do next. A specification describes the system, the work that needs to be completed, the constraints, and the conditions that should be true at the end.</em></p><p><em>That makes the specification the fuel for a long-running loop. I have given agents several tickets, asked them to complete the work, launch the application, visit the relevant pages and compare the result with the agreed mock-ups. The agent can run for hours because it has both a roadmap and a definition of done.</em></p><p><em>But a long specification is not automatically a good specification. If the plan is incomplete or wrong, the agent may follow it perfectly and still produce the wrong result. Garbage in, garbage out does not disappear because AI is doing the implementation.</em></p><p><em>For important or complicated plans, I often ask another agent&#8212;sometimes a different model, sometimes the same model in an isolated context&#8212;to look for missing assumptions and unclear requirements. A second set of eyes usually finds something.</em></p><p><em>The specification should also be visible to the human stakeholders who understand the intended outcome. An agent can critique a plan, but it cannot recover business context that nobody gave it.</em></p><h3>Can the same specification safely drive action and verification?</h3><p><em>Using one specification to tell the agent what to build and then using that same specification to judge the result is efficient, but it creates a shared point of failure.</em></p><p><em>If the specification misses a requirement, both the builder and the verifier may agree that the work is complete. The agent has built exactly what it was asked to build, but the requested system was wrong.</em></p><p><em>I still use the specification as the basis for verification, but I try to support it with independent evidence. That may include sample inputs from product management, approved mock-ups, existing tests, business rules, or review from an agent that did not participate in writing the implementation.</em></p><p><em>The agent should also be able to question the specification. If it finds a contradiction or discovers that an acceptance criterion cannot be satisfied, continuing blindly is not useful autonomy. It should report the problem and bring a person back into the loop.</em></p><p><em>The specification guides the work. It should not become something the system is forbidden to challenge.</em></p><h3>How do you verify an autonomous research loop?</h3><p><em>Research does not have the same natural completion signal as code. There may always be another source to read, another reference to follow, or another direction to investigate.</em></p><p><em>The research loops I have built usually begin with an agreed outline and a defined artifact, such as a document, image or presentation. The system can then validate whether the artifact follows the outline, whether the facts have been checked and whether it meets the required grammar, style and quality standards.</em></p><p><em>In that case, the verification gate is the rubric agreed upon before the research begins. The loop is complete when the artifact meets those requirements&#8212;not when the agent has exhausted every possible source.</em></p><p><em>Some research and experimentation also have external signals. In marketing, an agent can test different advertisements, measure clicks or conversions, and use that result to decide what to try next. AI could run far more of those experiments than a human marketing team could manage manually.</em></p><p><em>Open-ended research is more difficult. A rubric can tell us whether an artifact meets its requirements, but it cannot prove that no useful information remains undiscovered. We should be honest about that rather than pretending every kind of research can be given a perfectly deterministic finish line.</em></p><h3>What did production force you to change your mind about?</h3><p><em>Verification is expensive. Once you add adversarial subagents, judges, and feedback loops, a workflow may cost five or ten times more than a single model call.</em></p><p><em>But the useful comparison is not between a verified system and the cheapest possible workflow. It is between the cost of verification and the cost of delivering the wrong result.</em></p><p><em>If I can verify something with a compiler, test, or linter, I will use that. When I cannot, I may use a rubric and another model. That increases inference costs, but it also produces much better results. A brand violation, a broken production change, or an incorrect high-stakes answer can cost far more than the additional model calls needed to catch it.</em></p><p><em>If I had only one hour to design a new loop, I would begin with the verifier and the stopping conditions. What does done look like? What evidence proves it? How many attempts can the agent make? What happens if it stops improving?</em></p><p><em>I would also define the permissions and boundaries before allowing the loop to run. I would not run an unfamiliar autonomous agent with unrestricted access to the same machine that holds my credentials, personal files, or business systems. Use a sandbox, virtual machine, or managed environment, and give the agent only the tools it genuinely needs.</em></p><div><hr></div><p>Rick&#8217;s case for loop engineering is not that we should let agents run for as long as possible. It is that agents become more useful when the system around them can recognise success, detect failure and stop them safely when they are no longer making progress.</p><p>The model may perform the visible work, but the harness determines the conditions under which that work happens. It supplies the context, preserves the state, controls the tools and decides when a person needs to return.</p><p>Before giving an agent a longer task, a larger context window or more powerful tools, the most useful place to begin may be the end: decide what &#8216;done&#8217; means, how you will recognise it and what the system should do when it cannot get there.</p><div><hr></div><p style="text-align: center;"><em>Enjoyed today&#8217;s podcast and want more honest engineering conversations? </em></p><p style="text-align: center;"><em>Pull up a chair </em>&#8595;</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.agentengineering.co/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p>That&#8217;s it for this one. We&#8217;ll pick up the conversation next week.</p><p>Until then, keep building.</p><blockquote><p><span>Tanya D&#8217;cruz</span><br><em>Editor-in-Chief</em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[#7: The undeclared dependency]]></title><description><![CDATA[Every team building with AI coding agents is carrying a dependency it never declared. Here's what happens when you finally do.]]></description><link>https://newsletter.agentengineering.co/p/8-the-undeclared-dependency</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/8-the-undeclared-dependency</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 30 Jul 2026 14:30:07 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8e09a55e-8bbc-45d4-98f1-f694cec1b584_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="callout-block" data-callout="true"><h4><code>Editor&#8217;s note</code></h4><p>We&#8217;ve spent time over the past few weeks talking about prompts, context, harnesses, and loops. But there&#8217;s another part of the AI engineering stack that doesn&#8217;t get as much attention: the context your coding agents depend on.</p><p>This week, <strong><a href="https://www.linkedin.com/in/webmax/">Maxim Salnikov</a></strong> explores why agent context deserves to be treated like any other software dependency, and how Agent Package Manager (APM) aims to solve the configuration drift that&#8217;s emerging across AI development teams.</p><p>If you&#8217;d like to go deeper after reading, Maxim has also put together a <strong>free interactive guide</strong> on Agent Package Manager, which we&#8217;ve linked at the end of this article. I&#8217;ll leave the floor to him.</p></div><p><span>You know the feeling, because you&#8217;ve felt it before&#8230; in a different decade. A new engineer joins, and getting their environment right means a wiki page, a Slack thread, and copying files out of someone&#8217;s home directory. Two developers swear they have the same setup, but their tools behave differently, and nobody can say why.</span></p><p><span>Before npm, pip, and Cargo, that was simply how software felt. Dependencies lived wherever each developer happened to put them &#8212; a </span><code>lib/</code><span> folder copied between machines, a README that said &#8220;install these seven things first.&#8221; It mostly worked, until it didn&#8217;t. Package managers ended it by turning scattered files into a </span><em><span>declared</span></em><span> dependency: one manifest that names what you need, one lockfile that pins exactly what you got, one command that restores it anywhere.</span></p><p><span>The files that configure your AI coding agents are in that same pre-package-manager state right now.</span></p><h3><strong><span>The context you never treated as code</span></strong></h3><p><span>Call it your </span><em><span>agent context</span></em><span>: the instructions and coding standards, the reusable prompts, the skills and agent personas, the plugins, the Model Context Protocol (MCP) server configurations. It shapes the code an agent writes, how it reviews a diff, and which external tools it can reach. That makes it a real project dependency &#8212; one of the highest-leverage ones you have, because it multiplies across every line the agent touches.</span></p><p><span>Yet most teams hand-maintain it as loose files scattered across per-tool directories and personal machines, with no declared source of truth. So it drifts. One tool&#8217;s instructions say to use the Money value object; another still mentions raw decimals. Three copies of the same review prompt live in three home directories, each checking a </span><em><span>different</span></em><span> threat model. Neither is wrong. Neither is portable.</span></p><p><span>A shared README doesn&#8217;t save you: documentation describes </span><em><span>intent</span></em><span> but never restores </span><em><span>state</span></em><span>, so the drift returns the moment someone edits a local file. This is &#8220;works on my machine,&#8221; aimed squarely at your agents &#8212; and it&#8217;s invisible, because nothing broke. The build is green. The agent just quietly got a little less consistent, or a little more dangerous, on one laptop.</span></p><h3><strong><span>Three files and one habit</span></strong></h3><p><span>The fix is not a new methodology. It&#8217;s the boring, proven shape that worked the first time. </span><a href="https://github.com/microsoft/apm"><span>Agent Package Manager (APM)</span></a><span>, an open-source tool from Microsoft, borrows the manifest-plus-lockfile pattern and points it at agent context. There are three files and one habit:</span></p><ul><li><p><code>apm.yml</code><span> &#8212; the manifest. The single human-authored source of truth, naming the primitives and MCP servers your project depends on, and which tools to target.</span></p></li><li><p><code>apm.lock.yaml</code><span> &#8212; the lockfile. Machine-generated, never hand-edited, pinning every dependency to an exact source ref </span><em><span>and</span></em><span> a content hash, so two developers install byte-identical context.</span></p></li><li><p><code>apm-policy.yml</code><span> &#8212; install-time governance, for when you&#8217;re ready for it.</span></p></li><li><p><code>apm install</code><span> &#8212; the habit. The one command that reads the manifest and materializes the declared context into each tool&#8217;s native location.</span></p></li></ul><p><span>The move that makes it click is </span><em><span>materialization</span></em><span>. A manifest is a recipe, not a meal; apm install turns the declaration into the real files each tool reads natively &#8212; compiled into .github/ for Copilot, .claude/ for Claude Code, .cursor/ for Cursor &#8212; then stays out of the way at runtime. Your three-harness maintenance burden becomes one.</span></p><h3><strong><span>A 60-second quick start</span></strong></h3><p><span>Enough theory. Here is the entire loop against a real, public collection &#8212; </span><a href="https://github.com/webmaxru/web-ai-agent-skills#install-individual-skills"><span>web-ai-agent-skills</span></a><span>, a maintained set of skills for the browser&#8217;s built-in AI APIs. Every command below is in the </span><a href="https://microsoft.github.io/apm/"><span>APM docs</span></a><span>:</span></p><blockquote><p><code># 1. Install the CLI (macOS/Linux; use the PowerShell one-liner on Windows)</code></p><p><code>curl -sSL https://aka.ms/apm-unix | sh</code></p><p><code># 2. In your project, once:</code></p><p><code>apm init</code></p><p><code># 3. Add a skill as a pinned dependency:</code></p><p><code>apm install webmaxru/web-ai-agent-skills/skills/prompt-api</code></p></blockquote><p><span>That last line is the whole pitch. It resolves the skill straight from its git source, pins it in apm.lock.yaml by commit </span><em><span>and</span></em><span> content hash, and compiles it into the location your agents already read. Commit the two files, and the next developer who runs apm install gets byte-identical context. The Prompt API skill is a </span><em><span>declared dependency</span></em><span> now, not a snippet someone pasted from a stale gist &#8212; and swapping prompt-api for any other entry works the same way.</span></p><h3><strong><span>What actually changes</span></strong></h3><p><span>Onboarding stops being a ritual and becomes git clone then apm install &#8212; the new hire&#8217;s agent is configured identically to yours before their first coffee. Drift stops being something you find during a security review: the lockfile makes it impossible, reproducing known-good context down to the bytes. None of this makes your agents smarter. That&#8217;s the point &#8212; it makes them </span><em><span>consistent</span></em><span>, the precondition for every other improvement, since you can&#8217;t tune a context you can&#8217;t reproduce.</span></p><p><span>The uncomfortable part is that we already learned this lesson once. Agent context deserves the same declaration and review discipline you already demand of package.json or requirements.txt. The tools exist. The only question is how long you keep paying the drift tax before you declare the dependency you already have.</span></p><p><span>And if authoring that first manifest feels like one more chore, there is a fitting last move: let an agent do it. The </span><a href="https://github.com/webmaxru/ai-native-dev#agent-package-manager"><span>Agent Package Manager skill</span></a><span> teaches your coding agent to run </span><code>apm init</code><span>, add and pin dependencies, validate the manifest, and manage the lockfile &#8212; installed, of course, exactly like any other skill:</span></p><blockquote><p><code>apm install webmaxru/ai-native-dev/skills/agent-package-manager</code></p></blockquote><p><span>Which is the whole idea coming full circle: the package manager for agent context, installed and operated by the agent itself.</span></p><div class="callout-block" data-callout="true"><p>This article only scratches the surface. Maxim&#8217;s <em>The Missing Package Manager</em> expands on these ideas through an interactive guide that explains not just how to use Agent Package Manager, but why treating agent context as a declared dependency fundamentally changes the way teams build with AI coding agents.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://apm.isainative.dev/&quot;,&quot;text&quot;:&quot;Read the free guide&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://apm.isainative.dev/"><span>Read the free guide</span></a></p></div><p>That&#8217;s it for this one. We&#8217;ll pick up the conversation next week. </p><p>Until then, keep building.</p><blockquote><p>Tanya D&#8217;cruz<br><em>Editor-in-Chief</em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[🎙️ Episode 1: GraphRAG as an interface to the world]]></title><description><![CDATA[Podcast with David Knickerbocker on teaching GraphRAG agents when to stop, what to remember and how to investigate]]></description><link>https://newsletter.agentengineering.co/p/7-graphrag-as-an-interface-to-the</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/7-graphrag-as-an-interface-to-the</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 23 Jul 2026 14:30:51 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207954133/3a2748ce5bda0b2ef311de88ba082734.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>GraphRAG is often presented as the natural next step after conventional RAG: add a knowledge graph, give an agent access to it, and let the system investigate more complex questions.</p><p>But adding a graph does not remove the difficult questions. How does an agent know when it has gathered enough evidence? What happens when relationships change, memories become unreliable, or the graph itself sends the model down the wrong path? And when is a graph genuinely the right architecture rather than simply the more fashionable one?</p><p>For the first episode of the Agentic Engineering podcast, I spoke with <a href="https://www.linkedin.com/in/dkjapan/">David Knickerbocker</a>, founder and intelligence architect at Verdant Intelligence and author of <em><a href="https://www.amazon.com/Network-Science-Python-networks-analysis/dp/1801073694/ref=sr_1_1?crid=2ZBK48NRQ7S1V&amp;dib=eyJ2IjoiMSJ9.qkYJqZWvXvM46TpRPZVX00zD6PYxP4QqeyoXBSV48Wl85fVrWknFKeI8jRsNsNAWFnSHcTb3-rJ31-mB4ga3LeN61BxxjY5JTQ8i0NyFkzXrzsVcDmu5-oByCFGfZHakkqQztDx23Zj6I8gJtbksdUHzKW8aSu-Mf9gaB6QlL0b4CBu5cO7HrsE7c83C7xOJbkfSn6tkEtum5U5WGN1vCBuwwgM854x2WLmm4XGrfm0.u7sYg3HtA5ooYVENw5qK4evQecenoYXT_p8Kx1b8rl8&amp;dib_tag=se&amp;keywords=Network+Science+with+Python&amp;qid=1784646603&amp;sprefix=network+science+with+python%2Caps%2C325&amp;sr=8-1">Network Science with Python</a></em>. David has spent decades working across databases, cybersecurity, open-source intelligence, and network science. His answers return repeatedly to one idea: GraphRAG and agents are still software systems. If we want to trust them, we need to make their evidence visible, test their individual parts, and resist rebuilding systems that already work.</p><div><hr></div><h3>When do you genuinely need GraphRAG?</h3><p><em>I don&#8217;t see graphs as an inherently more complicated alternative to RAG. Networks already exist in nature, organisations, and social systems; graphs simply represent how things connect. They can feel difficult because fewer teams have experience with network science, but the basic idea&#8212;one thing connected to another&#8212;is straightforward.</em></p><p><em>That does not mean every team should replace its RAG system. If you have invested in a mature similarity-based system and it solves the problem well, there is no reason to solve the same problem twice. Use GraphRAG when relationships, paths, and connected evidence are central to the problem, and your team is prepared to work with them. If the existing system works, keep it.</em></p><h3>Does a natural-language GraphRAG endpoint remove complexity, or merely hide it?</h3><p><em>A natural-language endpoint should hide implementation work without hiding the evidence. I design systems to return the context behind an answer: the source material, the relevant data, and the reasoning behind a predicted relationship. For example, when one of my tools converts text into a graph, it preserves both the original passage and the reason an edge was created.</em></p><p><em>There will always be some trust involved when an agent uses an external service, just as there is when we use a search engine or API. That trust should develop through repeated, dependable use&#8212;not blind faith. Developers may not need to understand the database schema or write Cypher, but they should still receive enough information to verify what the endpoint returned.</em></p><div class="callout-block" data-callout="true"><p style="text-align: center;">If you&#8217;re liking this conversation so far, David is teaching a live, hands-on bootcamp on July 31. </p><p style="text-align: center;">In <em>Building Intelligent AI Agents with GraphRAG</em>, you&#8217;ll work with a production-ready natural-language GraphRAG endpoint and progressively build an agent that can break down complex questions, gather evidence across multiple graph queries, remember previous investigations, and extend its reasoning with specialised analytical tools. You&#8217;ll leave with practical patterns for building agents that work across connected knowledge rather than simply retrieving documents.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/building-intelligent-ai-agents-with-graphrag-tickets-1992756563525&quot;,&quot;text&quot;:&quot;Register here&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.eventbrite.co.uk/e/building-intelligent-ai-agents-with-graphrag-tickets-1992756563525"><span>Register here</span></a></p></div><h3>How does an investigative agent know when it has investigated enough?</h3><p><em>We need to teach agents something my managers used to tell me: good enough is good enough. An agent should have a defined purpose, a limited mandate, and a clear point at which it reports back. An agent checking for concerts does not need to investigate indefinitely; it can check the relevant venues and artists, return the results, and stop.</em></p><p><em>An open-source intelligence agent may need to search more widely, compare conflicting sources, and detect manipulation, so its definition of &#8220;enough&#8221; will be different. The stopping condition has to match the task. The important thing is to decide what good enough means before the agent begins, rather than allowing it to keep reasoning and spending without a boundary.</em></p><h3>How should GraphRAG represent information that changes over time?</h3><p><em>I treat graphs as temporal. Information changes, and the completeness of our information changes with it. Rather than overwriting an old relationship, I would preserve what was claimed, where it came from, and when it was recorded. The history is part of the information.</em></p><p><em>In my systems, the highest-level object is a claim rather than a fact because the internet contains many competing claims and facts require judgement. A graph should preserve those differences instead of making every relationship look timeless and definitive. Timestamps, sources, and changing versions allow an agent to reason about what was believed at a particular point rather than treating the latest entry as the only truth.</em></p><h3>How do we prevent an agent from poisoning its own memory?</h3><p><em>Agent memory needs the same quality assurance we expect from any other piece of software. If a weak inference is stored, the system should retain its source and confidence rather than retrieving it later as an established fact. Independent agents can check one another, although using the same model everywhere may reproduce the same blind spots.</em></p><p><em>The more granular the memory, the more granular the testing needs to be. If you cannot inspect, validate, correct, or test an agent&#8217;s memory, you cannot determine whether it is reliable. Agentic work does not remove the need for software engineering; it makes careful testing at the component level even more important.</em></p><h3>Can the knowledge graph, GraphRAG platform and agent really evolve separately?</h3><p><em>They can evolve separately if their dependencies are explicit, versioned and tested. I give the major components of my systems their own semantic versions so each can be developed, released and, when necessary, rolled forward without rebuilding everything at once. Iteration creates infrastructure; big-bang development usually creates demonstrations.</em></p><p><em>A major graph-schema change may require the GraphRAG interface to change, just as a database change could break an application long before generative AI existed. But if the agent communicates through a stable natural-language contract, it may continue asking the same questions while the implementation underneath evolves. Separation does not mean independence; it means managing the connections deliberately.</em></p><h3>Are we measuring GraphRAG correctly?</h3><p><em>Counting relevant retrieved facts is not enough if the answer depends on a complete chain of evidence. A system may retrieve three of four necessary facts and still produce a confident but fundamentally incomplete conclusion. Teams need to test whether the system identified every question it needed to ask and recovered the evidence required to answer each one.</em></p><p><em>That also means testing the layers separately. Does the language model decompose the original request correctly? Does the GraphRAG endpoint return the right data for each subquestion? Does the agent assemble that evidence into a supported answer? When something is missing, we should identify the layer that lost it instead of labelling the entire system inaccurate.</em></p><h3>Is the AI hallucinating, or is the graph misleading it?</h3><p><em>I use a phrase in my training: similarity is not sameness. A similarity-based system may retrieve extra information that looks related but does not support the answer. If the model incorporates that noise as evidence, it can produce a hallucinated conclusion.</em></p><p><em>GraphRAG can fail differently. A query may miss a relationship, fail to traverse a node type, or return an event without its time or source. That is not necessarily the model inventing information; it may be a problem in the data, schema, query, or traversal. The only way to tell is to inspect the evidence path and determine which layer introduced the error.</em></p><div><hr></div><p>David&#8217;s case for GraphRAG is not that a graph should replace every RAG pipeline. It is almost the opposite: keep the parts that already work, introduce graphs where relationships genuinely matter, and make every new layer earn its place.</p><p>That restraint matters as agents become more capable of searching, remembering, and acting without constant direction. More intelligence in the workflow also creates more places for evidence to become incomplete, outdated, or distorted. Reliability, therefore, depends less on whether a system carries the latest label and more on whether its builders can see what it did, test each part, and recognise when it has done enough.</p><p>The most useful place to begin may not be a large architectural migration. It may simply be a small experiment, built slowly enough that you can still understand what the agent sees and why it reaches the conclusions it does.</p><div><hr></div><p style="text-align: center;"><em>Enjoyed today&#8217;s podcast and want more engineering conversations? Pull up a chair.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.agentengineering.co/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p>That&#8217;s it for this one. We&#8217;ll pick up the conversation next week. </p><p>Until then, keep building.</p><blockquote><p>Tanya D&#8217;cruz<br><em>Editor-in-Chief</em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[#6: From prompt to harness to loop]]></title><description><![CDATA[Mona Mona on the agentic skill ladder no one told you about]]></description><link>https://newsletter.agentengineering.co/p/from-prompt-to-harness-to-loop</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/from-prompt-to-harness-to-loop</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 16 Jul 2026 15:02:33 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/4abc87d5-a325-4449-810c-9f3a655fe373_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="callout-block" data-callout="true"><h4><code>Editor&#8217;s note</code></h4><p>Last week, <a href="https://www.linkedin.com/in/kenhuang8/">Ken Huang</a> broke down <a href="https://www.agentengineering.co/p/5-the-three-disciplines-of-agentic">agentic engineering into three disciplines: harness, context, and loops</a>. But I couldn&#8217;t leave that discussion just at that.</p><p>Because if you talk to builders, it&#8217;s clear that the work is still moving. A couple of years ago, it was all about prompting. Then we got better at tweaking context. But now, we&#8217;re starting to move beyond that. Toward loops.</p><p>So this week, Mona picks this apart from a practitioner&#8217;s lens. I&#8217;ll leave the floor to her.</p></div><p>In 2023, everyone wanted to be a prompt engineer. In 2026, the head of Claude Code at Anthropic says he doesn&#8217;t prompt Claude anymore. His words: &#8220;My job is to write loops.&#8221;</p><p>That single shift<span> </span>from prompting an agent to designing the loop that prompts it<span> </span>is the biggest change in how professionals work with AI since ChatGPT launched. And most people haven&#8217;t caught up yet.</p><p>Let me walk you up the ladder, one level at a time.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9XmK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9XmK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!9XmK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!9XmK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!9XmK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9XmK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1305993,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.agentengineering.co/i/205624786?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9XmK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!9XmK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!9XmK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!9XmK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56a14d15-a143-4e8d-9a66-e6552313a9f1_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><strong><span>Level 1</span></strong></h3><h4><strong><span>Prompt engineering, the words you send</span></strong></h4><p>Prompt engineering is where all of us started. You learn that &#8220;write me a blog post&#8221; gets mediocre output, but &#8220;write a 900-word blog post for enterprise cloud architects, in a direct second-person voice, with a contrarian hook&#8221; gets something usable.</p><p>Prompt engineering is about <em>how you express the task</em>. Be specific. Give examples. Assign a role. Ask for a format. But it operates on one thing: the message you type into the box.</p><p>The limitation is obvious. A prompt is a single instruction to a model that has no idea who you are, what your codebase looks like, or what happened five minutes ago. You can only cram so much into words.</p><h3><strong><span>Level 2</span></strong></h3><h4><strong><span>Context engineering, everything the model sees</span></strong></h4><p>The next realization: the model&#8217;s output quality depends less on your clever phrasing and more on <em>what information it has access to</em> when it responds.</p><p>Context engineering is about curating everything in the model&#8217;s window &#8212; the instructions, the documents, the data, the tool definitions, the conversation history, the intermediate results. It asks a different question from prompt engineering:</p><ul><li><p><strong>Prompt engineering asks:</strong> how do I say it?</p></li><li><p><strong>Context engineering asks:</strong> what should the model know?</p></li></ul><p>This is why retrieval (RAG), memory systems, and structured tool outputs became the center of gravity for enterprise AI teams in 2024&#8211;2025. Teams needed to feed the model the right context at the right time, and what was just as important was keeping the wrong context out.</p><p>But context engineering has ceilings of its own:</p><ul><li><p><strong>The window is finite, and attention degrades before it fills.</strong> </p><p>Even massive context windows don&#8217;t mean equally useful attention across every token. Overload the window, and the model starts missing details buried in the middle. More context is not better context.</p></li><li><p><strong>It&#8217;s a snapshot, not a conversation with reality.</strong> </p><p>Context is assembled before the model responds. Real tasks surprise you mid-execution &#8212; a dependency is missing, an API returns something malformed, a test fails. No amount of upfront curation can anticipate what only shows up when you run the thing.</p></li><li><p><strong>The model can know everything and still do nothing.</strong></p><p>Context engineering makes the model informed. It doesn&#8217;t make it capable. It can&#8217;t execute your code, check the result, or fix what broke. Knowledge without the ability to act is a very well-read consultant with no hands.</p></li><li><p><strong>Curation doesn&#8217;t scale by hand.</strong></p><p>Someone has to decide what goes into the window for every task. Past a certain complexity, assembling the context itself needs to be automated, which is the door the next two levels walk through.</p></li></ul><h3><strong><span>Level 3</span></strong></h3><h4><strong><span>Harness engineering, the environment the agent runs in</span></strong></h4><p>Then agents arrived, and a prompt plus context stopped being enough.</p><p>An agent doesn&#8217;t just answer &#8212; it <em>acts</em>. It reads files, runs code, calls APIs, and checks its own work. For that, it needs an environment: tools it&#8217;s allowed to use, a filesystem it can touch, permissions that constrain it, feedback mechanisms that tell it whether the tests passed.</p><p>That environment is the <strong>harness</strong>. Harness engineering is designing the executable world around a single agent run &#8212; the system prompt, the tool set, the memory, the guardrails, the verification steps. If context engineering decides what the model <em>knows</em>, harness engineering decides what the model <em>can do</em> and how it finds out whether it did it well.</p><div><hr></div><p style="text-align: center;"><em>If you haven&#8217;t subscribed yet, pull up a chair.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.agentengineering.co/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p>If you&#8217;ve used Claude Code, Cursor, or any serious coding agent, you&#8217;ve benefited from harness engineering, whether you knew the term or not. The difference between an agent that flails and an agent that ships is usually not the model. It&#8217;s the harness.</p><p>But even a great harness has a human bottleneck: <em>you</em>. You&#8217;re still the one deciding what the agent works on next, reading the output, and firing the next run. One turn after another. The agent is a power tool, and you&#8217;re holding it the whole time.</p><h3><strong><span>Level 4</span></strong></h3><h4><strong><span>Loop engineering, the cycle that drives it all</span></strong></h4><p>Loop engineering removes you from the turn-by-turn seat.</p><p>Instead of prompting the agent, you build a small system (sometimes a shell script, sometimes a scheduled job, sometimes a few hundred lines of orchestration code) that runs the agent in a repeating cycle:</p><div class="callout-block" data-callout="true"><p><span>Discover work &#8594; dispatch it to an agent &#8594; verify the result &#8594; persist the state &#8594; decide the next action &#8594; repeat</span></p></div><p>You define the goal and the stopping condition. The loop does the iterating. It runs on a schedule (including while you sleep) or until the goal is met.</p><p>The idea crystallized fast in mid-2026. Peter Steinberger argued that the real skill had moved from prompting agents to designing their loops; Addy Osmani gave the practice its name and structure in an essay the following day; and inside Anthropic, the </p><p>Claude Code team was already describing their daily work the same way. The framing that stuck: the harness equips a <em>single</em> agent run &#8212; the loop is what keeps running the harness, spawning helper agents, checking results, and feeding itself the next task.</p><p>Each level wraps the one below it. The loop runs the harness. The harness carries the context. The context frames the prompt. Nothing on the ladder becomes obsolete. Rather, the leverage keeps moving up.</p><h3><span>How to schedule a loop</span></h3><ol><li><p><span>Open your project in Claude Code.</span></p></li><li><p><span>Use the </span><code>/loop</code><span> command followed by your interval and the prompt or skill you want to run.</span></p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mwCj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mwCj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!mwCj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!mwCj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!mwCj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mwCj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1374110,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.agentengineering.co/i/205624786?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mwCj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!mwCj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!mwCj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!mwCj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d903b36-3570-441f-a120-8ee17a8b9078_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The </span><code>/loop</code><span> command is the core, with three modes depending on what you provide:</span></p><ul><li><p><code>/loop 5m check the deploy</code><span>: Runs your prompt on a fixed schedule (Claude converts the interval to a cron expression)</span></p></li><li><p><code>/loop check the deploy (no interval)</code><span>: Claude picks the delay dynamically after each iteration, between one minute and one hour, waiting less when things are active and more when quiet</span></p></li><li><p><code>Bare /loop</code><span>: Runs a built-in maintenance prompt that continues unfinished work, tends the current PR (review comments, failed CI, merge conflicts), and does cleanup passes. You can replace this default with your own loop.md file at project level (</span><code>.claude/loop.md</code><span>) or user level (</span><code>~/.claude/loop.md</code><span>).</span></p></li></ul><h3><strong><span>But here&#8217;s the part nobody puts in the headline</span></strong></h3><p><strong>First: loops multiply whatever you give them.</strong> A well-designed loop multiplies a strong engineer. A badly designed one multiplies a bad decision just as fast &#8212; with less of you watching. Two people can build the identical loop: one uses it to move faster on work they deeply understand, the other uses it to avoid understanding the work at all. The loop can&#8217;t tell the difference. You can.</p><p><strong>Second: the unglamorous work is making it stop.</strong> Ask practitioners about loops and the first war stories you&#8217;ll hear aren&#8217;t about architecture: they&#8217;re about cost. Autonomous loops that burned through hundreds of dollars overnight. Companies imposing hard monthly caps on agent spend after blowing annual AI budgets in a quarter. The production rule is simple: every loop ships with hard guards (budget limits, iteration caps, verification gates, audit logs) or it doesn&#8217;t ship.</p><p>Everything in this article assumes one foundational skill: knowing how to build an agent in the first place. The loop is only as good as the agents it orchestrates.</p><p>That&#8217;s what my latest book, <em><strong>AI Agents on AWS</strong></em> (co-authored with Bunny Kaushik, published by Packt), teaches. You&#8217;ll learn how to design agents that reason, use tools, and act autonomously; how to wire up the harness around them &#8212; memory, guardrails, verification; and how to take them from a notebook experiment to production on AWS. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.amazon.com/AI-Agents-AWS-Beginners-building/dp/1806387212/ref=sr_1_1?crid=2OZZEUH8W9N84&amp;dib=eyJ2IjoiMSJ9.WoYhBzPfst4oLow-UhWGf60ojRiINAh11owNzjGZRMCFE7CKTNrNFfXvR1eCp2Di_ZZe8p3DsOm_Y1I9gLdzt_Zv8WPOUGeu5tjYQfzyNf-GY8w7ZAnTIrdne1VkDjJEEIlGOghmTSg1UhR1kwcI8sOthiS8mhNtSf4gG1DWxJAoQQeA9RZTYmBV-CzdL60K47mUN3L3VBpQ6IZ7dcW6Xj6t-CM5lf_1i8Aoikx7kiI.WnBM7Emalv-D2FboWQEA13RX_6dnBcNlus_TvEwb5A8&amp;dib_tag=se&amp;keywords=AI+Agents+on+AWS&amp;qid=1784202103&amp;sprefix=%2Caps%2C326&amp;sr=8-1&quot;,&quot;text&quot;:&quot;Add it to your shelf&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.amazon.com/AI-Agents-AWS-Beginners-building/dp/1806387212/ref=sr_1_1?crid=2OZZEUH8W9N84&amp;dib=eyJ2IjoiMSJ9.WoYhBzPfst4oLow-UhWGf60ojRiINAh11owNzjGZRMCFE7CKTNrNFfXvR1eCp2Di_ZZe8p3DsOm_Y1I9gLdzt_Zv8WPOUGeu5tjYQfzyNf-GY8w7ZAnTIrdne1VkDjJEEIlGOghmTSg1UhR1kwcI8sOthiS8mhNtSf4gG1DWxJAoQQeA9RZTYmBV-CzdL60K47mUN3L3VBpQ6IZ7dcW6Xj6t-CM5lf_1i8Aoikx7kiI.WnBM7Emalv-D2FboWQEA13RX_6dnBcNlus_TvEwb5A8&amp;dib_tag=se&amp;keywords=AI+Agents+on+AWS&amp;qid=1784202103&amp;sprefix=%2Caps%2C326&amp;sr=8-1"><span>Add it to your shelf</span></a></p><div class="callout-block" data-callout="true"><p><a href="https://www.linkedin.com/in/mona-mona/">Mona Mona</a> is a Senior Worldwide GenAI Solutions Architect at AWS, where she works closely with enterprise teams to design, deploy, and scale AI systems in production. Her work spans model customization, evaluation, and inference, with a focus on how performance, cost, and reliability interact in real-world systems.</p></div><div><hr></div>]]></content:encoded></item><item><title><![CDATA[#5: The three disciplines of agentic engineering]]></title><description><![CDATA[Ken Huang, founder of Agentic AI and author of OpenClaw AI in Production, maps the three disciplines shaping the future of production AI systems]]></description><link>https://newsletter.agentengineering.co/p/5-the-three-disciplines-of-agentic</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/5-the-three-disciplines-of-agentic</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 09 Jul 2026 16:01:10 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/351f55c2-4aef-4abf-ab99-9b27087f1a0b_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most teams first encounter AI agents through a demo. A user sends a request, the model reasons through it, a tool is called, and a useful answer comes back. It feels almost magical. For a moment, it is tempting to believe the hard part is simply choosing the right model and writing a better prompt. Production quickly proves otherwise.</p><p>When an agent has to serve real users across real channels, with real credentials, tools, permissions, memory, failures, latency, and cost constraints, the problem changes. It stops being only a prompting problem. It becomes an engineering problem.</p><h3>This is where agentic engineering begins.</h3><p>Agentic engineering recognizes that an AI agent is not just a model wrapped in an interface. It is a runtime system. It needs boundaries, identity, observability, recovery paths, and a way to explain what happened when something goes wrong. It must be designed like infrastructure, not like a clever script.</p><p>Three practices define this shift: harness engineering, context engineering, and loop engineering.</p><p><strong>Harness engineering</strong> is the discipline of building the runtime around the agent. It includes gateways, tools, permissions, hooks, retries, telemetry, memory, sandboxes, and recovery paths. The harness is what lets an agent act safely in a real system. Without it, autonomy becomes fragility.</p><p><strong>Context engineering</strong> is the discipline of selecting, shaping, compressing, refreshing, and protecting the information given to the model. Context is no longer just &#8220;stuff in the prompt.&#8221; It is an operational resource with a budget, lifecycle, security posture, and failure mode.</p><div class="callout-block" data-callout="true"><p style="text-align: center;"><em><strong>New to context engineering?</strong><br>We recently dedicated an entire issue to the topic, with <a href="https://www.linkedin.com/in/denis-rothman/">Denis Rothman</a> exploring how engineered context helps transform probabilistic models into more reliable AI systems.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://open.substack.com/pub/agenticengg/p/4-the-illusion-of-autonomous-agents?r=8770aj&amp;utm_campaign=post-expanded-share&amp;utm_medium=web&quot;,&quot;text&quot;:&quot;Catch up here&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://open.substack.com/pub/agenticengg/p/4-the-illusion-of-autonomous-agents?r=8770aj&amp;utm_campaign=post-expanded-share&amp;utm_medium=web"><span>Catch up here</span></a></p></div><p><strong>Loop engineering</strong> is the discipline of designing the cycles through which agentic systems act, learn, improve, and produce software. At the runtime level, this means structuring the agent&#8217;s observe-decide-act-evaluate cycle with termination conditions, policy gates, retry limits, rollback paths, and human escalation points.</p><h3>But loop engineering goes further.</h3><p>One loop is the recursive self-improvement loop. A production agent system should be able to detect weakness, diagnose failure, propose improvement, test that improvement, validate it against policy, and only then adopt it. This does not mean giving agents uncontrolled permission to rewrite themselves. It means creating bounded, observable, auditable improvement cycles where agents can help harden the system without bypassing engineering discipline.</p><p>Another loop is the continuous software factory. In agentic engineering, software development itself becomes a loop: requirements become designs, designs become code, code becomes tests, tests become deployment signals, deployment produces telemetry, telemetry informs the next requirement, and the cycle continues. Agents can participate at every stage, but the factory still needs gates, evidence, review, and rollback.</p><p>This is one of the central themes of my upcoming book, <em><a href="https://www.amazon.com/OpenClaw-Production-Architecture-engineering-practices/dp/1807785017/ref=mp_s_a_1_2?crid=30MNI5UVODOVM&amp;dib=eyJ2IjoiMSJ9.e5VfTq1ruCPX9UYX9-OxYN_lA0EbzhPnYMlNbpcUN7V5I-ABElWrh4ymwNdMEhFA0RSosYaGK5O5cavyvnO5KPZMY_vNMRGo8I-78K46BxabJwcA02bOSsc9o_xcINTKcXCgFU9MdLzqJwLE6t3iHhYurSV5y3vH9qPezWJ8DpKiLNtZMHHIEGj2WPzBGJ9Ko_LC20Cqie8MrT6UqaREEQ.TOcY1fv6gMynz6RjXyckjdMgxek5ok7MPLT4UrnzTLA&amp;dib_tag=se&amp;keywords=openclaw+ai+in+production+ken+huang&amp;qid=1783347522&amp;sprefix=openclaw+production+%2Caps%2C353&amp;sr=8-2">OpenClaw AI in Production</a></em>. OpenClaw treats agent behavior as something that must be routed, bounded, observed, and corrected through architecture. The Gateway becomes the control plane. The request pipeline separates context assembly, authorization, directive handling, tool execution, and model invocation. Hooks allow cross-cutting concerns such as telemetry, policy, billing, and remediation to operate outside the core agent logic.</p><p>That architecture matters because agent failures rarely appear as simple crashes. A stale memory, a slow tool, a partial outage, a malformed directive, or an over-permissive policy can quietly distort the loop. The agent may keep acting, but each step moves it further from safe and useful behavior.</p><h3>Production systems need to detect that drift.</h3><p>This is why observability for agents has to go beyond CPU, latency, and HTTP response codes. A 200 OK from a model provider does not tell us whether the agent used the right tool, respected the right policy, retrieved the right memory, or stopped at the right time. Agentic systems need semantic observability: traces that show the agent&#8217;s observable decisions without leaking sensitive data or private reasoning.</p><p>The deeper opportunity is self-correction. If an agent platform can detect degraded state, isolate the failing component, reduce privileges, pause risky workflows, retry safely, or route through a fallback path, then we move from passive monitoring to active resilience. If the software factory can learn from its own incidents, tests, deployments, and user feedback, then engineering itself becomes more adaptive.</p><p>The future will not be won by prompts alone. It will be built by engineers who understand harnesses, contexts, and loops - and who can turn model intelligence into systems that are secure, observable, resilient, and continuously improving.</p><p>OpenClaw AI in Production is my attempt to map that territory in detail: from Gateway-centered architecture and policy enforcement to distributed state, semantic observability, self-correcting stacks, fault injection, high-throughput design, and decentralized deployments.</p><p>Agentic engineering as a discipline is still young. But one thing is already clear: building agents that work in demos is very different from building agents that survive, learn, and improve in production.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5OEJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5OEJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!5OEJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!5OEJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!5OEJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5OEJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1875025,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.agentengineering.co/i/205575777?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5OEJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!5OEJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!5OEJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!5OEJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766564a3-3eb4-4118-830c-cbd542ac2bf9_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="callout-block" data-callout="true"><p><strong><a href="https://www.linkedin.com/in/kenhuang8/">Ken Huang</a></strong> is an AI researcher, author, educator, and founder of <a href="https://kenhuangus.substack.com/">Agentic AI</a> and DistributedApps.ai. He serves as CEO and Chief AI Officer at DistributedApps.ai, is an Adjunct Professor at the University of San Francisco, and co-chairs multiple AI safety initiatives at the Cloud Security Alliance and OWASP. A prolific author of books on AI, security, and distributed systems, his latest work, <em><a href="https://www.amazon.com/OpenClaw-Production-Architecture-engineering-practices/dp/1807785017/ref=mp_s_a_1_2?crid=30MNI5UVODOVM&amp;dib=eyJ2IjoiMSJ9.e5VfTq1ruCPX9UYX9-OxYN_lA0EbzhPnYMlNbpcUN7V5I-ABElWrh4ymwNdMEhFA0RSosYaGK5O5cavyvnO5KPZMY_vNMRGo8I-78K46BxabJwcA02bOSsc9o_xcINTKcXCgFU9MdLzqJwLE6t3iHhYurSV5y3vH9qPezWJ8DpKiLNtZMHHIEGj2WPzBGJ9Ko_LC20Cqie8MrT6UqaREEQ.TOcY1fv6gMynz6RjXyckjdMgxek5ok7MPLT4UrnzTLA&amp;dib_tag=se&amp;keywords=openclaw+ai+in+production+ken+huang&amp;qid=1783347522&amp;sprefix=openclaw+production+%2Caps%2C353&amp;sr=8-2">OpenClaw AI in Production</a></em>, explores the engineering practices required to build reliable AI agents at scale.</p></div><div><hr></div><p style="text-align: center;">This article is a glimpse into the ideas behind Ken&#8217;s <em>OpenClaw AI in Production</em>. If you&#8217;re interested in building AI systems that are resilient by design rather than optimistic by default, the book expands on the engineering principles that separate production platforms from impressive prototypes.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.amazon.com/OpenClaw-Production-Architecture-engineering-practices/dp/1807785017/ref=mp_s_a_1_2?crid=30MNI5UVODOVM&amp;dib=eyJ2IjoiMSJ9.e5VfTq1ruCPX9UYX9-OxYN_lA0EbzhPnYMlNbpcUN7V5I-ABElWrh4ymwNdMEhFA0RSosYaGK5O5cavyvnO5KPZMY_vNMRGo8I-78K46BxabJwcA02bOSsc9o_xcINTKcXCgFU9MdLzqJwLE6t3iHhYurSV5y3vH9qPezWJ8DpKiLNtZMHHIEGj2WPzBGJ9Ko_LC20Cqie8MrT6UqaREEQ.TOcY1fv6gMynz6RjXyckjdMgxek5ok7MPLT4UrnzTLA&amp;dib_tag=se&amp;keywords=openclaw+ai+in+production+ken+huang&amp;qid=1783347522&amp;sprefix=openclaw+production+%2Caps%2C353&amp;sr=8-2&quot;,&quot;text&quot;:&quot;Add it to your shelf&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.amazon.com/OpenClaw-Production-Architecture-engineering-practices/dp/1807785017/ref=mp_s_a_1_2?crid=30MNI5UVODOVM&amp;dib=eyJ2IjoiMSJ9.e5VfTq1ruCPX9UYX9-OxYN_lA0EbzhPnYMlNbpcUN7V5I-ABElWrh4ymwNdMEhFA0RSosYaGK5O5cavyvnO5KPZMY_vNMRGo8I-78K46BxabJwcA02bOSsc9o_xcINTKcXCgFU9MdLzqJwLE6t3iHhYurSV5y3vH9qPezWJ8DpKiLNtZMHHIEGj2WPzBGJ9Ko_LC20Cqie8MrT6UqaREEQ.TOcY1fv6gMynz6RjXyckjdMgxek5ok7MPLT4UrnzTLA&amp;dib_tag=se&amp;keywords=openclaw+ai+in+production+ken+huang&amp;qid=1783347522&amp;sprefix=openclaw+production+%2Caps%2C353&amp;sr=8-2"><span>Add it to your shelf</span></a></p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[#4: The illusion of autonomous agents (part 2)]]></title><description><![CDATA[Denis Rothman on context engineering as the missing layer for reliable AI agents]]></description><link>https://newsletter.agentengineering.co/p/4-the-illusion-of-autonomous-agents</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/4-the-illusion-of-autonomous-agents</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 02 Jul 2026 15:39:25 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a7164517-9c64-4c06-b35c-9d293c5d4f3c_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Last week, we explored why autonomous agents struggle to deliver the reliability enterprises expect. The problem isn&#8217;t simply that LLMs make mistakes, but that prompt engineering alone cannot overcome the structural limits of probabilistic reasoning. If reliability cannot emerge from prompts, then it must be engineered elsewhere. That &#8220;elsewhere&#8221; is context: not as a longer system prompt or larger context window, but as a structured, transparent architecture that guides how agents reason, communicate, and act. This is where context engineering begins.</p><p><span>Context complexity exists across five distinct levels, evolving from zero-context basic prompts to highly advanced semantic blueprints. Prompt engineering resides at the shallowest level, treating context as a mere text prefix. Context engineering, conversely, treats context as a structural, programmatic architecture.</span></p><p><span>To build a robust multi-agent system (MAS), we must transition from linear text parsing to multidimensional semantic structures. We can achieve this by implementing Semantic Role Labeling (SRL) to map complex data relationships natively. By utilizing the Model Context Protocol (MCP), we can define rigorous protocol message formats that specialist agents such as a Researcher, Writer, and Orchestrator use to communicate without ambiguity.</span></p><p><span>Instead of hoping a black-box model infers the correct workflow, we architect a semantic blueprint that explicitly guides the system&#8217;s reasoning process, completely decoupling the immutable enterprise data layer from the probabilistic reasoning layer.</span></p><h3><span>Architecting the glass-box context engine</span></h3><p><span>The antidote to black-box unpredictability is the context engine: a transparent, glass-box architecture of context and reasoning. In this framework, the engine&#8217;s core intelligence is compartmentalized into distinct, observable modules: the Planner, the Executor, and the Tracer.</span></p><ul><li><p><strong><span>The Planner:</span></strong><span> Acts as a meta-controller. It utilizes a discoverable Agent Registry to dynamically route tasks, intelligently switching between semantic search and strict data filtering based on the precise context of the request.</span></p></li><li><p><strong><span>The Executor:</span></strong><span> Relies on a Dual RAG architecture, simultaneously processing factual data via a strict Knowledge Base and procedural instructions via a Context Library.</span></p></li><li><p><strong><span>The Tracer:</span></strong><span> Allows human operators to parse the ExecutionTrace object to render token metrics, dependency resolutions, and view the internal &#8220;thinking&#8221; steps of the agent.</span></p></li></ul><p><span>Within a finite reaction field, agent decisions must be bounded by plausibility to avoid cognitive dissonance. We enforce these boundaries using operators like the Semantic Switch ($\sigma$) and the Triton Planner Kernel. These tools apply explicit, rule-bound kinematics over the continuous probabilities of the neural field, acting as the constraint satisfaction engines that pure LLMs lack.</span></p><h3><span>Enterprise guardrails: Scaling with Dual RAG</span></h3><p><span>Bringing these systems into production requires </span><strong><span>moving AI directly to the data.</span></strong><span> By leveraging architectures of modern enterprise databases with vector and relational support, we can build Sovereign AI systems that are dock-agnostic, portable multi-agent systems directly to immutable enterprise databases.</span></p><p><span>This paradigm shift enables hyper-contextual capabilities:</span></p><ul><li><p><strong><span>Hybrid querying:</span></strong><span> Combining strict SQL scalar filters (such as experience levels or salary caps) with semantic vector searches to generate grounded, policy-compliant recommendations.</span></p></li><li><p><strong><span>Converged Spatial-RAG and GraphRAG:</span></strong><span> Integrating physical geolocation (Oracle Spatial) and social dimensions (SQL Property Graphs) directly into the vector search pipeline to evaluate meaning, physical proximity, and relationship mappings simultaneously.</span></p></li><li><p><strong><span>Micro-context engineering:</span></strong><span> Deploying specific Summarizer agents to proactively manage API costs, reduce context overhead, and actively manage token limits.</span></p></li><li><p><strong><span>Policy-driven moderation:</span></strong><span> Implementing a two-stage moderation gatekeeper that flags anomalies, prevents data poisoning, and adapts to real-world legal compliance limits.</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JJES!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JJES!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!JJES!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!JJES!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!JJES!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JJES!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2455914,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.agentengineering.co/i/204686086?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JJES!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!JJES!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!JJES!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!JJES!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32481c16-f31c-43a2-9e4c-10398103133d_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><span>Conclusion: From cost center to value multiplier</span></h3><p><span>The debate over how to achieve true agentic reasoning will not be settled by building ever-larger black boxes or writing increasingly convoluted text prompts. The future belongs to those who recognize and engineer around the structural boundaries of artificial intelligence.</span></p><p><span>By embracing context engineering, we shift AI from an unpredictable stochastic experiment into a verifiable, transparent asset. Through glass-box architectures, Dual RAG pipelines, and the rigorous application of spatial and semantic boundaries, we can finally build human-centered AI systems that are genuinely business-ready. It is time to stop prompting the machine and start engineering the context.</span></p><div class="callout-block" data-callout="true"><p><strong><a href="https://www.linkedin.com/in/denis-rothman/"><span>Denis Rothman</span></a></strong><span> graduated from Sorbonne University and Paris-Diderot University, designing one of the very first word2matrix patented embedding and patented AI conversational agents. He began his career authoring one of the first AI cognitive NLP chatbots applied as an automated language teacher for Moet et Chandon and other companies. He authored an AI resource optimizer for IBM and apparel producers. He then authored an Advanced Planning and Scheduling (APS) solution used worldwide.</span></p></div><div><hr></div><p style="text-align: center;"><em>Less hype, more engineering. Pull up a chair.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.agentengineering.co/subscribe?"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[OpenAI Codex Bootcamp]]></title><description><![CDATA[Bootcamp &#183; Certificate Included &#183; Hands-on Labs &#183; AI Coding]]></description><link>https://newsletter.agentengineering.co/p/openai-codex-bootcamp</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/openai-codex-bootcamp</guid><dc:creator><![CDATA[Kunal Sawant]]></dc:creator><pubDate>Thu, 02 Jul 2026 05:55:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!pWyo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pWyo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pWyo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!pWyo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!pWyo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!pWyo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pWyo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png" width="728" height="382.2" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:173410,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.agentengineering.co/i/204577926?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pWyo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!pWyo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!pWyo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!pWyo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F712c02d8-fe82-477e-8871-9019a970a4ca_1200x630.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>AI-assisted development is becoming the new standard. Learn how to use OpenAI Codex effectively in real engineering projects through a live, instructor-led bootcamp.</p><p><strong>Date:</strong> Saturday, July 25, 2026<br><strong>Time:</strong> 9:00 AM &#8211; 3:00 PM EST<br><strong>Format:</strong> Live Virtual Bootcamp<br><strong>Hosted by:</strong> Packt Publishing</p><p><strong>What you&#8217;ll learn:</strong></p><ul><li><p>How to use OpenAI Codex as a practical coding partner</p></li><li><p>How to plan, build, debug, test, and review code with AI assistance</p></li><li><p>How to apply Codex across real-world developer workflows</p></li><li><p>How to complete hands-on projects and coding exercises</p></li><li><p>How to improve code quality, productivity, and delivery speed using AI-assisted development</p></li></ul><p><strong>How this helps your career:</strong><br>AI coding tools are becoming part of everyday software development. Learning how to work effectively with Codex can help developers become faster, more confident, and more valuable in AI-enabled engineering teams. This bootcamp helps participants strengthen their practical development workflow, improve productivity, and build skills that are increasingly relevant across modern software roles.</p><p><strong>Certificate included:</strong><br>Participants will receive a Packt certificate, which can be added to their LinkedIn profile, resume, or professional portfolio to showcase their learning and commitment to practical AI-assisted development.</p><p><strong>Best suited for:</strong><br>Technically professionals, developers, software engineers, technical leads, and AI builders who want to use OpenAI Codex to simplify coding workflows and improve engineering productivity.</p><p>This bootcamp helps you stay ahead so you can ship more, grow faster, and take on higher-impact work.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/openai-codex-bootcamp-tickets-1992048666185?aff=substack&quot;,&quot;text&quot;:&quot;Register Here&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.eventbrite.co.uk/e/openai-codex-bootcamp-tickets-1992048666185?aff=substack"><span>Register Here</span></a></p>]]></content:encoded></item><item><title><![CDATA[SUSE refuses to measure its engineers by how much code their agents write]]></title><description><![CDATA[Rick Spencer on why output, tokens, and lines of code tell you nothing, and what an open-source enterprise tracks instead]]></description><link>https://newsletter.agentengineering.co/p/suse-refuses-to-measure-its-engineers</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/suse-refuses-to-measure-its-engineers</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Tue, 30 Jun 2026 10:25:24 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/241f5215-5a9f-4825-a958-05d13383f442_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>As AI agents move into engineering workflows, new leaderboard metrics are tracking lines of code submitted, tokens consumed, and per-developer utilization. If agents are generating output, then output should be measured, compared across engineers, and ranked.</span></p><p><a href="https://www.linkedin.com/in/rickspencer3"><span>Rick Spencer</span></a><span>, General Manager for Technology and Product at </span><a href="https://www.suse.com/"><span>SUSE</span></a><span>, has looked hard at how the industry is measuring AI&#8217;s effect on engineering. &#8220;I consider that garbage vanity metrics,&#8221; he says, calling them unhelpful.</span></p><p><span>His argument for what to track instead is one of the more clarifying things an engineering leader can hear right now, because it separates the numbers that look like progress from the numbers that actually represent it.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Bhqf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Bhqf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Bhqf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Bhqf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Bhqf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Bhqf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2131552,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.agentengineering.co/i/204235791?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Bhqf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Bhqf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Bhqf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Bhqf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4dec581-739a-422f-8764-96c3baeee56f_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><span>Output is cheap; impact is what counts</span></h3><p><span>The core of Spencer&#8217;s position is a distinction between output and impact, and it matters because the two come apart precisely when AI enters the picture. AI makes output cheap as the lines of code, pull requests, and token counts all mount up when agents are doing the writing, which makes them exactly the wrong thing to measure if what you care about is value delivered. &#8220;We&#8217;re really tending away from measurements that measure output and utilization, and we&#8217;re trying to focus on impact,&#8221; he says. A leaderboard that ranks engineers by how much their agents produced does not tell you who is solving the hardest problems or keeping customers safe. It tells you who is generating the most volume, and in an AI-assisted world, that number is close to meaningless.</span></p><p><span>There is also a structural reason the standard tooling does not fit SUSE, and it applies to more organizations than it might first appear. Much of the available measurement tooling assumes a particular shape of company. &#8220;They really assume you&#8217;re a proprietary software company where everyone&#8217;s working on a single code base,&#8221; Spencer explains, &#8220;which is just not how an open-source enterprise works.&#8221; His engineers work across hundreds, sometimes thousands, of repositories, where the maintenance work on each one differs enormously. A per-developer comparison across that landscape measures the shape of the work far more than it measures the contribution of the engineer, which is why he treats developer-to-developer comparison as fundamentally low value rather than merely imperfect.</span></p><p><span>The reporting burden itself is part of his objection, and it is a point leaders setting up AI dashboards should sit with carefully. A measurement regime that requires engineers to generate weekly utilization reports spends the very time it claims to be optimizing. &#8220;I&#8217;d rather have them working than reporting,&#8221; Spencer says. The instrument meant to measure productivity eventually becomes a tax on it.</span></p><h3><span>What SUSE tracks instead</span></h3><p><span>Rejecting vanity metrics only helps if there is something better to put in their place. And Spencer shares how SUSE measures business impact in terms that connect directly to what customers actually receive. &#8220;How fast are CVEs being addressed, how fast are patches being backported, how fast are our L3 responses getting closed while maintaining the same NPS score,&#8221; he underscores, listing what his teams track. The common thread is that each one is an outcome the customer feels, not an activity the engineer performs. AI has been applied to exactly these areas, so measuring the speed and quality of those outcomes tells you whether the AI is doing anything worth its cost, which is the actual question worth asking.</span></p><p><span>This shift from output to outcome reframes what a metric is for in the first place. A CVE response time captures whether the organization is keeping customers safe faster than it used to. A backport speed captures whether stable releases are getting their fixes without the manual grind that used to gate them. These numbers move because the underlying work got genuinely better, not because more text was generated, and that is the property that makes them trustworthy. They are also far harder to game, because the only way to improve them is to actually improve the thing the customer depends on.</span></p><h3><span>Give managers visibility, not a leaderboard</span></h3><p><span>None of this means SUSE ignores cost or utilization entirely, and the distinction Spencer draws here is the one that keeps the approach from collapsing into either negligence or surveillance. The company is building dashboards that give engineering managers visibility into their team&#8217;s cost and utilization, but the purpose is coaching rather than ranking. The unit of analysis is the team, and the question it answers is diagnostic. Spencer gives the example of a manager with an eight-person team noticing the numbers and asking the right kind of question. &#8220;We&#8217;re burning a lot of tokens. What are we actually doing that&#8217;s burning that many tokens? I&#8217;m not sure we&#8217;re getting value out of that.&#8221; The inverse matters just as much, where purchased seats for a code assistant sit unused, and the manager asks whether there are places the team should be drawing value that it is currently leaving on the table.</span></p><p><span>The governance side of that picture, including how SUSE keeps agents and their costs inside a boundary it can stand behind, is covered in a companion piece,</span><a href="https://deepengineering.net/p/how-suse-runs-ai-without-losing-control"><span> How SUSE Runs AI Without Losing Control</span></a><span>.</span></p><p><span>The difference between this and a leaderboard is not subtle, and it is the heart of the leadership lesson. A leaderboard exposes individuals and turns measurement into a game engineers play against each other, a game Spencer is explicit has nothing to do with customer value. Team-level cost visibility used for coaching does the opposite. It gives a manager the information to guide the team toward better use of the tools without making any individual engineer feel watched. &#8220;We&#8217;re really trying to decentralize and allow engineering managers to guide their teams on getting the most value out of the AI,&#8221; he says, &#8220;without it becoming like a leaderboard game where developers feel like they&#8217;re exposed.&#8221; The data exists to help the manager help the team, not to rank the team against itself.</span></p><p><strong><span>The principle holding it together</span></strong></p><p><span>What makes Spencer&#8217;s approach more than a list of preferred numbers is the principle holding it together, which is that measurement should serve the work rather than distort it. Every choice he describes follows from that one idea. Impact comes before output because output is the thing AI inflates. Team-level diagnostics come before individual leaderboards, because the goal is coaching rather than competition. Business outcomes come before activity counts, because outcomes are what customers actually receive. The decentralization to engineering managers reflects the same conviction that the people closest to the work are best placed to judge whether the AI is helping, given the right information and trusted to use it well.</span></p><p><span>The deeper point for any leader standing up AI measurement is that the easy numbers and the useful numbers are not the same, and AI has widened the gap between them. The figures that are simplest to collect, lines of code, tokens, and per-head utilization, are the ones AI has made least meaningful. The figures that matter, the speed and quality of the outcomes customers depend on, take more thought to define and more care to track. Spencer&#8217;s argument is that the effort is the job. &#8220;Let&#8217;s focus on the impact,&#8221; he says, &#8220;the business impact, not on the utilization.&#8221; For engineering leaders deciding what belongs on a dashboard as agents reshape their teams, that is the distinction worth getting right before the vanity metrics calcify into the way the organization sees itself.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><em>Less hype, more engineering. Pull up a chair.</em>  </p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[#3: The illusion of autonomous agents (part 1)]]></title><description><![CDATA[Denis Rothman, author of Context Engineering for Multi-Agent Systems and RAG-Driven Generative AI, begins this two-part series by examining why today&#8217;s autonomous agents struggle to deliver reliable, enterprise-grade performance.]]></description><link>https://newsletter.agentengineering.co/p/the-illusion-of-autonomous-agents</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/the-illusion-of-autonomous-agents</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 25 Jun 2026 17:27:11 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/27026882-610c-478f-971b-e8ed13d4a6fc_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>The AI community is caught in a contradiction. On one hand, there is a push to deploy autonomous agents into enterprise environments, expecting them to reason, plan, and execute complex, multi-step workflows. On the other hand, the methodology relied upon to build these agents is overwhelmingly based on prompt engineering, which is merely an attempt to cajole predictable, deterministic behavior out of fundamentally stochastic, black-box LLMs.</span></p><p><span>This creates a pervasive dissonance: the expectation of industrial reliability built atop probabilistic generation. </span></p><p><span>The current noise suggests that if we simply scale the parameters, refine the system prompts, or throw more compute at the problem, true autonomous reasoning will spontaneously emerge.</span></p><p><span>To cut through this noise, we must confront an uncomfortable reality. Unconstrained probabilistic generation cannot serve as the kinematics for reliable robotic or enterprise execution. If we are to build truly agentic systems, we must move beyond the brittle, zero-context art of prompting and embrace the rigorous, transparent discipline of </span><strong><span>context engineering</span></strong><span>.</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Agentic Engineering! Subscribe to receive new posts every week :)</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3><span>The limits of the translation lattice</span></h3><p><span>At their core, LLMs do not guarantee structured reasoning across steps. They act as translation lattices operating on lossy text compressions, fundamentally disconnected from the non-verbal, experiential grounding that characterizes human thought. They are magnificent at semantic association but possess no innate understanding of physical or logical bounds. When an agent relies solely on an LLM&#8217;s internal weights to reason, it runs headfirst into the reality of the </span><strong><span>Probability Decay Theorem (PDT)</span></strong><span>, a multiplicative reliability collapse effect. The theorem was derived through multiple real-world implementations by the author.</span></p><p><span>The PDT governs the multiplicative collapse of chained rules. In single-turn interactions, an LLM might succeed 95% of the time. But in agentic workflows requiring multi-step execution, the probability of success does not simply carry an additive exception load PDT</span><sub><span>1</span></sub><span>. Instead, it collapses under multiplicative chained probability PDT</span><sub><span>2</span></sub><span>. A chain of ten probabilistic bets, each with a 95% success rate, yields an overall reliability that is entirely unacceptable for real-world enterprise operations.</span></p><p><span>If an autonomous agent executes a sequence of tasks where each step is 95% reliable, the multiplicative collapse of the Probability Decay Theorem dictates that it takes just 14 consecutive steps 0.95</span><sup><span>14</span></sup><span> </span>&#8776; <span>0.488 for the overall success rate to plummet below 50%, rendering the entire workflow less predictable than a blind coin toss. This shows that after enough chained steps, reliability rapidly degrades below usable thresholds, even when individual steps appear highly accurate.</span></p><p><span>LLMs operate on statistical associations in text without explicit grounding in physical or logical constraints. This means that mapping raw signals to lexical labels will always incur an irreducible loss due to the polysemy inherent in language. These models are structurally bounded by the imprecision of polysemy. Treating their hallucinations as mere software bugs to be patched by longer text prompts ignores the mathematical reality of the medium.</span></p><h3><span>The embodied reality factor</span></h3><p><span>To understand how to fix this, we must look at the lineage of modern embodied AI, tracing back to industrial automated systems. Physical robotics and real-world supply chains cannot operate on unconstrained probabilistic generation. Instead, algorithmic logic must be ruthlessly translated into the thermodynamic and spatial kinematics of the physical world. This absolute reliance on physical bounds forms the applied foundation for defining hard limitations such as spatial collisions, reaction times, and hardware constraints, which are not errors. They are the foundational constraint forces that dictate the accuracy of a signal. This principle inherently solves the simulation-to-reality (Sim2Real) transfer gap. If we want our digital agents to interact reliably with the real world, we must enforce digital representations of physical limits to ensure their reasoning remains strictly bounded by the reality factor of the environment.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!g69E!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!g69E!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!g69E!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!g69E!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!g69E!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!g69E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2455914,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.agentengineering.co/i/203580946?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!g69E!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!g69E!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!g69E!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!g69E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01b93bc4-55bc-4579-92a2-04ab82c01341_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>If reliability cannot emerge from probabilistic generation alone, then it has to come from somewhere else. That </span><em><span>somewhere else</span></em><span> isn&#8217;t another prompt. It&#8217;s context, engineered as a first-class component of the system.</span></p><p><span>That&#8217;s where we&#8217;ll begin in Part 2. See you next Thursday!</span></p><div class="callout-block" data-callout="true"><p><strong><a href="https://www.linkedin.com/in/denis-rothman/"><span>Denis Rothman</span></a></strong><span> graduated from Sorbonne University and Paris-Diderot University, designing one of the very first word2matrix patented embedding and patented AI conversational agents. He began his career authoring one of the first AI cognitive NLP chatbots applied as an automated language teacher for Moet et Chandon and other companies. He authored an AI resource optimizer for IBM and apparel producers. He then authored an Advanced Planning and Scheduling (APS) solution used worldwide.</span></p></div>]]></content:encoded></item><item><title><![CDATA[#2: Why a good answer doesn’t mean a good agent]]></title><description><![CDATA[Ammar Mohanna argues that most teams are measuring outcomes when they should be measuring behavior across the entire agent workflow]]></description><link>https://newsletter.agentengineering.co/p/agentic-engineering-2-why-a-good</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/agentic-engineering-2-why-a-good</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 18 Jun 2026 13:57:12 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e66e55bd-024d-4cef-8520-302fa93748fd_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>During a time when AI conversations are often louder than they are useful, <strong><a href="https://www.linkedin.com/in/ammarmohanna/">Ammar Mohanna, PhD</a></strong>, brings a refreshing perspective.</p><p>His career has moved fluidly between academia and industry, from teaching advanced AI courses at the American University of Beirut to advising teams on turning machine learning ideas into systems that can be trusted.</p><p>He is also known for his candid take on the current AI landscape, especially the gap between meaningful engineering and what he often calls <em>AI slop</em>.</p><p>In this conversation, Ammar challenges one of the most common assumptions in agent development: that a correct answer is evidence of a successful agent. He explains why reliability lies in the path an agent takes, not just in the result it produces, and why evaluation must evolve from output scoring to a discipline that measures behaviour and trustworthiness in production.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7gmb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7gmb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!7gmb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!7gmb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!7gmb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7gmb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1729113,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.agentengineering.co/i/202568466?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7gmb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!7gmb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!7gmb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!7gmb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420d5d51-10ba-4893-a7dc-be8b9414a091_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="pullquote"><p><strong><span data-color="#ff0000" style="color: rgb(255, 0, 0);">Read through to the end. We&#8217;ve left something extra for Agentic Engineering readers.</span></strong></p></div><h3><strong>Most teams think they&#8217;re evaluating agents, but they&#8217;re actually not. Where do you see the biggest illusion of evaluation today?</strong></h3><p><em>The biggest illusion is that teams think they are evaluating an agent when they are only evaluating the final answer.</em></p><p><em>That works well for a chatbot. But an agent is different. It plans, chooses tools, passes arguments, reads observations, retries, stops, and sometimes takes action. A final-answer score hides most of the actual failure surface.</em></p><p><em>An agent can produce a good-looking answer after calling the wrong tool, wasting ten steps, misreading a tool result, or ignoring a failed call. From the outside, the answer may look acceptable. From a reliability perspective, the run is not acceptable.</em></p><p><em>So the illusion is: &#8220;the answer looked right, therefore the agent worked.&#8221; However, what you need to know is whether the path was valid, efficient, grounded, and safe.</em></p><div><hr></div><h3><strong>You break evaluation into component, trajectory, outcome, and adversarial layers. Where do most teams underinvest, and what failures does that lead to?</strong></h3><p><em>Most teams underinvest in trajectory evaluation and adversarial evaluation.</em></p><p><em>Outcome evaluation is the easiest layer to reach for because the final answer is visible. Component evaluation is also fairly intuitive once tools are involved: did it choose the right tool, did it pass valid arguments, did the plan make sense?</em></p><p><em>Trajectory evaluation is harder because you need structured traces and assertions over the run itself. But this is where many production failures live: loops, duplicate calls, silent retries, no recovery after a tool failure, unnecessary detours, high latency, high token cost. Two agents can produce the same answer, but one gets there in four clean steps, and the other gets there through an expensive, brittle path. Output scoring treats them as equal. Production does not.</em></p><p><em>Adversarial evaluation is also underbuilt. Teams may try a few prompt-injection examples manually, but they rarely turn those attacks into a versioned regression suite. That leads to a false sense of safety. A guardrail can look very strong against the examples it was designed for, while still being fragile against slightly different payloads.</em></p><div><hr></div><h3><strong>Agent failures only show up after deployment. What&#8217;s the hardest failure mode to catch early, even with a good evaluation setup?</strong></h3><p><em>The hardest failures are the ones that look like successful runs.</em></p><p><em>A tool returns something plausible but incomplete. The agent takes a reasonable-looking path. The final answer is fluent. No exception is thrown. But the answer is weakly grounded, the evidence is stale, or the agent skipped a recovery step after a bad observation.</em></p><p><em>These failures are hard because they do not announce themselves as failures. They show up later as retries, edits after the answer, escalations, user abandonment, or quiet loss of trust.</em></p><p><em>The other hard category is drift. A hosted model changes, a tool schema changes, retrieval content shifts, or user traffic moves into a different distribution. Nothing &#8220;breaks&#8221; in the traditional software sense, but the agent becomes less reliable. This is why offline evals need to connect to production monitoring. A test suite is necessary, but it is not the whole system.</em></p><div><hr></div><h3><strong>LLM-as-a-judge is becoming a default pattern. Where does it actually work well, and where does it quietly break?</strong></h3><p><em>LLM-as-a-judge works well when the task is bounded, the rubric is explicit, and the judge has the evidence they need. It is useful for rubric-based scoring, regression checks, pairwise comparisons, and multi-dimensional outcome evaluation, especially when you calibrate it against human labels.</em></p><p><em>The important part is that the judge itself has to be evaluated. I would not trust a judge just because it is an LLM. I would look at correlation with human labels, agreement rates, mean absolute error, and performance by rubric dimension.</em></p><p><em>Where it quietly breaks is when teams use it as an uncalibrated oracle. Judges often reward verbosity, prefer answers in a certain style, miss missing citations, or give a strong score to an answer that is polished but not grounded. Overall scores can also hide weak dimensions. For example, a judge may be decent on safety or format, but poor on groundedness, which is often the dimension that matters most for a research or retrieval-heavy agent.</em></p><p><em>So I see LLM judges as useful evaluators, not authorities. They need rubrics, evidence, calibration, and periodic human audit.</em></p><div><hr></div><h3><strong>If you had to audit an agent system in production with very limited time, what signals or metrics would you look at first to decide if it&#8217;s reliable?</strong></h3><p><em>Aggregate success rate is often the last metric I look at. I would start with the traces behind the failures and the production signals that users generate when the agent is not working.</em></p><p><em>The first signals I would inspect are abandonment rate, retry rate, escalation rate, clarification rate, thumbs down, and edit-after-answer rate. Those are often more honest than a dashboard success metric.</em></p><p><em>Then I would look at trace-level reliability: number of tool calls, duplicate calls, loop-like behaviour, failed tool calls, recovery after failure, latency, and token cost. A reliable agent should not only get the answer right; it should get there through a path that is stable and explainable.</em></p><p><em>I would also check whether offline evals are tied to production: are failed production examples converted into regression tests? Are adversarial cases versioned? Are judge scores calibrated against human labels? Are there no-go gates for safety, groundedness, cost, latency, and step count?</em></p><p><em>With limited time, I am looking for one thing: whether the team has connected offline evaluation, online monitoring, and regression gates. If those are disconnected, reliability is usually more assumed than measured.</em></p><div><hr></div><p>As organizations continue to explore what AI can and should do, voices like <a href="https://www.linkedin.com/in/ammarmohanna/">Ammar&#8217;s</a> help bring the discussion back to the questions that matter: What problem are we really solving? Can the system be trusted? And are we building something meaningful, or simply adding more noise to an already crowded field?</p><p>If you&#8217;d like to continue exploring these ideas, Ammar will be speaking at <strong><a href="https://packt.link/JuExC">Agent Evals Bootcamp on June 27th</a></strong>, where he will turn these ideas into a hands-on framework for evaluating agents across tool use, planning, trajectories, outcomes, regressions, and adversarial failure modes before deployment. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2DCz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F265c1c36-9516-4457-86be-2b4fde4cd9ce_2160x897.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2DCz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F265c1c36-9516-4457-86be-2b4fde4cd9ce_2160x897.png 424w, https://substackcdn.com/image/fetch/$s_!2DCz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F265c1c36-9516-4457-86be-2b4fde4cd9ce_2160x897.png 848w, https://substackcdn.com/image/fetch/$s_!2DCz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F265c1c36-9516-4457-86be-2b4fde4cd9ce_2160x897.png 1272w, https://substackcdn.com/image/fetch/$s_!2DCz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F265c1c36-9516-4457-86be-2b4fde4cd9ce_2160x897.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2DCz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F265c1c36-9516-4457-86be-2b4fde4cd9ce_2160x897.png" width="728" height="302.3222222222222" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/265c1c36-9516-4457-86be-2b4fde4cd9ce_2160x897.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:897,&quot;width&quot;:2160,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:839703,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://agenticengg.substack.com/i/202568466?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8608d764-d7b4-4000-8445-6230a38a542b_2160x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2DCz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F265c1c36-9516-4457-86be-2b4fde4cd9ce_2160x897.png 424w, https://substackcdn.com/image/fetch/$s_!2DCz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F265c1c36-9516-4457-86be-2b4fde4cd9ce_2160x897.png 848w, https://substackcdn.com/image/fetch/$s_!2DCz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F265c1c36-9516-4457-86be-2b4fde4cd9ce_2160x897.png 1272w, https://substackcdn.com/image/fetch/$s_!2DCz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F265c1c36-9516-4457-86be-2b4fde4cd9ce_2160x897.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>You made it to the end, so here&#8217;s the promised bonus: an exclusive 40% discount for Agentic Engineering readers. Just be sure to register using the link below.</strong></figcaption></figure></div><div class="callout-block" data-callout="true"><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://packt.link/JuExC&quot;,&quot;text&quot;:&quot;GET TICKETS AT 40% OFF&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://packt.link/JuExC"><span>GET TICKETS AT 40% OFF</span></a></p></div>]]></content:encoded></item><item><title><![CDATA[#1: What happens when OpenClaw meets LangGraph?]]></title><description><![CDATA[Building AI agents that can operate beyond the chat window]]></description><link>https://newsletter.agentengineering.co/p/agentic-engineering-1-what-happens</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/agentic-engineering-1-what-happens</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 11 Jun 2026 16:01:57 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/08097fef-4822-47c6-916a-466be4005daa_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Typical agents (like those built on OpenAI Assistants, Copilot Studio, or custom developer frameworks) are often confined to isolated chat dashboards, command-line interfaces, or custom web portals. That&#8217;s the power of Claude Code and similar tools: they can communicate with the outside world, both to gather inputs for research and to deliver results. To connect agents to apps like WhatsApp or Slack, however, developers usually have to manually configure webhooks and middleware.</p><p>Of course, this won&#8217;t do for us. What&#8217;s the point of building custom AI pipelines if they sit locked in a dark room? To make that logic useful, we need to connect it to the real world. We build user authentication, set up secure access controls, write adapters for communication apps like Telegram or Discord, and configure reliable task schedulers.</p><p>By combining LangChain with the OpenClaw ecosystem via the LangClaw framework, you can bridge this gap in an afternoon.</p><div><hr></div><h2>Why combine OpenClaw and LangGraph?</h2><p>OpenClaw began as a configuration-first, out-of-the-box personal AI runtime. It provides developers with a pre-wired ecosystem that includes direct messaging connectors for platforms like WhatsApp and Telegram, persistent memory layers, and native browser automation tools.</p><p>Instead of writing code to manage how an agent communicates with a chat application or remembers previous conversations, users simply write markdown files (SKILL.md) to define functions. OpenClaw handles the operational runtime, API routing, and state storage automatically.</p><p>While this structure is ideal for standard automation, engineering teams often require deeper programmatic control over their agentic reasoning. They need complex, multi-step state machines, advanced semantic search pipelines, and conditional execution paths.</p><p>This is where LangChain excels, and it is exactly why the two systems are paired together.</p><div><hr></div><h2>LangClaw as the glue layer</h2><p>LangClaw acts as the architectural bridge between these two worlds. It is a declarative, Pythonic framework that compiles complex LangChain reasoning loops directly into an OpenClaw-compatible runtime engine.</p><p>Instead of managing external configuration markdown files, you use clean Python decorators to build tools, establish role-based access controls, and schedule autonomous background tasks.</p><p>If you&#8217;re the kind of person who reads something like this and immediately wants to see the code, we&#8217;ve documented a complete LangGraph + OpenClaw build that turns these ideas into a working corporate intelligence agent, covering everything from subagent orchestration and scheduled research tasks to permissions, messaging, and token-efficient utility commands.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://agenticengg.substack.com/p/building-a-corporate-intelligence&quot;,&quot;text&quot;:&quot;Get the playbook here&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://agenticengg.substack.com/p/building-a-corporate-intelligence"><span>Get the playbook here</span></a></p><div class="callout-block" data-callout="true"><div><hr></div><p><strong><a href="https://www.linkedin.com/in/ben-auffarth/">Ben Auffarth</a></strong> is the co-author of <strong><a href="https://www.amazon.com/Generative-LangChain-production-ready-applications-LangGraph/dp/1837022011/ref=sr_1_1?crid=1N2U9JTXXYJKW&amp;dib=eyJ2IjoiMSJ9.w464Pq2KWF9HkflZepVL1kpjAMAAVVcNM0Vj6JhU4srF73UiW9F4cMoNkqdZ7EY0UbEYA4pB_2TelPESVI59OrfEfn0MMu5RdNvqvPEl81OWbNoR8OI7f5hD9w-xFWFFxPKcy0njd16YX9ezuwq35h-gP85lOiGhEpw6NYUtiODyJzF6LI_hjIS90uKsMrKir6vetcEdeULCAAaCuLV5eSgxAoESWB_WNwR4wdnEo7I.fUY6feu7uWCCO9OLr6bF5vObyIAAy6AkXFbbzwEjlwk&amp;dib_tag=se&amp;keywords=generative+AI+with+langchain&amp;qid=1781192284&amp;sprefix=generative+ai+with+langchai%2Caps%2C312&amp;sr=8-1">Generative AI with LangChain</a></strong> and other books. He&#8217;s launched several funded startups and is working as a consultant with Chelsea AI Ventures.</p></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.agentengineering.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Agentic Engineering! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[OpenClaw + LangGraph playbook]]></title><description><![CDATA[Scheduling, subagents, permissions, and messaging in one system]]></description><link>https://newsletter.agentengineering.co/p/building-a-corporate-intelligence</link><guid isPermaLink="false">https://newsletter.agentengineering.co/p/building-a-corporate-intelligence</guid><dc:creator><![CDATA[Tanya D'cruz]]></dc:creator><pubDate>Thu, 11 Jun 2026 15:24:02 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/3707da9e-7dd4-4eeb-954f-61ba5ae3dc22_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>It&#8217;s easy to build an agent that talks. Building one that remembers things, sends messages, runs on a schedule, and generally makes itself useful is a different challenge. That&#8217;s where LangGraph and OpenClaw make an interesting combination.</p><p>Let&#8217;s build one.</p><div><hr></div><p>The first step is creating the primary agent and establishing its responsibilities through a system prompt.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;5782ba58-1ed1-4f3a-8fea-b161795797b7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">import os
from datetime import datetime
from langclaw import Langclaw
from langclaw.gateway.commands import CommandContext
# Initialize the master agent application interface
app = Langclaw(
 system_prompt=(
 &#8220;## Corporate Intelligence Agent\n&#8221;
 &#8220;You are a corporate intelligence analyst. You track market trends &#8220;
 &#8220;and draft precise outreach sequences based on current events.\n&#8221;
 &#8220;Delegate deep multi-source research tasks to the web-researcher subagent.&#8221;
 ),
)</code></pre></div><p>Tools allow the agent to access capabilities beyond language generation. In this example, the agent can retrieve market intelligence data.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7ca10de5-40ad-48a1-9090-954ffd446cd2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Custom Tool: Data Extraction Layer
@app.tool()
async def fetch_market_news(sector: str) -&gt; dict:
 &#8220;&#8221;&#8220;Extract breaking market developments for a specific commercial sector.&#8221;&#8220;&#8221;
 # This acts as a standard tool available to your core reasoning models
 return {
 &#8220;timestamp&#8221;: datetime.now().isoformat(),
 &#8220;sector&#8221;: sector,
 &#8220;headline&#8221;: &#8220;Warehouse Automation Demands Surge Amid Labor Reshuffling&#8221;,
 &#8220;body&#8221;: &#8220;Logistics providers face margin pressures, driving immediate infrastructure upgrades.&#8221;
 }</code></pre></div><p>Rather than placing all reasoning responsibilities inside a single agent, LangClaw allows specialised subagents to handle focused tasks.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;d01edd29-2bc2-41df-8aaa-0dd129cc8712&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext"># Subagent Delegation: Isolated Reasoning Context
app.subagent(
 &#8220;web-researcher&#8221;,
 description=&#8221;Performs deep multi-step web analysis and synthesis.&#8221;,
 system_prompt=&#8221;You are an investigative researcher. Identify technical pain points and list them.&#8221;,
 tools=[&#8221;web_fetch&#8221;, &#8220;web_search&#8221;],
 output=&#8221;channel&#8221;, # Streams output directly back to the active user channel
)</code></pre></div><p>Role-based permissions make it possible to expose different capabilities to different groups of users.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;bd9b40bb-a0ab-4d46-80d1-254f9690ec8c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext"># Role-Based Access Control: Security Gateways
app.role(&#8221;premium_analyst&#8221;, tools=[&#8221;*&#8221;])
app.role(&#8221;basic_user&#8221;, tools=[&#8221;fetch_market_news&#8221;])</code></pre></div><p>Agents do not need to wait for user prompts. Using cron schedules, they can execute tasks automatically.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5e46f0c3-50a6-4a39-837d-1856077cfe22&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Automated Execution: Background Scheduler
@app.cron(&#8221;0 8 * * 1-5&#8221;) # Executes Monday through Friday at 8:00 AM
async def automated_morning_brief():
 &#8220;&#8221;&#8220;Scans market trends autonomously and publishes digests to the user feed.&#8221;&#8220;&#8221;
 raw_data = await fetch_market_news(sector=&#8221;Logistics&#8221;)
 instruction = f&#8221;Synthesize this market report for the main channel: {raw_data[&#8217;body&#8217;]}&#8221;
 # Compile runtime structure down to LangGraph for invocation
 compiled_agent = app.compile()
 analysis = await compiled_agent.ainvoke(
 {&#8221;messages&#8221;: [{&#8221;role&#8221;: &#8220;user&#8221;, &#8220;content&#8221;: instruction}]},
 config={&#8221;configurable&#8221;: {&#8221;subagent&#8221;: &#8220;web-researcher&#8221;}}
 )
 # Broadcast the completed report directly to your active chat channel
 await app.gateway.send_message(
 channel=&#8221;management_feed&#8221;,
 text=f&#8221;&#128276; **Automated Market Digest**\n\n{analysis[&#8217;messages&#8217;][-1].content}&#8221;
 )</code></pre></div><p>Not every task requires an LLM. For predictable workflows, slash commands can bypass the model entirely.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b1e51f67-437c-4362-a291-f3202e250dde&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Slash Command: High-Speed Utility Routing
@app.command(&#8221;template&#8221;, description=&#8221;Generate static outreach copy without AI overhead&#8221;)
async def outreach_template_cmd(ctx: CommandContext) -&gt; str:
 &#8220;&#8221;&#8220;Returns a structured outreach framework directly, bypassing the LLM entirely.&#8221;&#8220;&#8221;
 args = ctx.message.text.split(&#8217;&#8221;&#8217;)
 client_name = args[1] if len(args) &gt; 1 else &#8220;Colleague&#8221;
 # Bypassing the LLM completely saves computation costs and provides sub-millisecond responses
 return (
 f&#8221;&#128233; **System Template Generated**\n\n&#8221;
 f&#8221;Hello {client_name},\n\n&#8221;
 f&#8221;I noted your team&#8217;s footprint in the logistics sector. Given recent automation trends, &#8220;
 f&#8221;I wanted to connect to share notes on how regional hubs are optimizing their margins.\n\n&#8221;
 f&#8221;Regards,\n[Internal System]&#8221;
 )</code></pre></div><p>Once the components are assembled, launching the runtime is straightforward.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;b10f0624-c3fc-490c-9b0a-14857282519f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">if __name__ == &#8220;__main__&#8221;:
 app.run()</code></pre></div><p>Once the runtime is launched, the agent becomes more than a collection of functions and workflows. It becomes a persistent service that can receive requests, execute tasks, and communicate results through your chosen channels.</p><p>Let&#8217;s see how that works.</p><div><hr></div><h3>What happens when the agent goes live</h3><p>When you invoke <code>app.run()</code>, LangClaw spins up a persistent asynchronous gateway server. It does not wait for a single terminal input; it establishes a permanent network listener.</p><p>Interaction happens entirely through your chosen deployment endpoints:</p><ul><li><p><strong>Production chat channels</strong>: By adding tokens for Telegram, Discord, or Slack into your .env configuration file, your script logs into those platforms as a live bot user.</p></li><li><p><strong>LLM invocations</strong>: When a user types a natural language question into the chat, LangClaw pipes the string into your LangChain model, runs the necessary code tools, and formats the response.</p></li><li><p><strong>The slash command route</strong>: When a user types <code>/template &#8220;Alex&#8221;</code>, the gateway identifies the leading slash character. It isolates the instruction, skips the LLM reasoning step, runs the native Python function, and returns the text.</p><div><hr></div></li></ul><h3>Hosting your agent for less than &#163;5 per month</h3><p>Deploying a traditional enterprise application often brings significant server management costs. However, because this framework relies on an event-driven architecture and asynchronous IO libraries, it requires minimal computing resources.</p><p>You can host this entire stack for under &#163;5 per month using modern cloud infrastructure.</p><h4>Virtual Private Servers (VPS)</h4><p>Platforms like DigitalOcean, Hetzner, or Linode offer entry-level Linux instances with 1GB of RAM and 1 CPU core. Because LangClaw uses a non-blocking asyncio event loop, a single micro-instance can easily manage hundreds of concurrent chat messages and automated cron tasks without breaking a sweat.</p><h4>Process Management</h4><p>To ensure your agent runs continuously, you wrap the script using a Linux system process supervisor like systemd or Supervisor. If the server reboots or encounters an unhandled API timeout, the operating system brings the agent back online instantly, providing 99.9% uptime with zero manual maintenance.</p><div><hr></div><p>We&#8217;ve set up an agent, we&#8217;ve talked about messaging and about the advantages of full control, about the hosting. Building your automation stack with this methodology yields three immediate operational improvements:</p><ul><li><p><strong>Zero-token utilities</strong>: By routing routine structural actions through slash commands, you bypass LLM token fees entirely. This provides instantaneous responses while eliminating API runtime costs for predictable tasks.</p></li><li><p><strong>Decoupled architecture</strong>: Your core business logic remains completely separate from your communication channels. If your team decides to migrate from Telegram to Slack, you alter a single environment variable. The underlying tools, subagents, and schedules require no modifications.</p></li><li><p><strong>Production readiness</strong>: You move from an abstract local script to an enterprise-grade agent. The system manages its own user security roles, schedules its own data retrieval cycles, and operates autonomously around the clock.</p></li></ul><p>Let us know how you get on!</p>]]></content:encoded></item></channel></rss>