<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Retrofit]]></title><description><![CDATA[Uncovering agentic applications in enterprise spaces, building mental models, and writing market maps ]]></description><link>https://www.theretrofit.ai</link><image><url>https://substackcdn.com/image/fetch/$s_!VtD5!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8154578e-4a95-4aba-8379-06ffa5883fc7_128x128.png</url><title>The Retrofit</title><link>https://www.theretrofit.ai</link></image><generator>Substack</generator><lastBuildDate>Wed, 02 Sep 2026 10:56:34 GMT</lastBuildDate><atom:link href="https://www.theretrofit.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[The Retrofit]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[theretrofit@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[theretrofit@substack.com]]></itunes:email><itunes:name><![CDATA[The Retrofit]]></itunes:name></itunes:owner><itunes:author><![CDATA[The Retrofit]]></itunes:author><googleplay:owner><![CDATA[theretrofit@substack.com]]></googleplay:owner><googleplay:email><![CDATA[theretrofit@substack.com]]></googleplay:email><googleplay:author><![CDATA[The Retrofit]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Compliance AI Market Map]]></title><description><![CDATA[Three layers of compliance work - and the path to autonomy]]></description><link>https://www.theretrofit.ai/p/the-compliance-ai-market-map</link><guid isPermaLink="false">https://www.theretrofit.ai/p/the-compliance-ai-market-map</guid><dc:creator><![CDATA[The Retrofit]]></dc:creator><pubDate>Tue, 11 Aug 2026 13:39:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!eJBm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most conversations about &#8220;AI compliance&#8221; flatten a large market into a single category. In practice, compliance software automates three different kinds of work, each with different buyers, incumbents, data, and technical difficulty:</p><ol><li><p><strong>Regulatory intelligence - know the rules.</strong> Ingest, track, and interpret regulations and regulatory change, then turn them into structured, monitored obligations.</p></li><li><p><strong>Compliance decisioning - decide the case.</strong> Apply external regulations, internal policies, and prior decisions to a customer, transaction, document, communication, vendor, or questionnaire.</p></li><li><p><strong>Compliance operations - execute the workflow.</strong> Gather evidence, route cases, capture approvals, file, remediate, and preserve an audit trail.</p></li></ol><p>The three layers also reflect how the market has developed. Regulatory intelligence is the oldest and most mature category; regtech companies have sold regulatory data and change-management software for more than a decade. Decisioning is where many AI-native startups begin because case-level review is bounded, measurable, and well suited to language models. Operations is the current frontier, as agents become capable of handling variable, unstructured, multi-step work across documents and systems.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.theretrofit.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><strong>Autonomous execution</strong> is starting to appear in specific, lower-risk workflows, but broad autonomy barely exists. The constraint is partly model reliability and partly whether institutions, regulators, and customers are willing to let a machine make and act on a consequential decision.</p><p>I think that value and defensibility generally increase as a product owns more of the decision-to-action loop. A company does not need to begin with regulatory intelligence or move through every layer in order. But products that move from providing information to owning decisions, execution, and governed authority capture more of the workflow, more institutional memory, and more control over how compliance work gets done.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eJBm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eJBm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png 424w, https://substackcdn.com/image/fetch/$s_!eJBm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png 848w, https://substackcdn.com/image/fetch/$s_!eJBm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png 1272w, https://substackcdn.com/image/fetch/$s_!eJBm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eJBm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png" width="1456" height="942" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:942,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:332868,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.theretrofit.ai/i/209717629?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eJBm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png 424w, https://substackcdn.com/image/fetch/$s_!eJBm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png 848w, https://substackcdn.com/image/fetch/$s_!eJBm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png 1272w, https://substackcdn.com/image/fetch/$s_!eJBm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f152005-d6ab-4a39-bf73-0d06fd1c8885_2720x1760.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h4>1. Why compliance AI is suddenly investable</h4><p><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;a16z&quot;,&quot;id&quot;:2315700,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!-aGV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff698a0c5-1fee-40a7-a33c-80609431ae31_400x400.png&quot;,&quot;uuid&quot;:&quot;b24974c2-fdd9-4829-b97d-f0a3aec3b773&quot;}" data-component-name="MentionToDOM"></span> &#8217;s <em><a href="https://a16z.com/everything-everywhere-is-compliance/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Everything, Everywhere Is Compliance</a></em>, published by James da Costa and Angela Strange in May 2026, makes the macro case. The United States has more than 400,000 compliance officers representing over $40 billion in annual labor spend. The work is also becoming harder to staff, with annual churn above 20% while regulatory complexity continues to grow.</p><p>The market was already large. What changed is how much of the work software can perform. Vision-language models can interpret documents that traditional OCR could only extract, while frontier models have improved at reading rules, applying them to facts, and identifying exceptions. Computer-use and longer-horizon agents extend this beyond review, allowing systems to carry context across documents, databases, and legacy applications. Benchmarks are not proof of production reliability, but they show that the addressable workflow is expanding.</p><p>The buyer calculus is changing as well. Faster KYC improves onboarding, faster marketing review increases speed to market, and better investigations reduce backlogs. Compliance software is no longer sold only as a way to reduce headcount or avoid fines. <a href="https://www.ycombinator.com/rfs/?status=new&amp;utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack#ai-native-compliance-infrastructure">YC&#8217;s Fall 2026 Requests for Startups</a> reflects that shift, calling for AI-native infrastructure that replaces fragmented tools and manual workflows.</p><p>a16z describes compliance as regulation, the software trying to codify it, and the people connecting the two. The three layers in this piece describe which part of that human work a product automates; the autonomy axis describes how much responsibility moves to the system.</p><div><hr></div><h4>2. The three layers of compliance AI</h4><p><em>How AI moves from interpreting regulation to executing compliance work</em></p><p><strong>Layer 1 &#8212; Regulatory intelligence: know the rules</strong></p><blockquote><p>Regulatory intelligence turns a constantly changing body of regulation into something a compliance team and eventually an agent can act on. The workflow begins with horizon scanning but continues through a consolidated rules inventory, applicability analysis, obligation extraction, and mapping to internal policies, procedures, controls, risks, and owners. Provenance has to remain intact throughout, because the company needs to know where an obligation came from, which version was in force, and what changed.</p><p><strong>Regulatory source &#8594; obligation &#8594; policy &#8594; procedure &#8594; control</strong></p></blockquote><p>This is the oldest and most mature layer. Regtech companies have collected, classified, and distributed regulatory information for more than a decade, so the market already has established buyers, large incumbents, and recognizable product categories.</p><p>It is also consolidating. <a href="https://cube.global/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">CUBE</a> acquired <a href="https://cube.global/resources/news/cube-acquires-global-regulatory-intelligence-businesses-from-thomson-reuters?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Thomson Reuters Regulatory Intelligence and Oden</a> in 2024. The combined company now covers more than 750 jurisdictions and employs close to 250 regulatory subject-matter experts and legal and compliance professionals alongside its AI products. Other established players include <a href="https://www.wolterskluwer.com/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Wolters Kluwer</a>, <a href="https://corlytics.com/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Corlytics</a>, in which Verdane made a <a href="https://verdane.com/verdane-partners-with-global-regulatory-software-leader-corlytics/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">majority investment</a> in 2024, and <a href="https://www.archerirm.com/press-releases/archer-acquires-compliance.ai-to-drive-ai-powered-regulatory-compliance-and-risk-management?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Compliance.ai</a>, which Archer acquired in the same year.</p><p><a href="https://www.bloombergindustry.com/press-releases/bloomberg-industry-group-regology/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Bloomberg Industry Group&#8217;s acquisition of Regology</a> in June 2026 follows the same logic. Bloomberg combined its legal, tax, and government information with software that monitors regulatory change, determines which rules apply to a particular company, and connects those changes to internal policies, risks, and controls.</p><p>AI changes this layer by turning regulatory content into company-specific work, moving the product from identifying the rules toward deciding how they apply.</p><p><a href="https://www.ascentregtech.com/our-difference/change-management/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Ascent</a> is a good example. Its platform creates company-specific obligation inventories and maps them to internal policies, procedures, and controls. Its <a href="https://www.ascentregtech.com/blog/ascent-technologies-acquires-horizon-scanning-solution-provider-waymark/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">acquisition of Waymark</a> in 2024 combined horizon scanning with obligations management, extending the product from detecting regulatory change to showing where that change matters inside the institution.</p><p><strong>Layer 2 &#8212; Compliance decisioning: decide the case</strong></p><blockquote><p>Compliance decisioning applies external regulations, internal policies, and prior decisions to a specific customer, transaction, document, communication, vendor, or questionnaire. The system evaluates the facts and produces a determination or draft supported by evidence, rationale, and citations. Where the answer is uncertain or the risk exceeds a defined threshold, it escalates the case to a human reviewer.</p><p><strong>Case facts + rules + policy + precedent &#8594; determination + rationale + evidence</strong></p></blockquote><p>The distinction from regulatory intelligence is simple. Regulatory intelligence determines what the business must do in general; compliance decisioning determines what it should do in this case. Its central value is replacing hours spent searching policies, source documents, evidence, and prior decisions with a grounded first determination that a reviewer can verify.</p><p>This is a natural starting point for AI-native companies because the work can be bounded and evaluated. A buyer can compare the system&#8217;s conclusions with those of experienced reviewers, measure reductions in review time, and test whether similar cases produce consistent results. The product does not initially need authority to approve or file anything. It can create value by completing the first pass and sending exceptions to the appropriate reviewer.</p><p>I would draw the line here: a decisioning product applies governed rules, company policy, or precedent to a specific set of facts and produces an answer that can be checked against a human reviewer&#8217;s. That distinguishes compliance decisioning from generic document generation or retrieval.</p><p>This decisioning pattern is appearing across several compliance workflows.</p><p><strong>Marketing and communications review.</strong> Companies including <a href="https://www.norm.ai/platform/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Norm Ai</a>, <a href="https://www.sedric.ai/solutions/marketing?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Sedric</a>, and <a href="https://www.hadrius.com/solutions/marketing?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Hadrius</a> translate requirements such as UDAAP, SEC and FINRA rules, regulatory guidance, and internal company policy into review criteria that can be applied to financial-services marketing. The systems review advertisements, emails, websites, social posts, and other customer communications for prohibited claims, missing disclosures, unsupported statements, and policy violations.</p><p>Life sciences has a specialized version of the same workflow called medical, legal, and regulatory review. <a href="https://www.veeva.com/resources/veeva-acquires-copli-launches-veeva-falcon-mlr-to-accelerate-content-review/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Veeva&#8217;s acquisition of Copli and launch of Falcon MLR</a> in June 2026 brought agentic review of promotional and medical content into Veeva&#8217;s existing life-sciences platform. The acquisition validates marketing review as a decisioning wedge, but it also shows the strategic risk: an incumbent that already owns the content and approval workflow can absorb the review product.</p><p><strong>DDQs, RFPs, and security questionnaires.</strong> Products such as <a href="https://1up.ai/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">1up</a>, <a href="https://www.arphie.ai/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Arphie</a>, and <a href="https://tribble.ai/security-questionnaires/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Tribble</a> generate answers from approved company materials, policies, prior responses, and supporting evidence. The value is not simply writing prose faster. It is determining which source is authoritative, retrieving previously approved language, attaching evidence to each answer, and routing low-confidence responses to the person qualified to approve them.</p><p><strong>Third-party risk.</strong> <a href="https://www.coverbase.com/solutions/third-party-risk-management?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Coverbase</a> analyzes questionnaires, contracts, and third-party evidence against a company&#8217;s controls, surfaces exceptions, and produces a risk assessment. Its broader platform also handles vendor intake, evidence collection, follow-ups, continuous monitoring, and remediation. The risk determination sits in decisioning, while the surrounding vendor lifecycle pushes the product into compliance operations.</p><p><strong>KYC, KYB, and financial-crime review.</strong> <a href="https://arva.ai/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Arva AI</a> applies an institution&#8217;s policies across customer and business onboarding, sanctions screening, and <a href="https://arva.ai/product/transaction-monitoring-ai?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">transaction-monitoring alerts</a>. It assembles the relevant identity, ownership, transaction, and screening data, produces a risk determination supported by evidence, and routes uncertain or higher-risk cases to human investigators.</p><p><a href="https://www.parcha.ai/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Parcha</a> initially offered a similar product focused on automating KYC, KYB, and compliance reviews for banks and fintechs. It has since retired that product line and relaunched as <a href="https://grep.ai/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Grep AI</a>, applying the agent architecture it developed in financial-services compliance to a broader set of research-intensive, high-stakes workflows.</p><p>At the same time, the identity platforms underneath these products are moving into the same decisioning and workflow layer. <a href="https://withpersona.com/product/workflows?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Persona</a>, which <a href="https://www.prnewswire.com/news-releases/persona-raises-200m-at-2b-valuation-to-build-the-verified-identity-layer-for-an-agentic-ai-world-302442649.html?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">raised $200 million at a $2 billion valuation</a>, has expanded from identity verification into automated decisions, follow-ups, workflows, and case management. </p><p>I think this is the more important takeaway from Parcha&#8217;s pivot. The KYC and KYB platforms underneath standalone AI reviewers are expanding into decisioning and workflows themselves, leaving less room to build a durable moat at the review layer. The larger opportunity is to own the case, the workflow, or proprietary data or to take the underlying agent architecture into a broader market, as Parcha did.</p><p><strong>Product and trade compliance.</strong> <a href="https://www.usetrava.com/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Trava</a> and <a href="https://www.complir.io/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Complir</a> use the same underlying architecture: translate a body of regulation into machine-readable logic, map it to product data, and make a product-level compliance determination.</p><p><strong>Audit and regulatory submissions.</strong> <a href="https://denki.ai/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Denki</a> collects evidence, tests controls, and produces traceable workpapers for internal-audit and compliance programs. <a href="https://www.ritivel.com/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Ritivel</a> turns clinical data and prior submissions into reproducible, source-linked drafts for life-sciences filings. Both begin with decisioning but move into operations as they assemble the evidence and produce the auditor- or regulator-ready document.</p><p>The durable asset in decisioning is not access to the model. It is the system built around it: the mapping between regulations and company policy, the evidence supporting each conclusion, domain-specific evaluations, and the institutional memory created as reviewers accept, reject, or modify recommendations. Capturing why a reviewer overrode the system is especially important, because that is how individual judgments become reusable precedent rather than disappearing into an audit log.</p><p>Decisioning is therefore both an attractive entry point and an unstable place to stop. It creates measurable value without requiring the product to act autonomously, but it becomes more embedded when it also gathers evidence, routes exceptions, manages approvals, and records the final outcome. That is the natural path from decisioning into compliance operations.</p><p><strong>Layer 3 &#8212; Compliance operations: execute the workflow</strong></p><blockquote><p>Compliance operations automates the work that compliance analysts and operations teams perform to move a case from intake to closure. This includes gathering evidence, researching across systems, requesting missing information, routing the case, capturing approvals, preparing filings, managing remediation, and preserving an audit trail. Decisioning provides the judgment within the case; operations carries the case through the governed workflow.</p><p><strong>Intake &#8594; evidence &#8594; decision &#8594; approval or action &#8594; audit record</strong></p></blockquote><p>Workflow automation is not new. Traditional systems can move a case through a predefined sequence, but they generally depend on structured inputs and explicit rules for every step. When a document is incomplete, an answer is ambiguous, or the next action depends on information scattered across several systems, a person has historically carried the context and determined what to do next.</p><p>Agents expand the kind of work software can perform. They can read unstructured documents, retrieve information from multiple systems, determine which evidence is missing, adapt the next step, and operate the applications through which compliance work gets completed. The workflow still needs defined permissions, approval gates, and escalation rules, but it no longer has to anticipate every possible path in advance.</p><p>Several companies illustrate how this layer is developing.</p><p><strong>Financial-crime operations.</strong> <a href="https://www.unit21.ai/products/ai-agent?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Unit21</a> began as a fraud and AML platform and is adding agents inside the system where alerts, investigations, and regulatory filings already live. Its agents gather evidence, investigate cases against the institution&#8217;s procedures, recommend dispositions, draft regulator-ready narratives, and produce a transparent log of the data accessed and steps performed. The institution continues to define its risk appetite, thresholds, approval requirements, and escalation paths.</p><p><strong>SOX testing.</strong> <a href="https://www.petual.ai/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Petual</a> imports a company&#8217;s risk and control matrix, maps evidence to the relevant controls and samples, executes the prescribed testing procedures, and generates audit-ready workpapers. Deviations are surfaced for review and remediation. The determination whether the control operated effectively is decisioning; collecting the evidence, performing the test, documenting the result, and moving exceptions toward resolution are compliance operations.</p><p><strong>Financial licensing.</strong> <a href="https://www.brico.ai/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Brico</a> manages the operational work surrounding state financial licenses, including applications, renewals, periodic reports, amendments, and status tracking. Licensing is a particularly clear operations workflow because the work extends far beyond determining which licenses are required. Companies must assemble information from across the business, complete forms, track jurisdiction-specific deadlines, respond to deficiencies, and maintain each license after approval.</p><p>There are two routes into compliance operations: incumbents can add agents to their existing systems of record, while agent-first companies can absorb the workflow, integrations, and case data around them. <a href="https://a16z.com/everything-everywhere-is-compliance/?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">a16z</a> frames the buyer&#8217;s choice as keeping the incumbent as a backend, rebuilding internally, or buying an AI-native replacement. <a href="https://www.ycombinator.com/rfs/?status=new&amp;utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack#ai-native-compliance-infrastructure">YC</a> similarly points toward AI-native products that consolidate fragmented compliance tools.</p><p>I think the two routes converge. Incumbents will add agents, while agent-first companies will build toward systems of record. Both want to own not only the final decision but how the work was performed: the evidence, policy, exceptions, approvals, and outcome. That execution history becomes institutional memory and makes the product the place where compliance work is executed, governed, and remembered.</p><div><hr></div><h4>3. The trust stack: what gives AI permission to act</h4><p>The three layers describe what a compliance product does. The trust stack describes what it needs to do that work reliably, safely, and within the institution&#8217;s authority. The more judgment, workflow, and consequential action a product assumes, the more of this infrastructure it needs.</p><p><strong>Understanding the work</strong></p><ul><li><p><strong>Document competence:</strong> Read PDFs, Word documents, spreadsheets, screenshots, IDs, filings, questionnaires, and marketing assets without losing the structure or context that controls the decision.</p></li><li><p><strong>Company context and master-data mapping:</strong> Connect regulations, evidence, and decisions to the company&#8217;s legal entities, customers, products, accounts, jurisdictions, controls, owners, and systems. Without that mapping, the system may understand the rule but apply it to the wrong part of the business.</p></li></ul><p><strong>Grounding the decision</strong></p><ul><li><p><strong>Traceability:</strong> Tie every material conclusion to the evidence, internal policy, and regulatory language supporting it. The system needs to show not only which sources it retrieved, but why they governed the decision.</p></li><li><p><strong>Institutional memory:</strong> Preserve prior filings, decisions, approved and rejected language, exceptions, reversals, and precedent. That history must remain linked to the policies and rules in effect at the time.</p></li><li><p><strong>Auditability:</strong> Record who or what decided, when, under which policy and model version, and any human review or override.</p></li></ul><p><strong>Governing the system</strong></p><ul><li><p><strong>Permissions and information barriers:</strong> Treat access as a first-class control. The system should retrieve information, expose cases, and use tools only within the boundaries of the user, entity, jurisdiction, and task.</p></li><li><p><strong>Authority:</strong> Distinguish between permission to read, recommend, escalate, approve, file, reclassify, contact a customer, and change a system of record. Each action requires its own authorization boundary.</p></li><li><p><strong>Human control:</strong> Support review, correction, rejection, and escalation when evidence is incomplete, policy is ambiguous, risk is high, or human judgment is required by law or internal policy.</p></li></ul><p><strong>Executing the workflow</strong></p><ul><li><p><strong>Workflow orchestration:</strong> Sequence the work, route cases, manage handoffs, enforce approval gates, and escalate exceptions from intake through final disposition. Historically, this connective tissue was a person carrying context between documents, inboxes, databases, and applications.</p></li><li><p><strong>Operational competence:</strong> Perform the concrete actions inside the workflow-redline a document, populate a template, preserve formatting, request missing information, upload evidence, submit a filing, export an artifact, and write the outcome back to the correct system.</p></li></ul><p>A regulatory-intelligence product may only need a subset of these capabilities. A decisioning product needs most of them in at least a basic form to produce reliable case-level judgments. An operations product and especially one acting with limited human involvement needs the full trust stack because its mistakes can propagate into customer communications, filings, approvals, and systems of record.</p><p>The trust stack converts model capability into operational permission. It is what allows a compliance product to move from advising to doing, and eventually to acting within defined limits.</p><div><hr></div><h4>4. The autonomy frontier</h4><blockquote><p>Autonomous compliance closes the loop. The system detects an issue, gathers context, makes a determination, takes an authorized action, records its rationale, and updates the relevant systems of record. Humans define the guardrails and supervise exceptions rather than executing every routine case.</p><p><strong>Detect &#8594; investigate &#8594; decide &#8594; act &#8594; record &#8594; monitor</strong></p></blockquote><p>This is different from rules-based automation, where the trigger, logic, and response are specified in advance. An autonomous agent can navigate a variable path, determine what evidence it needs, choose which tools to use, and resolve ambiguity. The hard part is not giving it the ability to act; it is making sure it stops when it lacks the evidence, confidence, or authority to continue.</p><p><a href="https://www.unit21.ai/agentic-ai-maturity?utm_source=theretrofit&amp;utm_medium=substack&amp;utm_campaign=ai_compliance_stack">Unit21 and Chartis Research&#8217;s agentic AI maturity spectrum</a> frames autonomy as a progression:</p><ol><li><p><strong>Rules-based automation:</strong> Predefined conditions trigger deterministic workflows.</p></li><li><p><strong>AI-assisted:</strong> The system recommends actions, but humans perform the work and decide.</p></li><li><p><strong>Agentic with human oversight:</strong> Agents execute the workflow, while humans approve material decisions.</p></li><li><p><strong>Autonomous within governed limits:</strong> The system completes permitted decisions and actions, while humans supervise exceptions and quality.</p></li></ol><p>These levels are configured by workflow, jurisdiction, risk tier, and action. A system might close a low-risk false positive, recommend an outcome on a medium-risk case, and require explicit approval before rejecting a customer or submitting a regulatory filing. &#8220;Human in the loop&#8221; is not one setting; it is a set of controls around individual actions.</p><p>The constraint is both technical and institutional. The system has to remain reliable, recognize incomplete evidence, and produce a defensible record of what it did. The institution has to determine what can be delegated, scope the agent&#8217;s authority, and assign accountability when it errs. Autonomy will therefore arrive workflow by workflow, and the ability to act safely and prove why the action was appropriate may become one of the deepest moats in compliance.</p><div><hr></div><h4>5. Where founders should start</h4><p>The best entry point is usually a narrow, expensive workflow with a visible output and a human process against which the product can be evaluated.</p><p><strong>Regulatory intelligence</strong> works as a starting point when rules change frequently across jurisdictions and the product can propagate those changes into policies, controls, risks, and owners. A better summarizer is not enough in a mature and consolidating market. The product needs to determine applicability and connect regulatory change to company-specific work.</p><p><strong>Compliance decisioning</strong> works when there is a high-volume review task with a clear human benchmark: marketing review, KYC analysis, alert triage, medical review, product classification, vendor assessment, or due-diligence questionnaires. Decisioning makes it relatively easy to demonstrate value, but it is also easier for an incumbent to bundle if the startup does not capture the surrounding evidence, precedent, or workflow.</p><p><strong>Compliance operations</strong> works when the process spans multiple systems and produces a durable artifact: an investigation, filing, license, audit workpaper, approval record, or completed onboarding case. This is where a startup has the clearest path to owning the system of record for the work, because the product captures not only the result but the evidence, approvals, exceptions, and execution history behind it.</p><p><strong>Autonomy should be a direction, not a day-one claim.</strong> A company earns the right to automate consequential actions workflow by workflow, after proving reliability and building the necessary permissions, evaluation systems, escalation paths, and audit controls. The strongest initial use cases are usually routine, high-volume decisions with clear policy boundaries and a safe path for escalation.</p><p>The founder test is simple:</p><blockquote><p>Does the product only tell the customer what is compliant, or is it becoming the place where compliant work is executed, governed, and remembered?</p></blockquote><p>The first can be a valuable tool. The second has a stronger path to becoming infrastructure.</p><div><hr></div><h4>6. From compliance answers to compliance actions</h4><p>Not every company needs to span all three layers. Regulatory-intelligence and decisioning products can remain valuable on their own, but a company selling an agentic future needs to own more of the loop: the workflow, evidence, permissions, decision history, and eventually the authority to act.</p><p>The more of that loop a product owns, the harder it becomes to replace. Most compliance AI today helps companies determine what is compliant; the next generation will make the work compliant, and eventually keep it that way.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.theretrofit.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The LLM Inference Stack for AI Agents Is Being Rebuilt. Here's the Map. ]]></title><description><![CDATA[Why inference is becoming an application-specific systems problem-and the software stack emerging underneath it.]]></description><link>https://www.theretrofit.ai/p/the-llm-inference-stack-for-ai-agents</link><guid isPermaLink="false">https://www.theretrofit.ai/p/the-llm-inference-stack-for-ai-agents</guid><dc:creator><![CDATA[The Retrofit]]></dc:creator><pubDate>Fri, 17 Jul 2026 03:19:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7a-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>So far, most companies building AI products have been in what I call the <strong>accept and consume</strong> wave.</p><p>The playbook was simple: pick a frontier model, call the API, and build your product around whatever latency, pricing, throughput, and rate limits the provider offered. That wasn&#8217;t because companies didn&#8217;t care about inference. It was because they were trying to build useful products as quickly as possible. Foundation model APIs made that incredibly easy. They handled the infrastructure, gave developers instant access to frontier models, and continuously upgraded those models over time. If paying a little more for inference meant shipping months faster, it was an easy trade.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.theretrofit.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>That tradeoff is starting to break.</p><p>We&#8217;re entering the second wave: <strong>own and optimize</strong>.</p><p>As AI products mature, inference is no longer something you simply consume - it&#8217;s becoming something you optimize. Not just because it&#8217;s expensive, but because inference is becoming an application-specific systems problem. A voice AI company, a coding copilot, a research agent, and a document processing pipeline are all calling an LLM, but they have completely different latency, throughput, concurrency, and cost requirements. The default API has to optimize for everyone. Your application doesn&#8217;t.</p><p>That naturally raises the question: <strong>why now?</strong></p><p>What changed over the last two or three years that suddenly made inference something companies want to own instead of simply consume?</p><p>It wasn&#8217;t one breakthrough. It was a handful of shifts that all happened at roughly the same time.</p><div><hr></div><h4>What Changed?</h4><p><em>The forces reshaping how companies consume inference</em></p><p><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">01 &#8212; The Workload Changed</span></p><p><strong>From Conversational to Agentic</strong></p><p>Most AI applications up until this point were conversational or workflow-based. A user would submit a prompt, maybe the application would make a couple of LLM calls behind the scenes, and then return an answer. Teams spent most of their time improving prompts, adding examples, tweaking outputs, or chaining together a handful of predefined steps. Token usage was relatively predictable because the application controlled the workflow.</p><p>Agentic systems are different.</p><p>A single user prompt is no longer a single request. It&#8217;s the beginning of an execution loop. The model reasons, plans, decides whether it needs to use tools, retrieves context, calls those tools, evaluates the results, and repeats until the task is complete. If you&#8217;ve used Claude Code or Codex, you&#8217;ve already seen this. You ask it to fix a bug, and behind the scenes it searches your codebase, opens files, writes code, runs tests, retries when something fails, and keeps going until it has a solution.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c7a-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c7a-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png 424w, https://substackcdn.com/image/fetch/$s_!c7a-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png 848w, https://substackcdn.com/image/fetch/$s_!c7a-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png 1272w, https://substackcdn.com/image/fetch/$s_!c7a-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c7a-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png" width="1456" height="947" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:947,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:137068,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://theretrofit.substack.com/i/205436014?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c7a-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png 424w, https://substackcdn.com/image/fetch/$s_!c7a-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png 848w, https://substackcdn.com/image/fetch/$s_!c7a-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png 1272w, https://substackcdn.com/image/fetch/$s_!c7a-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f45b07d-2a27-4f25-b446-6382e61e0c36_1654x1076.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 1. A chat interaction usually ends after one or two model calls. An agent loops through planning, tool use, evaluation, and retries until the task is complete.</em></p><p>The important shift is simple: <strong>inference no longer scales with prompts. It scales with the amount of work an agent performs on your behalf.</strong> Every reasoning step, tool call, retry, or planning step can trigger another inference request. Some tasks might finish in a handful of model calls. Others can require hundreds.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Wxli!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Wxli!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png 424w, https://substackcdn.com/image/fetch/$s_!Wxli!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png 848w, https://substackcdn.com/image/fetch/$s_!Wxli!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png 1272w, https://substackcdn.com/image/fetch/$s_!Wxli!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Wxli!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png" width="1456" height="713" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:713,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:106392,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://theretrofit.substack.com/i/205436014?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Wxli!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png 424w, https://substackcdn.com/image/fetch/$s_!Wxli!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png 848w, https://substackcdn.com/image/fetch/$s_!Wxli!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png 1272w, https://substackcdn.com/image/fetch/$s_!Wxli!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3608af5-30c7-42b9-87d0-b8ade477918e_1580x774.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 2. The economics change quickly. What was pennies for a chat session becomes meaningful infrastructure spend once agents continuously execute work.</em></p><p>The surprising part is <strong>where those tokens actually go</strong>.</p><p>Most people assume they&#8217;re paying for the final answer. In reality, a large portion of the inference bill is everything required to get there: conversation history being carried forward, system prompts, tool definitions, planning, retries, and intermediate reasoning. As workloads become more agentic, the overhead around the answer grows alongside the answer itself.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UEAl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UEAl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png 424w, https://substackcdn.com/image/fetch/$s_!UEAl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png 848w, https://substackcdn.com/image/fetch/$s_!UEAl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png 1272w, https://substackcdn.com/image/fetch/$s_!UEAl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UEAl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png" width="1456" height="403" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:403,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:90700,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://theretrofit.substack.com/i/205436014?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UEAl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png 424w, https://substackcdn.com/image/fetch/$s_!UEAl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png 848w, https://substackcdn.com/image/fetch/$s_!UEAl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png 1272w, https://substackcdn.com/image/fetch/$s_!UEAl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc413a0f-0f4c-4f4d-bdb3-44072e1f79ef_1526x422.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 3. In an agentic workflow, much of the inference budget is spent on orchestration rather than the final user-visible output.</em></p><p><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">02 &#8212; The Constraints Changed</span></p><p>The workload changed. The optimization target changed with it.</p><p>Once companies started building agents instead of simple chatbots, they realized there isn&#8217;t one definition of &#8220;good&#8221; inference anymore. Every application is trying to optimize for a different combination of latency, throughput, and cost, which means there isn&#8217;t a single inference stack that&#8217;s optimal for everyone.</p><p>Underneath, every inference request has two phases. The first is <strong>prefill</strong>-processing the input prompt before the model can start generating a response. This determines <strong>Time to First Token (TTFT)</strong>, or how long the user waits before seeing the first word. The second is <strong>decode</strong>-generating the response one token at a time. This determines <strong>Time Per Output Token (TPOT)</strong>, or how quickly the rest of the response streams back. Most optimizations in the serving stack are trying to improve one or find a better balance between the two, without blowing up cost.</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NmFY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NmFY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png 424w, https://substackcdn.com/image/fetch/$s_!NmFY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png 848w, https://substackcdn.com/image/fetch/$s_!NmFY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png 1272w, https://substackcdn.com/image/fetch/$s_!NmFY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NmFY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png" width="1456" height="893" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:893,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:128383,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://theretrofit.substack.com/i/205436014?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!NmFY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png 424w, https://substackcdn.com/image/fetch/$s_!NmFY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png 848w, https://substackcdn.com/image/fetch/$s_!NmFY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png 1272w, https://substackcdn.com/image/fetch/$s_!NmFY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F198fc9e2-7c38-4772-8a19-257911a5583d_1610x988.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 4. Different AI applications optimize for different combinations of latency, throughput, and cost. There isn&#8217;t a universally optimal inference stack, only one that&#8217;s optimal for the workload.</em></p><p>This is where the tradeoffs become obvious. A voice AI company wants the model speaking in a couple hundred milliseconds because every delay feels awkward. A coding copilot also cares deeply about latency, but it has to maintain that experience across thousands of concurrent users. A research agent running overnight is happy to wait a few extra minutes if it cuts inference costs in half. A batch document processing pipeline doesn&#8217;t care about latency at all, it wants maximum throughput and GPU utilization.</p><p>They&#8217;re all building agentic applications. They&#8217;re all running inference. But they&#8217;re solving completely different optimization problems.</p><p>That&#8217;s why the default API stops being enough. Foundation model providers have to optimize for millions of developers across millions of different workloads. Your application only has one workload. The closer inference becomes to your product, the more those default tradeoffs stop making sense.</p><p><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">03 &#8212; Open source caught the frontier</span></p><p>DeepSeek, Qwen, Llama, Kimi K2, GLM- the open ecosystem has matured quickly. There isn&#8217;t one &#8220;best&#8221; model anymore. Different models lead on different benchmarks and different workloads. Some are better at reasoning, others at coding, multilingual tasks, or long-context retrieval. That gives companies something they didn&#8217;t really have two years ago: choice.</p><div><hr></div><h4>Where is inference actually being optimized?</h4><p>By this point, it&#8217;s clear inference is no longer just an API call. It&#8217;s an optimization problem. So where is all of this innovation actually happening?</p><p>The answer is across the entire stack. Some companies are rebuilding hardware. Others are rebuilding the software layer sitting on top of it. Both matter, but they operate on very different time horizons.</p><div><hr></div><h4>Hardware (Brief, But Honest)</h4><p>Inference isn&#8217;t just changing software. It&#8217;s changing hardware too.</p><p>Training and inference stress hardware differently. Training is dominated by massive parallel computation. Inference is dominated by latency, memory bandwidth, and serving requests efficiently. That&#8217;s why we&#8217;re seeing a new generation of inference-first silicon alongside GPUs.</p><p>Three themes stand out.</p><p><strong>First, specialized inference chips are becoming real.</strong> Companies like <a href="https://groq.com/">Groq</a>, <a href="https://www.cerebras.ai/">Cerebras</a>, and <a href="https://www.d-matrix.ai/">d-Matrix</a> proved there was room for architectures optimized specifically for inference rather than general-purpose GPU workloads. Whether those companies remain independent matters less than the fact that the industry has validated the need for inference-first hardware.</p><p><strong>Second, hardware and models are starting to be designed together.</strong> Google has been doing this with TPUs for years, and OpenAI's reported Jalape&#241;o effort points in the same direction. General-purpose GPUs have to support every workload. Custom silicon only has to optimize for one. The competitive advantage is no longer just the chip or the model, it's how well the entire stack works together.</p><p><strong>Third, inference is pushing hardware toward full-stack systems.</strong> NVIDIA&#8217;s answer is Rubin. It isn&#8217;t just a new GPU, it&#8217;s an entire rack designed as one system, combining CPUs, GPUs, networking, interconnects, and software. That&#8217;s because inference isn&#8217;t just a compute problem anymore. Models like DeepSeek and Qwen3 constantly move tokens between experts running on different GPUs. The chip still matters, but so does everything connecting those chips together. The competitive advantage is no longer the accelerator. It&#8217;s the system around it.</p><p>Purpose-built inference hardware will absolutely matter.</p><p>But hardware refresh cycles happen over years. Software improvements ship every week.</p><p><strong>That&#8217;s why the rest of this article focuses on the software stack. Today, it&#8217;s where companies are finding the biggest improvements in cost, latency, and throughput.</strong></p><div><hr></div><h4>Software Stack </h4><p>Inference isn&#8217;t one optimization problem anymore. It&#8217;s a stack of them.</p><p>Every layer exists because it&#8217;s solving a different bottleneck. Some make models portable. Others squeeze more work out of GPUs. Others manage deployments or optimize agent runtimes. Together, they determine the latency, throughput, and cost of every inference request.</p><p>Here&#8217;s how I think about the software stack, starting closest to the hardware and moving upward toward the application.</p><p><strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">Layer 1 </span></strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">&#8212;</span><strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);"> Model Runtime<br></span></strong><em>Purpose: Portable execution<br>Solves: Deployment fragmentation</em></p><p>This layer sits closest to the hardware. One of the messiest parts of AI today is that every piece of hardware wants your model in a different format. NVIDIA has one optimization stack. Apple has another. TPUs have another. CPUs have another. The model might stay the same, but the deployment pipeline changes every time.</p><p><strong><a href="https://muna.ai/">Muna</a></strong><br>Compiles Python inference code into optimized native executables, removing much of the runtime overhead and deployment complexity. Instead of shipping Python, containers, and a long chain of dependencies, you deploy a lightweight binary that starts faster and runs efficiently across different hardware.</p><p><strong><a href="https://www.modular.com/">Modular</a></strong><br>Takes a broader approach by rebuilding the runtime itself. Think of <strong>MAX</strong> as the engine that runs AI models and <strong>Mojo</strong> as the language developers use to build AI software. Together, they&#8217;re trying to hide hardware-specific complexity so the same AI software can execute efficiently across CPUs, GPUs, and other accelerators.</p><p><strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">Layer 2 </span></strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">&#8212;</span><strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);"> Model Serving Engine</span></strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);"><br></span><em>Purpose: Maximize GPU utilization<br>Solves: GPU underutilization</em></p><p>This is where inference economics are won or lost.</p><p>Running a model efficiently is much harder than just loading it onto a GPU. Requests arrive at different times, prompts have different lengths, outputs finish at different times, and GPUs end up spending more time waiting than doing useful work. The job of the serving layer is to keep that hardware as busy as possible.</p><p>That&#8217;s where techniques like <strong>continuous batching</strong>, <strong>disaggregated prefill</strong>, and <strong>custom kernels</strong> (e.g. paged attention) come in.</p><p>Instead of serving one request at a time, they continuously pack new requests onto the GPU, reuse memory more efficiently, and dramatically improve utilization. The model doesn&#8217;t change. The hardware doesn&#8217;t change. You simply get more tokens out of the same GPU, which means lower cost and higher throughput.</p><p><strong><a href="https://inferact.ai/">Inferact</a></strong><br>vLLM has become the default serving engine for production LLM deployments. If you&#8217;re serving chatbots, copilots, or general-purpose AI applications at scale, chances are you&#8217;re either using vLLM or evaluating it. Inferact is building the commercial layer around that ecosystem.</p><p><strong><a href="https://www.radixark.com/">RadixArk</a></strong><br>SGLang started with efficient structured generation but has quickly become a favorite for agentic and RL workloads, where models are reasoning, using tools, and running complex execution loops. RadixArk is commercializing that next generation of serving infrastructure.</p><p>Both companies were founded by many of the researchers behind the underlying open-source projects. It&#8217;s a familiar playbook: build the infrastructure in the open, then build the company around operating it at production scale.</p><blockquote><p><strong>Market signal:</strong> Both companies are betting that the serving layer becomes valuable in the same way Red Hat built a business around Linux - by commercializing critical open infrastructure rather than owning it.</p></blockquote><p><strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">Layer 3 </span></strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">&#8212;</span><strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);"> Deployment Platform<br></span></strong><em>Purpose: Run inference reliably in production<br>Solves: Production complexity.</em></p><p>Getting a model to run efficiently is only half the problem. You still have to deploy it, scale it, monitor it, and keep it running as traffic changes. That&#8217;s the job of the deployment layer. These companies sit one layer above the serving engine, handling GPU provisioning, autoscaling, deployments, routing, observability, and everything else needed to operate inference in production.</p><p><strong><a href="https://www.baseten.co/">Baseten</a></strong><br>Built for production inference. Models stay warm, serve traffic continuously, and come with the operational tooling you&#8217;d expect from a production service. If your model is powering a customer-facing application with real latency and uptime requirements, this is the type of platform you build on.</p><p><strong><a href="https://modal.com/">Modal</a></strong><br>Built for on-demand compute. Instead of keeping GPUs running 24/7, infrastructure spins up when your code runs and disappears when it&#8217;s done. It&#8217;s a much better fit for batch jobs, fine-tuning, parallel workloads, and agents that don&#8217;t need persistent endpoints.</p><p>The interesting part is that even though both abstract away GPU infrastructure, they just optimize for completely different workload patterns. Baseten assumes your model is always serving users. Modal assumes compute should exist only when there&#8217;s work to do.</p><blockquote><p><strong>Market signal:</strong> The line between deployment platforms and managed inference clouds is already starting to blur. Companies like Baseten, Fireworks, and Together are expanding vertically across the stack, increasingly bundling serving, deployment, and infrastructure into a single platform.</p></blockquote><p><strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">Layer 6 </span></strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">&#8212;</span><strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);"> Managed API layer</span><br></strong><em>Purpose: Consume optimized inference without owning the stack<br>Solves: Infrastructure ownership</em></p><p>Most companies don&#8217;t want to own the inference stack, but they also don&#8217;t want the limitations of a one-size-fits-all API.</p><p>Running models efficiently means making decisions about serving engines, deployments, routing, GPU provisioning, batching, autoscaling, and infrastructure. Companies like <strong>Fireworks</strong>, <strong>Together AI</strong>, and <strong>Doubleword</strong> take ownership of those decisions so developers don&#8217;t have to. You still interact with a familiar OpenAI-compatible API, but underneath the platform is constantly optimizing where your model runs, how requests are scheduled, how aggressively they&#8217;re batched, and how the underlying infrastructure is tuned for latency, throughput, and cost.</p><p><strong><a href="https://fireworks.ai/">Fireworks AI</a></strong><br>Built from the infrastructure up. Owns the serving layer, GPU infrastructure, and deployment stack to deliver highly optimized managed inference.</p><p><strong><a href="https://www.together.ai/">Together AI</a></strong><br>Started as the cloud for open-source models and has expanded into a full-stack inference platform spanning training, fine-tuning, deployment, and inference.</p><p><strong><a href="https://doubleword.ai/">Doubleword</a></strong><br>Takes a different approach. Instead of asking developers to think about infrastructure, it asks them to describe the workload- realtime, async, or batch- and figures out the most efficient way to execute it underneath.</p><p>The interesting part is that they&#8217;re all converging toward the same developer experience. Fireworks and Together started by abstracting infrastructure. Doubleword started by abstracting workload orchestration. Different implementations, increasingly similar abstraction.</p><blockquote><p><strong>Market signal:</strong> The bottom of the stack is fragmenting into specialized layers. The top of the stack is converging back into managed platforms that bundle those layers together behind a familiar API.</p></blockquote><p><strong><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">Layer 5 </span><span>&#8212;</span><span data-color="#4a86e8" style="color: rgb(74, 134, 232);">Agent-Native Inference<br></span></strong><em>Purpose: Optimize long-running agent workloads <br>Solves: Long-running execution</em></p><p>This is probably the newest layer in the inference stack.</p><p>Everything we&#8217;ve talked about so far assumes inference is the thing being optimized. But long-running agents expose another bottleneck. They don&#8217;t just make model calls&#8212;they execute code, call tools, wait for inference, resume execution, and repeat that loop for hours. Optimizing inference alone leaves a lot of efficiency on the table.</p><p>That&#8217;s why I think we&#8217;re starting to see a new category emerge.</p><p><strong><a href="https://www.sailresearch.com/">Sail Research</a></strong></p><p>I first came across Sail Research when they were positioning themselves as another inference provider. The pitch was familiar: cheaper inference, better utilization, infrastructure built for AI agents.</p><p>Their recent product launches made me realize they&#8217;re becoming something different.</p><p>Instead of optimizing inference in isolation, Sail couples inference and execution together. <strong>Sailboxes</strong> are persistent execution sandbox environments where agents actually run. When an agent blocks on inference, the sandbox automatically pauses and resumes when the response arrives. You&#8217;re no longer paying for allocated compute while the agent is simply waiting.</p><p>They&#8217;ve since added <strong>Voyages</strong>, an observability layer for monitoring long-running agent workflows. Together, inference, execution, and telemetry become one system.</p><p>I think that&#8217;s what makes this layer interesting. It&#8217;s no longer just about serving models faster. It&#8217;s about optimizing the entire lifecycle of an agent.</p><blockquote><p><strong>Market signal:</strong> Agent runtimes are still an emerging category, but they&#8217;re one of the first infrastructure layers built specifically for long-running AI agents rather than conversational AI</p></blockquote><div><hr></div><p>That&#8217;s my attempt at mapping where inference is headed.</p><p>It took me way longer than I expected because the market is evolving incredibly fast. New companies seem to pop up every week, the boundaries between categories are already starting to blur, and everyone is trying to solve a slightly different piece of the same problem.</p><p>The interesting part isn&#8217;t any one company.</p><p>It&#8217;s that inference has quietly become its own software industry.</p><p>I think we&#8217;re still very early.</p><div><hr></div><p><em>Further reading &amp; acknowledgements: I'd highly recommend reading <a href="https://www.baseten.co/inference-engineering/">Inference Engineering </a>by Philip Kiely and the Baseten team; </em><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Paolo Perrone&quot;,&quot;id&quot;:12567301,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4b3dca43-c984-41cc-a31a-b210ec08bddb_1024x1024.jpeg&quot;,&quot;uuid&quot;:&quot;64e8320b-5248-4340-962c-482642b97006&quot;}" data-component-name="MentionToDOM"></span> <em>'s <a href="https://theaiengineer.substack.com/p/vllm-vs-ollama-vs-sglang-vs-tensorrt">comparison of vLLM, Ollama, SGLang, and TensorRT-LLM</a> on The AI Engineer; <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Chris Zeoli&quot;,&quot;id&quot;:5502193,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd31796b-f22a-4af8-83fe-b10d81a17736_837x837.png&quot;,&quot;uuid&quot;:&quot;9d184f70-a239-4e9d-afa8-c9a6b03a2efe&quot;}" data-component-name="MentionToDOM"></span>&#8217;s and the team at <a href="https://www.wing.vc/content/the-ai-inference-stack">Wing VC</a>,  for some of the most thoughtful writing on inference I've read. Thanks also to my friend <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Yusuf Olokoba&quot;,&quot;id&quot;:42423943,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/0a4f22c5-f9e2-4bc1-91cd-fea771410624_3024x3024.jpeg&quot;,&quot;uuid&quot;:&quot;fa867a2b-8a7a-43f1-af4c-f878c0ab0c9f&quot;}" data-component-name="MentionToDOM"></span>, founder of <a href="https://www.muna.ai/">Muna</a>, for the many conversations that helped shape this framework, especially around how the different layers fit together.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.theretrofit.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Governing AI Agents Is Harder Than You Think. Here's The Map]]></title><description><![CDATA[Documentation, sandboxes, and system prompts aren't AI governance. Here's what is.]]></description><link>https://www.theretrofit.ai/p/governing-ai-agents-is-harder-than</link><guid isPermaLink="false">https://www.theretrofit.ai/p/governing-ai-agents-is-harder-than</guid><pubDate>Mon, 29 Jun 2026 02:07:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BQIg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most enterprise AI deployments today aren&#8217;t fully autonomous. There&#8217;s usually a human in the loop. The agent proposes a plan, chooses which tools to call, and surfaces its reasoning or actions for approval before executing.</p><p>That works today because the number of decisions is still manageable. But as agents become more capable, asking a human to validate every tool call, every API request, and every execution-time decision quickly stops being practical. At some point, the human becomes the bottleneck, and you&#8217;ve lost the benefit of having an autonomous system in the first place.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.theretrofit.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>That&#8217;s where I think the real governance problem starts.</p><p><a href="https://www.gartner.com/en/newsroom/press-releases/2026-05-26-gartner-says-applying-uniform-governance-across-ai-agents-will-lead-to-enterprise-ai-agent-failure">Gartner recently predicted that by 2027</a>, 40% of enterprises will demote or decommission autonomous AI agents because governance gaps won&#8217;t become apparent until after production incidents. I don&#8217;t think we have to wait for full autonomy to run into those problems.</p><div><hr></div><h4>If you&#8217;ve used Claude Code, you&#8217;ve already experienced the problem</h4><p>You just probably haven&#8217;t thought about it as an <strong>AI agent governance problem</strong>.</p><p>You give Claude Code a task, and a minute later it&#8217;s read half your codebase, called a handful of tools, maybe spun up a sub-agent, edited a file, hit an API, and somehow found a config file you forgot existed. Most of that happened before you had a chance to follow along.</p><p>That&#8217;s not because Claude Code is doing anything wrong. It&#8217;s because that&#8217;s how capable AI agents work. They gather context, adapt their plan, choose which tools to call, and make decisions based on what they learn while they&#8217;re executing.</p><p>What struck me is that almost all of those decisions happened after I gave the agent the task.</p><p>That&#8217;s the governance problem.</p><div><hr></div><h4>Today, most AI agent governance is front-loaded</h4><p>Most teams still focus on governing what an AI agent starts with, even though the most important decisions happen after execution begins. That model works well for deterministic software because the execution path is largely known in advance. AI agents are different. They gather new context, adapt their plan, and continuously make decisions while they're running.</p><p>The natural instinct is to apply least privilege, just as we do for humans. Give an agent the minimum access it needs. Scope the tools it can use. Issue short-lived credentials. That&#8217;s still good security, and it isn&#8217;t going away.</p><p>But least privilege isn&#8217;t the whole story anymore. It tells you what an agent <em>could</em> access, not whether a specific action should happen in the context of the task it&#8217;s executing.</p><p>An agent might have permission to GitHub, Jira, and Salesforce. The harder question is whether it should update <em>this</em> issue, modify <em>this</em> repository, or read <em>this</em> customer record while acting on behalf of <em>this</em> user.</p><p><strong>That&#8217;s a runtime decision.</strong></p><p>Right now, most teams think about AI agent governance in three layers.</p><p><strong>Tier 1: Static governance.</strong> Acceptable use policies. Approved model lists. Compliance documents legal signed off on. These define what agents are <em>supposed</em> to do. They say nothing about what agents actually do.</p><p><strong>Tier 2: Structural governance.</strong> Tool allowlists. MCP connectors. Sandboxed APIs. Better than documentation because you&#8217;ve narrowed the blast radius, but it&#8217;s still a configuration decision made before the agent starts.</p><p><strong>Tier 3: Instructional governance.</strong> <code>CLAUDE.md</code>. <code>AGENTS.md</code>. <code>.cursor/rules</code>. System prompts. The agent harness. &#8220;Never access production.&#8221; &#8220;Always ask before deleting.&#8221; This is where most sophisticated teams are today.</p><p>The problem is that all three layers stop before execution.</p><p>Once the agent starts reasoning, choosing tools, and interacting with systems, the first three layers have already done their job. They shape the agent&#8217;s behavior, but they don&#8217;t enforce it.</p><p>That&#8217;s why I think there&#8217;s a fourth layer.</p><p><strong>Tier 4: Runtime governance.</strong> Instead of evaluating policy once at setup, policy is evaluated at the moment of every tool call. Before the agent touches an external system, the runtime has the context of the task, the identity of the user the agent is acting for, the requested action, and the applicable policies. It can make a deterministic decision: allow or deny.</p><p><strong>That's the point where governance becomes enforcement instead of guidance.</strong></p><p></p><p><strong>                                       Before the agent touches anything</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BQIg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BQIg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png 424w, https://substackcdn.com/image/fetch/$s_!BQIg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png 848w, https://substackcdn.com/image/fetch/$s_!BQIg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png 1272w, https://substackcdn.com/image/fetch/$s_!BQIg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BQIg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png" width="1312" height="868" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:868,&quot;width&quot;:1312,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115596,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://theretrofit.substack.com/i/202975598?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BQIg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png 424w, https://substackcdn.com/image/fetch/$s_!BQIg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png 848w, https://substackcdn.com/image/fetch/$s_!BQIg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png 1272w, https://substackcdn.com/image/fetch/$s_!BQIg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78fdfb8b-ef03-4706-a85f-dfe6400ec376_1312x868.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h4>What runtime governance of AI agents actually looks like</h4><p>The easiest way I&#8217;ve found to think about it is to start with how we already govern people, then ask what changes when the actor is an AI agent.</p><p>The same principles still apply: identity, authorization, and access. The difference is that those decisions now have to account for who delegated the task, what the agent is doing, and what it&#8217;s trying to access at that moment.</p><h5><strong>Identity - Who is this agent?</strong></h5><p>Agents don&#8217;t look like users. They spin up, perform a task, call a few tools, and disappear. Runtime governance starts by giving every agent a real identity that&#8217;s tied to who delegated the task, what it&#8217;s doing, and what scope it was given. Without that, everything that follows becomes difficult to attribute.</p><h5><strong>Runtime authorization - Should this action be allowed?</strong></h5><p>The second capability is runtime authorization.</p><p>Every tool call becomes a policy decision. Before an agent touches an external system, something has to decide whether that specific action should be allowed for this user, through this agent, on this resource, in the context of the task it&#8217;s currently executing.</p><p>What&#8217;s interesting is that the industry seems to have independently converged on where that decision belongs. Claude Code, Codex, Gemini CLI, and Microsoft&#8217;s Agent Governance Toolkit all expose hooks that let developers inspect or intervene while an agent is running.</p><p>The enforcement point already exists.</p><p>The challenge is making those decisions with enough context to determine whether the action is actually authorized.</p><h5><strong>Access - How does the agent get the permissions it needs?</strong></h5><p>The final capability is access.</p><p>Long-lived API keys, service accounts, and <code>.env</code> files don&#8217;t fit well with autonomous agents. Instead, access should be granted just in time, scoped to the specific action being performed, and expire as soon as that work is complete.</p><p><strong>Taken together, that&#8217;s what runtime governance looks like: identity that can attribute every action, authorization that evaluates every tool call in context, and access that&#8217;s granted only for the work the agent is performing.</strong></p><div><hr></div><h4>The Emerging AI Agent Governance Stack</h4><p>One thing that stood out to me while researching this space is that nobody is building <strong>AI agent governance</strong> as a single product.</p><p>Instead, the market is emerging one layer at a time. Companies are building different parts of what increasingly looks like an <strong>AI agent runtime governance stack</strong>: <strong>agent identity</strong>, <strong>runtime policy evaluation</strong>, <strong>access &amp; credentials</strong>, and <strong>audit</strong>. Once you look at the landscape through that lens, it becomes much easier to understand.</p><p>What&#8217;s also interesting is that the market is splitting into two camps. Some companies are building <strong>agent-native infrastructure</strong> from scratch. Others are extending the <strong>identity and security infrastructure</strong> enterprises already rely on.</p><h5><strong>Agent Identity</strong></h5><p>Identity is where I think the market gets interesting. There seem to be two approaches emerging.</p><p>The incumbents already own human identity. Okta, Microsoft Entra, Auth0, Ping Identity, and others aren't going away. If enterprises are going to deploy millions of AI agents, it makes sense that those platforms will eventually need to understand agents alongside people. The question is whether agent identity becomes another object inside existing IAM systems, or whether it requires an entirely new identity layer.</p><p><a href="https://www.keycard.ai/">Keycard</a> is the clearest example. Rather than treating an agent as another service account, it gives agents first-class identities tied to the human who delegated the task, the workload they&#8217;re executing, and the scope they were granted. It federates with existing IAM systems instead of replacing them, then adds delegation, workload attestation, and short-lived credentials on top.</p><p>There&#8217;s also an adjacent category that&#8217;s becoming increasingly important: <strong>non-human identity (NHI) governance</strong>.</p><p>Companies like <a href="https://www.oasis.security/">Oasis Security</a> aren&#8217;t creating identities for new AI agents. They&#8217;re helping enterprises discover and govern the machine identities they already have-service accounts, OAuth applications, API keys, and other credentials that autonomous agents increasingly inherit. As AI agents become another consumer of machine identities, NHI governance starts looking like a foundational piece of the runtime stack.</p><h5><strong>Runtime Policy</strong></h5><p>Once an agent has an identity, every tool call becomes a policy decision.</p><p>This is where companies like <a href="https://www.cerbos.dev/">Cerbos</a> and <a href="https://www.osohq.com/">Oso</a> fit.</p><p><strong>Cerbos</strong> approaches the problem as an authorization policy engine. Rather than embedding permission logic throughout application code, it evaluates every authorization request against declarative policies at runtime. For AI agents, that means every tool call can be evaluated independently before it's allowed to execute.</p><p><strong>Oso</strong> approaches the problem from the application authorization side. It helps developers define fine-grained authorization based on roles, relationships, and context, making it possible to evaluate what an AI agent should be allowed to do on behalf of a user instead of simply inheriting all of that user's permissions.</p><p>Microsoft&#8217;s open-source Agent Governance Toolkit is another interesting signal. Rather than inventing a new runtime, Microsoft is extending existing agent frameworks with runtime policy enforcement. To me, that&#8217;s a sign they see <strong>runtime policy evaluation</strong> becoming foundational infrastructure rather than a product differentiator.</p><h5><strong>Access &amp; Credentials</strong></h5><p>Identity answers <strong>who</strong>.</p><p>Policy answers <strong>whether</strong>.</p><p>Something still has to answer <strong>how an AI agent obtains access</strong>.</p><p>That&#8217;s where companies like <a href="https://arcade.dev/">Arcade</a> and <a href="https://nango.dev/">Nango</a> fit, although they&#8217;re solving different parts of the problem.</p><p><strong>Arcade</strong> focuses on delegated, just-in-time access for AI agents. Instead of giving an agent a long-lived API key or service account, it obtains a short-lived credential on behalf of the user, scopes it to the specific action, and expires it when the work is complete.</p><p><strong>Nango</strong> approaches the problem from the OAuth and API integration layer. It provides infrastructure for OAuth flows, token refresh, and credential management across hundreds of third-party APIs while allowing organizations to retain control of their own credential infrastructure.</p><p><strong>Both are moving away from standing credentials toward credentials that exist only for the duration of the work an AI agent is performing.</strong></p><h5><strong>Audit</strong></h5><p>Audit is less of a standalone product category than an outcome of the other three working together.</p><p>If identity, policy, and credentials are all evaluated at runtime, an audit trail naturally falls out of the system. You know who delegated the task, which policy evaluated the request, what credential was issued, and every action the agent took.</p><p>Companies like Keycard, Cerbos, and Arcade all contribute different parts of that picture - identity, policy decisions, and execution events, but together they make it possible to reconstruct what happened after the fact.</p><p>That&#8217;s the interesting part to me.</p><p>These companies aren&#8217;t building competing products. They&#8217;re building different layers of what looks increasingly like the runtime stack for AI agents.</p><div><hr></div><h4>Executive Order 14409 Signals Where AI Agent Governance Is Headed</h4><p>On June 2, 2026, President Trump signed <a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/?utm_source=chatgpt.com">Executive Order 14409</a>. Most of the attention went to <a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/?utm_source=chatgpt.com#:~:text=Sec.%203,including%20frontier%20models.">Section 3</a>, which targets frontier model developers.</p><p>I think<a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/?utm_source=chatgpt.com#:~:text=Sec.%204,or%20unlawful%20purpose."> Section 4</a> is the part enterprises should be paying attention to.</p><p>It directs the Attorney General to prioritize enforcement of existing laws, including the Computer Fraud and Abuse Act (CFAA), against people who use AI to gain unauthorized access to computer systems. It explicitly calls out AI agents.</p><p>The Executive Order doesn&#8217;t create a new legal standard for enterprises.</p><p>What it does signal is that authorization is becoming the question.</p><p>Not:</p><blockquote><p>Did we tell the agent what to do?</p></blockquote><p>But:</p><blockquote><p>Can we prove this action was authorized?</p></blockquote><p></p><p><strong>To me, that's where AI agent governance is headed. Not more prompts or policy documents, but runtime identity, authorization, access control, and auditability that can answer that question every time an AI agent takes an action.</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.theretrofit.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>