<?xml version="1.0" encoding="utf-8"?><?xml-stylesheet type="text/xsl" href="rss.xsl"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>southpawriter Blog</title>
        <link>https://southpawriter.com/blog</link>
        <description>southpawriter Blog</description>
        <lastBuildDate>Tue, 17 Mar 2026 00:00:00 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <item>
            <title><![CDATA[The Quiet Part]]></title>
            <link>https://southpawriter.com/blog/the-quiet-part</link>
            <guid>https://southpawriter.com/blog/the-quiet-part</guid>
            <pubDate>Tue, 17 Mar 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Three weeks of silence, no dramatic exit, no pivot announcement. Just a brain that needed to stop sprinting. Here's what happened and why I'm not apologizing for it.]]></description>
            <content:encoded><![CDATA[<p>You may have noticed it got quiet around here.</p>
<p>No farewell post. No "exciting announcement." No carefully worded "I'm pivoting to..." thread with a blue-sky emoji. Just... silence. Three weeks of it, which in blog-time is roughly equivalent to leaving a shopping cart in the middle of the grocery aisle and walking out of the store.</p>
<p>I owe you an explanation. Or, more accurately, I don't owe you anything, but I'm going to give you one anyway because I spent the last three weeks staring at my projects and feeling <em>nothing</em>, and it turns out that's worth talking about.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-wall">The Wall<a href="https://southpawriter.com/blog/the-quiet-part#the-wall" class="hash-link" aria-label="Direct link to The Wall" title="Direct link to The Wall" translate="no">​</a></h2>
<p>Burnout doesn't arrive like a deadline. It doesn't send calendar invites. It shows up like a slow leak in a tire; you don't notice it until you're standing in a parking lot wondering why everything feels slightly wrong and harder than it should be.</p>
<p>Here's what mine looked like: I'd open Lexichord. Stare at the code. Close the tab. Open FractalRecall. Read three lines of my own design spec. Close that tab too. Check my email. Make coffee. Open Lexichord again. Stare harder this time, as if intensity might substitute for interest. It didn't.</p>
<p>This went on for days before I admitted what was happening.</p>
<p>I wasn't stuck on a technical problem. I wasn't blocked by a dependency or a design flaw or a missing NuGet package. The projects were fine. <em>I</em> was the part that wasn't working.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-burnout-isnt">What Burnout Isn't<a href="https://southpawriter.com/blog/the-quiet-part#what-burnout-isnt" class="hash-link" aria-label="Direct link to What Burnout Isn't" title="Direct link to What Burnout Isn't" translate="no">​</a></h2>
<p>I want to be precise about this because the internet has turned "burnout" into a catch-all for everything from genuine exhaustion to "I didn't feel like doing stuff this week." Those aren't the same thing. Laziness is a choice. Burnout is what happens when you've been making the opposite choice for too long.</p>
<p>Fourteen blog posts in thirteen days. Multiple research projects running in parallel. A benchmark methodology. A .NET library. An AI orchestration tool with a growing test corpus. A dungeon crawler because apparently my brain's idea of "relaxation" is modeling combat state machines.</p>
<p>None of that was forced on me. I loved all of it. That's the part nobody warns you about; the burnout you have to watch for isn't the kind that comes from doing work you hate. It's the kind that comes from doing work you love at a pace your brain can't sustain.</p>
<p>Sprinting is fine. Sprinting <em>indefinitely</em> is not a strategy.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-permission-problem">The Permission Problem<a href="https://southpawriter.com/blog/the-quiet-part#the-permission-problem" class="hash-link" aria-label="Direct link to The Permission Problem" title="Direct link to The Permission Problem" translate="no">​</a></h2>
<p>There's this specific flavor of guilt that technical people carry around like a loyalty card they forgot to throw away. The project is interesting. The work is meaningful. Nobody's making you do it. So if you stop, what does that say about you?</p>
<p>It says you're a person with a nervous system. That's it. That's the whole answer.</p>
<p>But knowing that intellectually and <em>feeling</em> it are two different things. I spent the first week of this break doing the worst possible version of resting: not working on projects, but also not <em>not</em> thinking about them. Sitting on the couch with my laptop open to a file I refused to read, accomplishing the impressive feat of neither relaxing nor being productive. A perfectly balanced failure state.</p>
<p>Somewhere around week two, the guilt machinery finally ran out of fuel. I stopped opening the laptop. I played games that someone <em>else</em> designed. I read books that had nothing to do with embeddings or metadata or retrieval-augmented anything. I went outside, which I can report is still there and largely unchanged.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-i-learned-sort-of">What I Learned (Sort Of)<a href="https://southpawriter.com/blog/the-quiet-part#what-i-learned-sort-of" class="hash-link" aria-label="Direct link to What I Learned (Sort Of)" title="Direct link to What I Learned (Sort Of)" translate="no">​</a></h2>
<p>I'm not going to package this up with a bow. I don't have a five-step framework for recovering from burnout, and if someone tries to sell you one, that person is selling you a content calendar disguised as therapy.</p>
<p>But I noticed some things.</p>
<p>The projects didn't collapse. Lexichord is right where I left it. FractalRecall's design spec didn't rewrite itself out of spite. The llms.txt research is still sitting in my evidence inventory, patient as ever, waiting for me to come back and ask it more questions I won't like the answers to.</p>
<p>Nothing caught fire. The world continued to not depend on my publishing schedule. This is both humbling and, in a quieter way, freeing.</p>
<p>I also noticed that the docs-first compulsion, the one that drives me to document everything before building it, applies here too. I couldn't just <em>take a break</em>. I had to understand why I needed one, what the failure mode was, and what the recovery criteria looked like. Yes, I am aware that analyzing your own burnout with the same rigor you'd apply to a design spec is itself a symptom. I contain multitudes. (And a changelog.)</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-happens-now">What Happens Now<a href="https://southpawriter.com/blog/the-quiet-part#what-happens-now" class="hash-link" aria-label="Direct link to What Happens Now" title="Direct link to What Happens Now" translate="no">​</a></h2>
<p>I'm back. Provisionally. Not at the pace I was running before, because that pace was the problem.</p>
<p>The projects haven't changed. The work is still interesting. What <em>has</em> changed is that I've given myself permission to not publish fourteen posts in two weeks and call that "normal." It was never normal. It was adrenaline masquerading as a workflow.</p>
<p>I'll be writing here again. Probably not daily. The llms.txt research still has findings worth publishing. Lexichord still needs its test corpus finished. FractalRecall still has an unresolved question about prefix artifact contamination that I think about in the shower. (This is not a metaphor. I literally think about embedding prefix artifacts in the shower. I may need additional hobbies.)</p>
<p>But the pace is going to be different, and I'm not going to apologize for that. Three weeks of quiet isn't a failure. It's what happens when someone who writes documentation about their documentation process finally admits that the documentation subject this time is themselves.</p>
<p>I'm still here. I'm still building.</p>
<p>I just needed to stop sprinting long enough to remember why I started running.</p>]]></content:encoded>
            <category>Opinion</category>
            <category>Career</category>
        </item>
        <item>
            <title><![CDATA[Embedding Models Don't Read Your Metadata (But They Should)]]></title>
            <link>https://southpawriter.com/blog/embeddings-dont-read-metadata</link>
            <guid>https://southpawriter.com/blog/embeddings-dont-read-metadata</guid>
            <pubDate>Thu, 26 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Your embedding model ignores YAML metadata. Say the same thing in English and retrieval ranking quality jumps 16.5%. FractalRecall exploits that gap.]]></description>
            <content:encoded><![CDATA[
<p>Here's a sentence your <a class="" href="https://southpawriter.com/docs/glossary/library/embedding">embedding</a> model understands perfectly well:</p>
<blockquote>
<p><em>"This is a canonical faction document from the post-Glitch era describing cultural practices and political structure."</em></p>
</blockquote>
<p>And here's functionally identical information that your embedding model treats as random noise:</p>
<div class="language-yaml codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:#2a2a2a;--prism-color:#f92aad"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-yaml codeBlock_bY9V thin-scrollbar" style="background-color:#2a2a2a;background-image:#34294f;color:#f92aad;text-shadow:0 0 2px #100c0f, 0 0 5px #dc078e33, 0 0 10px #fff3"><code class="codeBlockLines_e6Vv"><span class="token-line" style="background-image:#34294f;color:#f92aad;text-shadow:0 0 2px #100c0f, 0 0 5px #dc078e33, 0 0 10px #fff3"><span class="token key atrule" style="color:#f4eee4;text-shadow:0 0 2px #393a33, 0 0 8px #f39f0575, 0 0 2px #f39f0575">canon</span><span class="token punctuation" style="color:#ccc">:</span><span class="token plain"> </span><span class="token boolean important" style="color:#f4eee4;text-shadow:0 0 2px #393a33, 0 0 8px #f39f0575, 0 0 2px #f39f0575;font-weight:bold">true</span><span class="token plain"></span><br></span><span class="token-line" style="background-image:#34294f;color:#f92aad;text-shadow:0 0 2px #100c0f, 0 0 5px #dc078e33, 0 0 10px #fff3"><span class="token plain"></span><span class="token key atrule" style="color:#f4eee4;text-shadow:0 0 2px #393a33, 0 0 8px #f39f0575, 0 0 2px #f39f0575">domain</span><span class="token punctuation" style="color:#ccc">:</span><span class="token plain"> faction</span><br></span><span class="token-line" style="background-image:#34294f;color:#f92aad;text-shadow:0 0 2px #100c0f, 0 0 5px #dc078e33, 0 0 10px #fff3"><span class="token plain"></span><span class="token key atrule" style="color:#f4eee4;text-shadow:0 0 2px #393a33, 0 0 8px #f39f0575, 0 0 2px #f39f0575">era</span><span class="token punctuation" style="color:#ccc">:</span><span class="token plain"> post</span><span class="token punctuation" style="color:#ccc">-</span><span class="token plain">glitch</span><br></span><span class="token-line" style="background-image:#34294f;color:#f92aad;text-shadow:0 0 2px #100c0f, 0 0 5px #dc078e33, 0 0 10px #fff3"><span class="token plain"></span><span class="token key atrule" style="color:#f4eee4;text-shadow:0 0 2px #393a33, 0 0 8px #f39f0575, 0 0 2px #f39f0575">topics</span><span class="token punctuation" style="color:#ccc">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#ccc">[</span><span class="token plain">culture</span><span class="token punctuation" style="color:#ccc">,</span><span class="token plain"> politics</span><span class="token punctuation" style="color:#ccc">]</span><br></span></code></pre></div></div>
<p>Same facts. Same document. Different embedding behavior. The YAML blob gets processed as four disconnected <a class="" href="https://southpawriter.com/docs/glossary/library/token">tokens</a> with no semantic weight. The natural language sentence gets encoded as a rich set of contextual signals that tell the model what this document <em>is</em>, what it's <em>about</em>, and how it relates to the kind of questions someone might ask.</p>
<p>The gap between the metadata your system knows and the context your embeddings encode is the single biggest free improvement sitting in most <a class="" href="https://southpawriter.com/docs/glossary/library/retrieval-augmented-generation">RAG</a> pipelines. Almost nobody exploits it.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-problem-with-metadata-in-modern-rag">The Problem With Metadata in Modern RAG<a href="https://southpawriter.com/blog/embeddings-dont-read-metadata#the-problem-with-metadata-in-modern-rag" class="hash-link" aria-label="Direct link to The Problem With Metadata in Modern RAG" title="Direct link to The Problem With Metadata in Modern RAG" translate="no">​</a></h2>
<p>Every serious retrieval system has metadata. Document type, creation date, author, category, topic tags, authority level--the structural information that tells you what a document <em>is</em> as opposed to what it <em>says</em>. This metadata is valuable. Everyone agrees it's valuable. It powers filters, facets, access control, and sorting.</p>
<p>But here's the thing: your embedding model never sees it.</p>
<p>When you chunk a document, embed it, and store it in a <a class="" href="https://southpawriter.com/docs/glossary/library/vector-store">vector store</a>, the embedding is computed from the <em>text content</em> of the chunk. The metadata sits in a separate field--available for post-retrieval filtering, but invisible to the semantic similarity computation that determines what gets retrieved in the first place.</p>
<p>This means your retrieval system has two disconnected brains:</p>
<ol>
<li class=""><strong>The vector search brain</strong> knows what the document <em>says</em> but not what it <em>is.</em></li>
<li class=""><strong>The metadata filter brain</strong> knows what the document <em>is</em>. It only gets consulted after the vector search has already decided what's relevant.</li>
</ol>
<p>The vector brain retrieves candidates based on semantic similarity. The filter brain then removes candidates that don't match the metadata criteria. This is fine for simple cases--"find documents about bears, but only from the fauna category." The vector search finds bear-related content, the filter removes anything that's not fauna.</p>
<p>But it fails for anything nuanced, because the filter can only <em>remove</em> candidates. It can't <em>boost</em> them. It can't say "this document is not just about bears, it's a <em>canonical bestiary entry</em> about bears, which is exactly what a question about bear taxonomy is looking for." That kind of reasoning requires the metadata to be inside the embedding--part of the semantic representation itself, not an afterthought applied post-retrieval.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-fractalrecall-actually-does">What FractalRecall Actually Does<a href="https://southpawriter.com/blog/embeddings-dont-read-metadata#what-fractalrecall-actually-does" class="hash-link" aria-label="Direct link to What FractalRecall Actually Does" title="Direct link to What FractalRecall Actually Does" translate="no">​</a></h2>
<p>FractalRecall's approach is simple: take the metadata, convert it to a natural language sentence, and prepend it to the chunk before embedding.</p>
<p>That's it. That's the intervention.</p>
<p>A chunk that previously looked like this to the embedding model:</p>
<blockquote>
<p><em>"The Rune-Bear's patrol cycle operates on a 22.4-hour rhythm, following territorial boundaries marked by degraded authentication beacons. When the beacons fire, the creature pauses, waits for a response that will never come, and resumes its circuit."</em></p>
</blockquote>
<p>Now looks like this:</p>
<blockquote>
<p><em>"[Canonical bestiary entry from the post-Glitch era, covering URSA-class autonomous fauna in the Asgard-Midgard border region. Authority: confirmed field observation, verified by Scriptorium-Primus.] The Rune-Bear's patrol cycle operates on a 22.4-hour rhythm, following territorial boundaries marked by degraded authentication beacons. When the beacons fire, the creature pauses, waits for a response that will never come, and resumes its circuit."</em></p>
</blockquote>
<p>Same content. But the embedding model now encodes not just <em>what the text says</em> but <em>what kind of document it comes from, how authoritative it is, and what domain it belongs to.</em></p>
<p>When someone queries "What are the most dangerous autonomous creatures in the Asgard border region?", the enriched chunk is a better semantic match--not because the text content changed, but because the metadata prefix tells the embedding model that this chunk is about autonomous fauna in the Asgard region, which is exactly what the query is asking about.</p>
<p>The embedding model can't read YAML. But it can read English. And it turns out that's all you need.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-numbers-briefly">The Numbers (Briefly)<a href="https://southpawriter.com/blog/embeddings-dont-read-metadata#the-numbers-briefly" class="hash-link" aria-label="Direct link to The Numbers (Briefly)" title="Direct link to The Numbers (Briefly)" translate="no">​</a></h2>
<p>The D-22 experiment tested this approach with a single metadata layer (domain classification). <a class="" href="https://southpawriter.com/docs/glossary/library/ndcg">NDCG@10</a> (ranking quality) improved 16.5%. <a class="" href="https://southpawriter.com/docs/glossary/library/recall">Recall</a> (the fraction of relevant documents actually found) jumped 27.3%. I've written about these numbers <a class="" href="https://southpawriter.com/blog/43-percent-disappeared">in detail</a>, including the uncomfortable part where 43% of chunks silently overflowed the token limit and the metrics improved anyway.</p>
<p>The core thesis (metadata as co-embedded context) produces strong results. The engineering challenge of doing it without exceeding token limits, without creating prefix artifacts, and without drowning the actual content in structural boilerplate is where the real work lives. That's why this is still a research project.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-this-works-the-intuition">Why This Works (The Intuition)<a href="https://southpawriter.com/blog/embeddings-dont-read-metadata#why-this-works-the-intuition" class="hash-link" aria-label="Direct link to Why This Works (The Intuition)" title="Direct link to Why This Works (The Intuition)" translate="no">​</a></h2>
<p>Embedding models are trained on natural language. They have strong priors about what sentences mean, how topics relate, and what kind of context modifies what kind of content. When you give them a sentence like "Canonical bestiary entry from the post-Glitch era," they encode:</p>
<ul>
<li class=""><strong>Canonical</strong> → authoritative, primary source, verified</li>
<li class=""><strong>Bestiary entry</strong> → creature description, fauna, biological</li>
<li class=""><strong>Post-Glitch era</strong> → temporal context, specific period</li>
</ul>
<p>These encoded signals create semantic bridges between the chunk and queries that use related language. A query about "authoritative sources on creatures" now has a shorter vector distance to this chunk--not because the creature description matches, but because the <em>prefix</em> matches.</p>
<p>This is not magic. It's leveraging something the model already knows how to do (process natural language context) and giving it context it didn't previously have access to.</p>
<p>It's the spine label on a library book. Without it, a librarian has to open every book and read every page to find what a patron needs. The metadata prefix doesn't change what's inside. It tells you what kind of book it is <em>before</em> you open it.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-this-means-for-production-systems">What This Means for Production Systems<a href="https://southpawriter.com/blog/embeddings-dont-read-metadata#what-this-means-for-production-systems" class="hash-link" aria-label="Direct link to What This Means for Production Systems" title="Direct link to What This Means for Production Systems" translate="no">​</a></h2>
<p>If you're running a RAG pipeline in production, what does your embedding model actually know about your documents?</p>
<p>It knows the text content of each chunk. That's it. It has no concept of:</p>
<ul>
<li class=""><strong>Document type.</strong> Is this a policy document? A product manual? A customer email?</li>
<li class=""><strong>Authority level.</strong> Is this the definitive source, or a draft?</li>
<li class=""><strong>Temporal context.</strong> Current or historical? Does it supersede something?</li>
<li class="">Where does this chunk fall in the document hierarchy? Section header, conclusion, appendix--the embedding has no idea.</li>
</ul>
<p>All of this information exists in your system. Some of it lives in explicit metadata fields. Some of it could be inferred from the document structure. All of it is invisible to your embeddings.</p>
<p>FractalRecall's thesis is that making this information visible is the highest-leverage improvement most RAG systems aren't making. Convert it to natural language, include it in the embedding input, and the model can finally use what your system already knows.</p>
<p>The caveats are real:</p>
<ul>
<li class=""><strong>Token limits matter.</strong> A 200-word prefix on a 300-word chunk means you're spending 40% of your embedding capacity on metadata rather than content. This is <a class="" href="https://southpawriter.com/blog/43-percent-disappeared">the overflow problem D-22 discovered the hard way</a>.</li>
<li class=""><strong>Prefix quality matters.</strong> "This is a document" adds nothing. "Canonical policy document approved by Legal, superseding version 3.2, covering employee termination procedures for remote workers" adds a lot. The prefix needs to be specific, concise, and genuinely informative.</li>
<li class=""><strong>Schema design matters.</strong> You can't encode metadata you don't have. If your documents aren't tagged with domain, authority, temporal context, and structural position, you need to build that taxonomy first. The enrichment is only as good as the metadata it draws from.</li>
<li class=""><strong>Evaluation matters.</strong> (See: <a class="" href="https://southpawriter.com/blog/check-engine-light">Your RAG Pipeline Has a Check Engine Light</a>.) You need to measure whether the enrichment actually improves retrieval for <em>your</em> queries against <em>your</em> corpus. My 16.5% on Aethelgard is not your 16.5% on your data.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-bigger-claim">The Bigger Claim<a href="https://southpawriter.com/blog/embeddings-dont-read-metadata#the-bigger-claim" class="hash-link" aria-label="Direct link to The Bigger Claim" title="Direct link to The Bigger Claim" translate="no">​</a></h2>
<p><strong>The future of retrieval is not better embedding models. It's better embedding inputs.</strong></p>
<p>The models are already very good at processing natural language. They understand topics, relationships, authority, temporality, and context--when you give them that information in a form they can process. The bottleneck isn't the model's capability. It's the information we choose to give it.</p>
<p>Every document in your system has structural context that the embedding model can't access. Making that context accessible is a simple intervention with outsized impact. It doesn't require a new model, a new architecture, or a new infrastructure. It requires a string concatenation and a willingness to rethink what "the document" means when you hand it to an embedding function.</p>
<p>The metadata is already there. The model is already capable of understanding it.</p>
<p>The only thing missing is the sentence.</p>
<hr>
<p><em>FractalRecall is an active research project exploring metadata as co-embedded context for retrieval. For experiment details, see the <a class="" href="https://southpawriter.com/projects/fractalrecall">project page</a>. For the story of what happens when enrichment goes wrong, see <a class="" href="https://southpawriter.com/blog/43-percent-disappeared">I Added Context to My Embeddings and 43% of My Data Disappeared</a>. For the evaluation framework that keeps this research honest, see <a class="" href="https://southpawriter.com/blog/check-engine-light">Your RAG Pipeline Has a Check Engine Light</a>.</em></p>]]></content:encoded>
            <category>FractalRecall</category>
            <category>Research</category>
            <category>RAG</category>
            <category>Embeddings</category>
        </item>
        <item>
            <title><![CDATA[Your RAG Pipeline Has a Check Engine Light. You're Ignoring It.]]></title>
            <link>https://southpawriter.com/blog/check-engine-light</link>
            <guid>https://southpawriter.com/blog/check-engine-light</guid>
            <pubDate>Wed, 25 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Most production RAG systems have no evaluation framework. No decision engine, no degradation tracking, no rollback criteria. Here's what one looks like.]]></description>
            <content:encoded><![CDATA[
<p>I ran a retrieval experiment that returned perfect zeros across all 36 queries, and every automated check I'd built said "statistically significant." The decision engine considered seven criteria, passed two of them, and issued a NO-GO. The pipeline caught the problem. Not me--the pipeline.</p>
<p>Here's what scares me: most production <a class="" href="https://southpawriter.com/docs/glossary/library/retrieval-augmented-generation">RAG</a> systems don't have a pipeline like that. They don't have decision criteria. They don't have rollback thresholds. They don't have a concept of "this retrieval result is wrong and we should know about it automatically." They ship a model, run some spot checks, and move on to the next sprint.</p>
<p>Your RAG pipeline has a check engine light. You just never installed it.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-dashboard-you-dont-have">The Dashboard You Don't Have<a href="https://southpawriter.com/blog/check-engine-light#the-dashboard-you-dont-have" class="hash-link" aria-label="Direct link to The Dashboard You Don't Have" title="Direct link to The Dashboard You Don't Have" translate="no">​</a></h2>
<p>What I built for <a class="" href="https://southpawriter.com/projects/fractalrecall">FractalRecall</a>'s D-23 experiment isn't perfect. (It caught a real problem but also passed metrics that were wrong.) It exists, though, which puts it ahead of most production systems I've seen.</p>
<p>The framework evaluates seven criteria before issuing a GO or NO-GO:</p>
<ol>
<li class=""><strong>Aggregate Improvement.</strong> Did the overall metrics improve compared to baseline? Not "did they change"--did they move in the right direction by a meaningful amount?</li>
<li class=""><strong>Statistical Significance.</strong> Was the improvement unlikely to be random? (This is where things get interesting, because "statistically significant" and "correct" are not the same thing. More on that in a moment.)</li>
<li class=""><strong>Query Degradation Rate.</strong> What percentage of individual queries got <em>worse</em>? An aggregate improvement can hide widespread degradation if a few queries improved dramatically.</li>
<li class=""><strong>Per-Type Analysis.</strong> Did performance improve across different query types (factual, relational, comparative, authority), or only in some categories?</li>
<li class=""><strong>Effect Size.</strong> Is the improvement practically meaningful, not just mathematically detectable?</li>
<li class=""><strong>Overflow Rate.</strong> Did the enrichment process silently discard data? How much? Is that acceptable?</li>
<li class=""><strong>Stability.</strong> Are the results consistent across runs, or do they fluctuate?</li>
</ol>
<p>Seven criteria. A majority must pass for a GO decision, with hard vetoes on degradation rate and overflow. The framework runs automatically after every experiment.</p>
<p>This took about a day to build. It has caught two genuine problems <a class="" href="https://southpawriter.com/blog/context-windows-are-a-lie">(the D-22 overflow and the D-23 metric bug)</a>. It has saved me from publishing results I would have had to retract. And it is, as far as I can tell, more evaluation infrastructure than most companies apply to their production RAG systems.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-nobody-builds-the-dashboard">Why Nobody Builds the Dashboard<a href="https://southpawriter.com/blog/check-engine-light#why-nobody-builds-the-dashboard" class="hash-link" aria-label="Direct link to Why Nobody Builds the Dashboard" title="Direct link to Why Nobody Builds the Dashboard" translate="no">​</a></h2>
<p>I have theories about this, and none of them are flattering.</p>
<p><strong>Retrieval quality is invisible.</strong> If your chatbot returns a wrong answer, the user might not notice. They didn't know the right answer--that's why they asked. The feedback loop that exists for other software bugs (crash reports, error logs, angry support tickets) barely exists for retrieval errors. The system fails silently, and silence is comfortable. (This also explains the "it works on my test queries" approach: developers run five queries they already know the answers to and call it validated. Confirmation bias with a search bar.)</p>
<p><strong>Evaluation is boring.</strong> Building a RAG pipeline is interesting. Choosing an <a class="" href="https://southpawriter.com/docs/glossary/library/embedding">embedding</a> model! Tuning chunk sizes! Experimenting with <a class="" href="https://southpawriter.com/docs/glossary/library/reranking">re-ranking</a>! Evaluation is where you write tests, compute metrics, and stare at tables of numbers. The vegetables of machine learning and the bane of my existence.</p>
<p><strong>The metrics are hard.</strong> They have jargon names--<a class="" href="https://southpawriter.com/docs/glossary/library/precision">precision</a>, <a class="" href="https://southpawriter.com/docs/glossary/library/recall">recall</a>, <a class="" href="https://southpawriter.com/docs/glossary/library/ndcg">NDCG</a>, <a class="" href="https://southpawriter.com/docs/glossary/library/mrr">MRR</a>--but the questions they ask are simple. Did you return the right stuff? Did you find <em>all</em> of it? Is the best result near the top? How far does the user scroll before hitting something useful? The problem is that "retrieval quality" is not one thing. A system can return only relevant results but miss half the relevant documents. It can ace factual queries and completely miss relational ones. Understanding the metrics means understanding those tradeoffs, and the tradeoffs require caring enough to learn.</p>
<p>I suspect the boring one is strongest. Evaluation doesn't feel like progress. It feels like homework.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-statistically-significant-actually-means-a-cautionary-tale">What Statistically Significant Actually Means (A Cautionary Tale)<a href="https://southpawriter.com/blog/check-engine-light#what-statistically-significant-actually-means-a-cautionary-tale" class="hash-link" aria-label="Direct link to What Statistically Significant Actually Means (A Cautionary Tale)" title="Direct link to What Statistically Significant Actually Means (A Cautionary Tale)" translate="no">​</a></h2>
<p>D-23 is the single best argument I have for why evaluation infrastructure matters.</p>
<p>I ran 36 queries against a corpus enriched with eight layers of structural context: domain, entity type, authority status, temporal era, relationships, section heading, and sequence position. Every query returned zero relevant documents. Every metric (precision, recall, NDCG, MRR) came back as 0.000.</p>
<p>The statistical test I ran (comparing D-23 to D-22 baseline) returned a p-value (a confidence score for "this result isn't random") of approximately 0.0000000000000000007. That is not a typo. It was reporting near-absolute certainty that the results were different from baseline.</p>
<p>And it was correct! The results <em>were</em> different from baseline. They were zero. All of them. Because I had a bug in my metric computation where chunk IDs included a <code>#chunk_002</code> suffix that didn't match the expected document filenames.</p>
<p>The statistical test did exactly what it was designed to do: determine whether two distributions were likely to be different. It is not designed to determine whether your code is correct. It is not designed to determine whether the result makes sense. It is designed to answer one narrow question about two sets of numbers, and it answered that question with extreme confidence.</p>
<p>If my evaluation framework had consisted only of "run the statistical test and check if p &lt; 0.05 (the standard threshold for 'probably not random')," I would have concluded that multi-layer enrichment catastrophically degrades retrieval quality. That conclusion would have been supported by a p-value that most journals would kill for. And it would have been <em>spectacularly</em> wrong.</p>
<p>The GO/NO-GO framework caught it--not through the significance test, but through the degradation rate criterion. When 100% of queries degrade, something is broken, regardless of what the p-value says. The dashboard has multiple indicators because no single indicator is sufficient.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-your-dashboard-should-actually-check">What Your Dashboard Should Actually Check<a href="https://southpawriter.com/blog/check-engine-light#what-your-dashboard-should-actually-check" class="hash-link" aria-label="Direct link to What Your Dashboard Should Actually Check" title="Direct link to What Your Dashboard Should Actually Check" translate="no">​</a></h2>
<p>None of this is theoretical. These are checks that have caught real bugs in my own work.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-baseline-comparison-non-negotiable">1. Baseline Comparison (Non-Negotiable)<a href="https://southpawriter.com/blog/check-engine-light#1-baseline-comparison-non-negotiable" class="hash-link" aria-label="Direct link to 1. Baseline Comparison (Non-Negotiable)" title="Direct link to 1. Baseline Comparison (Non-Negotiable)" translate="no">​</a></h3>
<p>Every change to your retrieval pipeline (new embedding model, different chunk size, updated re-ranking) should be compared against a frozen baseline using the same queries and the same ground truth. Not "we ran some queries and it seemed better." A structured comparison with numbers.</p>
<p>If you don't have ground truth (a set of queries with known-correct answers), build it. Manually, if you have to. Twenty well-curated queries with verified answers is worth more than a thousand unverified spot checks.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-per-query-degradation-rate">2. Per-Query Degradation Rate<a href="https://southpawriter.com/blog/check-engine-light#2-per-query-degradation-rate" class="hash-link" aria-label="Direct link to 2. Per-Query Degradation Rate" title="Direct link to 2. Per-Query Degradation Rate" translate="no">​</a></h3>
<p>Aggregate metrics lie. They average out the disasters.</p>
<p>If your new pipeline improves average recall (the fraction of relevant content it actually finds) from 0.72 to 0.78 but degrades 40% of individual queries, you have not improved your system. You have improved your system for some users while making it worse for others. The aggregate number is a press release. The degradation rate is the truth.</p>
<p>My threshold: if more than 25% of queries degrade, the change does not ship. Period. Even if the aggregate improves. Especially if the aggregate improves--because that means a small number of dramatic improvements are masking widespread damage.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="3-overflow--data-loss-tracking">3. Overflow / Data Loss Tracking<a href="https://southpawriter.com/blog/check-engine-light#3-overflow--data-loss-tracking" class="hash-link" aria-label="Direct link to 3. Overflow / Data Loss Tracking" title="Direct link to 3. Overflow / Data Loss Tracking" translate="no">​</a></h3>
<p>If your pipeline involves any transformation that can discard data (token limits, chunk size constraints, format conversion, deduplication), track the loss rate. I <a class="" href="https://southpawriter.com/blog/43-percent-disappeared">learned this the hard way</a>: silent data loss can hide behind improving metrics.</p>
<p>Assert on your chunk counts. Before indexing: N documents. After indexing: N documents. If those numbers differ, stop and find out why.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="4-type-level-breakdowns">4. Type-Level Breakdowns<a href="https://southpawriter.com/blog/check-engine-light#4-type-level-breakdowns" class="hash-link" aria-label="Direct link to 4. Type-Level Breakdowns" title="Direct link to 4. Type-Level Breakdowns" translate="no">​</a></h3>
<p>Don't just compute overall precision. Compute precision by query type, by document type, by topic. A system that's amazing at factual queries and terrible at relational queries is not "pretty good overall"--it's broken for relationship queries and nobody noticed because the factual queries pulled the average up.</p>
<p>In my experiments, authority queries ("Which documents are canonical?") behaved differently from relational queries ("What factions are connected to the Dvergr?"). Aggregate metrics hid this. Type-level breakdowns revealed it.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="5-the-veto">5. The Veto<a href="https://southpawriter.com/blog/check-engine-light#5-the-veto" class="hash-link" aria-label="Direct link to 5. The Veto" title="Direct link to 5. The Veto" translate="no">​</a></h3>
<p>Every framework needs a hard stop--a condition where the change does not ship regardless of any other metric. For me, it's 100% query degradation (which catches catastrophic bugs like D-23's) and overflow above 50% (which catches silent data loss like D-22's).</p>
<p>Your vetoes will be different. But you need them. Without a veto, there is always a way to rationalize shipping a broken change. "The aggregate improved." "The p-value is significant." "We're behind on the roadmap." A veto is a firewall against motivated reasoning.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-boring-part-is-the-important-part">The Boring Part Is the Important Part<a href="https://southpawriter.com/blog/check-engine-light#the-boring-part-is-the-important-part" class="hash-link" aria-label="Direct link to The Boring Part Is the Important Part" title="Direct link to The Boring Part Is the Important Part" translate="no">​</a></h2>
<p>I know this is not a sexy blog post. It doesn't have a dramatic twist or a surprising result. It says: build a dashboard, check your metrics, don't trust any single number.</p>
<p>But here's what I keep coming back to: I have a <a class="" href="https://southpawriter.com/blog/document-is-database">research project</a> with three experiments, a corpus of 77 documents, and 36 queries. Small-scale, personal, low-stakes. And even at that scale, evaluation infrastructure caught two bugs that would have produced wrong conclusions.</p>
<p>If my little retrieval experiment needs a seven-criterion decision framework to avoid publishing nonsense, what does your production RAG system need? The one serving real users, handling real queries, influencing real decisions.</p>
<p>Whatever it is, you probably don't have it yet. And the reason you don't is not that you can't build it. It's that you haven't decided it's worth building. The vibes feel good. The spot checks pass. The users aren't complaining.</p>
<p>The check engine light is off because it was never wired in.</p>
<p>Wire it in.</p>]]></content:encoded>
            <category>Opinion</category>
            <category>Research</category>
            <category>RAG</category>
        </item>
        <item>
            <title><![CDATA[Five Projects, One Realization: The Document Is the Database]]></title>
            <link>https://southpawriter.com/blog/document-is-database</link>
            <guid>https://southpawriter.com/blog/document-is-database</guid>
            <pubDate>Tue, 24 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Five AI projects, one pattern: they all treat documents as structured knowledge systems, not content delivery vehicles. The filing cabinet analogy, expanded.]]></description>
            <content:encoded><![CDATA[
<p>I didn't plan a portfolio. I planned a Markdown file. Then another one. Then five projects materialized around them like ice crystals on a cold window, each shaped by the same principle I didn't recognize until project number four. Apparently I need to build the same insight multiple times before I notice I keep building it.</p>
<p>The insight: <strong>documents are not content delivery vehicles. They are structured knowledge systems.</strong> Almost every AI tool in production today throws away the structure and keeps only the content. That's like buying a filing cabinet, dumping all the folders on the floor, and asking someone to find last quarter's tax return by feeling the texture of the paper.</p>
<p>I know this because I've now built five projects that all, in their own way, try to fix that mistake.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-lineup">The Lineup<a href="https://southpawriter.com/blog/document-is-database#the-lineup" class="hash-link" aria-label="Direct link to The Lineup" title="Direct link to The Lineup" translate="no">​</a></h2>
<p>Let me introduce them in birth order. (I'm using "born" loosely. A project is "born" when it gets its first Markdown file. The code comes later. Sometimes much later.)</p>
<p><strong><a class="" href="https://southpawriter.com/projects/llmstxtkit">LlmsTxtKit</a></strong> parses <a class="" href="https://southpawriter.com/docs/glossary/library/llms-txt">llms.txt</a> files, a standard where websites publish Markdown summaries of their content so AI systems can understand them without crawling every page. I built a C#/.NET library that fetches, parses, validates, caches, and generates context from these files. The idea is elegant: instead of making an AI read your whole website, hand it a curated summary. Give it the filing cabinet with the folders intact.</p>
<p><strong><a class="" href="https://southpawriter.com/projects/docstratum">DocStratum</a></strong> validates those same files against the spec. Think ESLint, but for a Markdown standard defined by a blog post. If LlmsTxtKit is "here's how to read the file," DocStratum is "here's whether the file was written correctly."</p>
<p><strong><a class="" href="https://southpawriter.com/projects/fractalrecall">FractalRecall</a></strong> is where things got interesting. Most retrieval systems (the "R" in <a class="" href="https://southpawriter.com/docs/glossary/library/retrieval-augmented-generation">RAG</a>) chop documents into chunks, embed them as vectors, and search by similarity. The chunks are orphans. They know <em>what they say</em> but not <em>what they are</em>. FractalRecall's thesis is that if you tell the <a class="" href="https://southpawriter.com/docs/glossary/library/embedding">embedding</a> model what kind of document a chunk came from--its domain, its authority status, its temporal context--retrieval quality improves. I tested this. It does. Twenty-four tokens of structural context improved my retrieval quality by 16.5%. That's less text than this sentence.</p>
<p><strong><a class="" href="https://southpawriter.com/projects/haiku-protocol">Haiku Protocol</a></strong> attacks the opposite end: instead of adding context, it compresses content. A <a class="" href="https://southpawriter.com/docs/glossary/library/controlled-natural-language">Controlled Natural Language</a> system that transforms verbose prose into dense, machine-readable strings. Same information, fewer tokens. If your <a class="" href="https://southpawriter.com/docs/glossary/library/context-window">context window</a> is functionally 8K tokens despite the marketing department claiming 128K, every token you save on content is a token you can spend on structure.</p>
<p><strong><a class="" href="https://southpawriter.com/projects/chronicle">Chronicle</a></strong> ties it together. I didn't realize that until embarrassingly late. Chronicle treats worldbuilding lore like a software codebase: Markdown files in a Git repo, YAML <a class="" href="https://southpawriter.com/docs/glossary/library/frontmatter">frontmatter</a> for metadata, deterministic validation for consistency, FractalRecall for <a class="" href="https://southpawriter.com/docs/glossary/library/semantic-search">semantic search</a>. Version-controlled fiction with CI/CD for your canon.</p>
<p>Five projects. Five different problems. One pattern I kept accidentally rediscovering.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-pattern">The Pattern<a href="https://southpawriter.com/blog/document-is-database#the-pattern" class="hash-link" aria-label="Direct link to The Pattern" title="Direct link to The Pattern" translate="no">​</a></h2>
<p>Every one of these projects treats a <em>document</em>--not a database row, not a JSON object, not a feature vector--as the fundamental unit of knowledge. And every one insists that the <em>structure</em> of that document carries meaning the AI pipeline has an obligation to preserve.</p>
<p>LlmsTxtKit: your website is a structured document. Give AI the structure, not just the text.</p>
<p>DocStratum: that structure has rules. Verify them.</p>
<p>FractalRecall: when you embed that document for retrieval, the structure should travel with it.</p>
<p>Haiku Protocol: when you compress it, compress the prose. Not the structure.</p>
<p>Chronicle: manage these documents with the same rigor you'd give source code, because they <em>are</em> source code--for knowledge.</p>
<p>I could dress that up, but the pattern is blunt enough to say flat. The common thread isn't AI, isn't embeddings or context windows or Markdown parsing. It's a conviction that <strong>documents are databases</strong>--that a well-structured Markdown file with YAML frontmatter carries more retrievable intelligence than a row in PostgreSQL, because it carries both the content <em>and</em> the organizational context that makes the content findable.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-this-matters">Why This Matters<a href="https://southpawriter.com/blog/document-is-database#why-this-matters" class="hash-link" aria-label="Direct link to Why This Matters" title="Direct link to Why This Matters" translate="no">​</a></h2>
<p>Here's what frustrates me about the current RAG ecosystem.</p>
<p>The standard workflow: take your documents, chop them into 512-token chunks, embed them as vectors, store them in a vector database, find the nearest neighbors at query time. Simple. Elegant. And it throws away everything that made those documents <em>documents</em>.</p>
<p>When you chunk a technical manual, you lose the chapter structure. Chunk a worldbuilding corpus, you lose the distinction between canonical lore and speculative drafts. Chunk a legal contract, you lose the hierarchy of clauses and subclauses that determines what's binding and what's illustrative. The chunks are semantically meaningful fragments floating in a void, stripped of the logical intelligence that we human authors spent hours building into the document's structure.</p>
<p>Then we bolt metadata back on <em>after the fact</em>--as database columns, filter fields, post-retrieval classification steps--and wonder why retrieval quality plateaus. We're reconstructing information that was right there in the original document. Before we destroyed it.</p>
<p>That's the filing cabinet problem. The structure was there. We threw it on the floor. Now we're building increasingly sophisticated AI systems to figure out which pile the tax return is in. We could have just kept the folders.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-document-as-first-class-citizen">The Document as First-Class Citizen<a href="https://southpawriter.com/blog/document-is-database#the-document-as-first-class-citizen" class="hash-link" aria-label="Direct link to The Document as First-Class Citizen" title="Direct link to The Document as First-Class Citizen" translate="no">​</a></h2>
<p>The alternative is simpler than it sounds.</p>
<p>Treat the document as a first-class citizen in the AI pipeline. Not as a source of text to be extracted and discarded, but as a structured knowledge object whose organization carries meaning at every stage.</p>
<p>LlmsTxtKit does this at the publishing stage: give AI systems the document's structure directly instead of making them reconstruct it from HTML. DocStratum does it at validation: enforce structural consistency so the AI can trust what it receives. FractalRecall does it at embedding: encode structural context into the vector itself, so retrieval is structurally aware from the start. Haiku Protocol does it at compression: preserve structural relationships even when reducing token count. Chronicle does it at management: version-control documents with the same discipline we give code.</p>
<p>None of these ideas are individually revolutionary. But I haven't seen anyone connect them into a coherent pipeline. I think the reason is cultural: the AI industry sees documents as <em>input</em>—raw material to be processed and discarded. I see them as <em>infrastructure</em>. The load-bearing walls of a knowledge system.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-accidental-architecture">The Accidental Architecture<a href="https://southpawriter.com/blog/document-is-database#the-accidental-architecture" class="hash-link" aria-label="Direct link to The Accidental Architecture" title="Direct link to The Accidental Architecture" translate="no">​</a></h2>
<p>I want to be clear: I did not plan this. I didn't sit down and decide to build five complementary projects forming a coherent document-centric AI pipeline. I built LlmsTxtKit because I was frustrated that AI couldn't read my website. DocStratum because the llms.txt files I was parsing were full of spec violations. FractalRecall because embedding retrieval kept returning the wrong documents. Haiku Protocol because context windows are smaller than advertised and I was angry about it. Chronicle because my worldbuilding corpus was a mess and I have very specific feelings about version control.</p>
<p>Each project solved a real problem. The pattern emerged <em>after</em> the projects existed.</p>
<p>I wrote the documentation first--obviously--but I wrote the unifying thesis last. That's probably the most honest thing a documentation-first developer has ever admitted.</p>
<p>The through-line, now that I can see it: <strong>the document is the database.</strong> The structure is the schema. The metadata is the index. The content is the data. Build your AI pipeline to respect that--from ingestion to retrieval to compression to delivery--and you get better results than treating text as an undifferentiated stream of tokens.</p>
<p>I have the experiment results to prove it. Those are stories for upcoming posts.</p>
<hr>
<p><em>This is the first in a series connecting the <a class="" href="https://southpawriter.com/projects#ai-and-llm-research">AI and LLM Research projects</a>. Next: how 24 tokens of metadata improved retrieval by 16.5%--and why losing 43% of my data somehow made things better.</em></p>]]></content:encoded>
            <category>Opinion</category>
            <category>llms.txt</category>
            <category>Documentation</category>
        </item>
        <item>
            <title><![CDATA[I Added Context to My Embeddings and 43% of My Data Disappeared]]></title>
            <link>https://southpawriter.com/blog/43-percent-disappeared</link>
            <guid>https://southpawriter.com/blog/43-percent-disappeared</guid>
            <pubDate>Mon, 23 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[I prepended 24 tokens of metadata to each chunk. 43% overflowed the token limit and vanished. Retrieval quality went up anyway. I have questions.]]></description>
            <content:encoded><![CDATA[
<p>In <a class="" href="https://southpawriter.com/blog/context-windows-are-a-lie">Part 1</a>, I mentioned the D-22 experiment almost as an aside. Twenty-four tokens of metadata prefix, 16.5% improvement in ranking quality, 27.3% recall jump. Good numbers. Clean story.</p>
<p>I left out the part where 43% of my data vanished.</p>
<p>Not "performed poorly." Not "returned lower-quality results." <em>Vanished.</em> Ninety-four of 218 chunks silently dropped from the index because I added one sentence of context and didn't do the arithmetic on what that sentence would cost. The <a class="" href="https://southpawriter.com/docs/glossary/library/embedding">embedding</a> pipeline didn't warn me. ChromaDB didn't complain. I only noticed because I'm the kind of person who checks row counts after every insert. (This is not a personality trait. It's scar tissue.)</p>
<p>The results improved anyway. That's the part I need to explain.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="silent-overflow">Silent Overflow<a href="https://southpawriter.com/blog/43-percent-disappeared#silent-overflow" class="hash-link" aria-label="Direct link to Silent Overflow" title="Direct link to Silent Overflow" translate="no">​</a></h2>
<p>Embedding models have a <a class="" href="https://southpawriter.com/docs/glossary/library/token">token</a> budget. <code>nomic-embed-text-v1.5</code> accepts up to 8,192 tokens per input. My chunks were sized for a 1,024-token target window. Most fit comfortably.</p>
<p>Then I added the prefix.</p>
<p><em>"Domain: faction. Entity: Iron-Banes Alliance. Canon: true."</em></p>
<p>Twenty-four tokens on average. The chunks that were already near their limit tipped over. The embedding pipeline didn't raise an exception or log a warning. It just skipped them. Ninety-four chunks, gone. The index built successfully, the queries ran, the metrics came back looking healthy. If I hadn't compared the chunk count in ChromaDB against the source count in my data directory, I'd have published the D-22 results without ever knowing half the corpus was missing.</p>
<p>This is the failure mode nobody warns you about in <a class="" href="https://southpawriter.com/docs/glossary/library/retrieval-augmented-generation">RAG</a> tutorials. Not "bad results." Invisible data loss that makes your results look <em>better</em> than they should.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-confound">The Confound<a href="https://southpawriter.com/blog/43-percent-disappeared#the-confound" class="hash-link" aria-label="Direct link to The Confound" title="Direct link to The Confound" translate="no">​</a></h2>
<p>Here's the D-22 data, which I presented in Part 1 without enough caveats:</p>
<table><thead><tr><th>Metric</th><th>D-21 (Baseline)</th><th>D-22 (Prefix)</th><th>Change</th></tr></thead><tbody><tr><td>Recall@10</td><td>0.720</td><td>0.917</td><td><strong>+27.3%</strong></td></tr><tr><td>NDCG@10</td><td>0.706</td><td>0.823</td><td><strong>+16.5%</strong></td></tr><tr><td>Precision@5</td><td>0.383</td><td>0.417</td><td>+8.8%</td></tr><tr><td>MRR</td><td>0.845</td><td>0.861</td><td>+1.9%</td></tr></tbody></table>
<p>Every metric improved. With 43% of the corpus missing.</p>
<p>I sat with this for a while. Then I sat with it some more. Then I wrote a findings document, because that's apparently how I process confusion.</p>
<p>The problem: I can't tell you how much of this improvement came from the metadata enrichment and how much came from accidentally deleting the worst chunks. Two things happened simultaneously, and I didn't instrument the experiment to separate them.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="three-explanations">Three Explanations<a href="https://southpawriter.com/blog/43-percent-disappeared#three-explanations" class="hash-link" aria-label="Direct link to Three Explanations" title="Direct link to Three Explanations" translate="no">​</a></h2>
<p><strong>The dropped chunks were noise.</strong> Maybe the chunks that overflowed were the weakest ones, already too long, too dense, too full of prose that added bulk without adding signal. Dropping them was accidental data cleaning. The corpus got leaner and leaner corpora retrieve better.</p>
<p>Plausible. Also uncomfortable. If losing 43% of my data improves results, that's less "my enrichment strategy is brilliant" and more "my <a class="" href="https://southpawriter.com/docs/glossary/library/chunking">chunking</a> strategy was terrible."</p>
<p><strong>The metadata was doing the heavy lifting.</strong> Without the prefix, the embedding model knew a chunk was about armor, warfare, and territorial disputes. With the prefix, it knew this was <em>a canonical faction document about a specific organization</em>. The semantic fingerprint sharpened. Queries about factions found faction documents. Queries about authority found canonical documents. The prefix acted as a disambiguator, not just a label.</p>
<p>This is what I believe. This is what matters for <a class="" href="https://southpawriter.com/projects/fractalrecall">FractalRecall</a>'s thesis.</p>
<p><strong>Both.</strong> The enrichment improved the surviving chunks <em>and</em> the overflow removed the weakest chunks. Two effects, additive. The improvement was real but inflated by a confounding variable I can't measure because I didn't log which chunks overflowed.</p>
<p>This is probably the honest answer.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-24-tokens-replace">What 24 Tokens Replace<a href="https://southpawriter.com/blog/43-percent-disappeared#what-24-tokens-replace" class="hash-link" aria-label="Direct link to What 24 Tokens Replace" title="Direct link to What 24 Tokens Replace" translate="no">​</a></h2>
<p>The prefix is three facts in one sentence:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:#2a2a2a;--prism-color:#f92aad"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:#2a2a2a;background-image:#34294f;color:#f92aad;text-shadow:0 0 2px #100c0f, 0 0 5px #dc078e33, 0 0 10px #fff3"><code class="codeBlockLines_e6Vv"><span class="token-line" style="background-image:#34294f;color:#f92aad;text-shadow:0 0 2px #100c0f, 0 0 5px #dc078e33, 0 0 10px #fff3"><span class="token plain">Domain: faction. Entity: Iron-Banes Alliance. Canon: true.</span><br></span></code></pre></div></div>
<p>A domain classification. An entity name. An authority flag. Roughly the same token cost as "once upon a time, in a land far, far away," except this version actually helps retrieval.</p>
<p>Compare that to what those tokens replaced in the chunks that overflowed:</p>
<blockquote>
<p><em>The Iron-Banes Alliance maintains a complex hierarchical structure that has evolved significantly over the centuries, reflecting both its martial origins and its subsequent...</em></p>
</blockquote>
<p>Twenty-four tokens of throat-clearing. It told the embedding model the chunk was about <em>some kind of hierarchy that evolved</em>. The prefix told it the chunk was about <em>a specific canonical faction</em>. The model doesn't care about your prose style. It cares about information density.</p>
<p>This is the lesson I keep circling back to from <a class="" href="https://southpawriter.com/blog/context-windows-are-a-lie">Part 1</a>: in information retrieval, what you <em>tell</em> the system is more valuable than what you <em>show</em> it. RAG pipelines hand the model raw text and say "figure it out." Embedding models are good at this. But inferring topic, domain, entity type, and authority from 974 tokens of narrative is harder than just being told in 24 tokens of structured prefix.</p>
<p>It's the difference between handing someone a novel and asking "what genre?" versus writing "MYSTERY" on the spine. The novel contains everything needed to answer the question. The label is faster.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-id-do-differently">What I'd Do Differently<a href="https://southpawriter.com/blog/43-percent-disappeared#what-id-do-differently" class="hash-link" aria-label="Direct link to What I'd Do Differently" title="Direct link to What I'd Do Differently" translate="no">​</a></h2>
<p>Three things, in order of how much sleep they've cost me.</p>
<p><strong>Log the overflow.</strong> I know 94 of 218 chunks overflowed. I don't know <em>which</em> 94. Were they systematically different from the survivors? Longer? From specific document types? Without IDs, I can't separate the enrichment signal from the overflow cleaning signal. D-23 logs everything.</p>
<p><strong>Reserve the token budget.</strong> Dropping 43% of chunks is not a strategy. It was an arithmetic oversight: I sized chunks for 1,024 tokens, then added prefix tokens that pushed some of them over. D-23 addresses this with a <code>prefix_reserve</code> mechanism that pre-allocates token budget for enrichment, guaranteeing zero overflow. (It produced zero overflow. It also produced all-zero metrics, because I introduced a different bug entirely. Research is glamorous.)</p>
<p><strong>Start with the confound, not the result.</strong> I presented the D-22 numbers in Part 1 as a clean improvement story. They're not clean. The improvement is real, probably, but the magnitude is inflated by uncontrolled data loss. I should have led with the caveat. I'm leading with it now.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-actual-takeaway">The Actual Takeaway<a href="https://southpawriter.com/blog/43-percent-disappeared#the-actual-takeaway" class="hash-link" aria-label="Direct link to The Actual Takeaway" title="Direct link to The Actual Takeaway" translate="no">​</a></h2>
<p>Twenty-four tokens of the <em>right</em> information outperformed 974 tokens of undifferentiated narrative. The prefix told the embedding model <em>what the topic was</em>. The raw text told it <em>about</em> the topic. That's a different kind of information, and embedding models use it better than most RAG engineers expect.</p>
<p>Also: check your chunk counts after indexing. Always. Learn from my 43%.</p>
<hr>
<p><em>This is Part 2 of the <a class="" href="https://southpawriter.com/blog/tags/research">Research Notebooks</a> series. Part 1: <a class="" href="https://southpawriter.com/blog/context-windows-are-a-lie">Context Windows Are a Lie</a>. Next: the D-23 multi-layer experiment, where I ran a perfectly executed pipeline and every single metric was wrong.</em></p>]]></content:encoded>
            <category>FractalRecall</category>
            <category>Research</category>
            <category>RAG</category>
            <category>Embeddings</category>
        </item>
        <item>
            <title><![CDATA[Google Said No to llms.txt. Five Google Teams Didn't Get the Memo.]]></title>
            <link>https://southpawriter.com/blog/google-didnt-get-the-memo</link>
            <guid>https://southpawriter.com/blog/google-didnt-get-the-memo</guid>
            <pubDate>Sun, 22 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Google's top search voices called llms.txt a dead meta tag. Then I found llms.txt files on five Google developer documentation properties and an AGENTS.md that tells AI to read them.]]></description>
            <content:encoded><![CDATA[
<p>The timeline is where the joke lives.</p>
<p><strong>April 2025.</strong> Google's John Mueller <a href="https://www.searchenginejournal.com/google-says-llms-txt-comparable-to-keywords-meta-tag/544804/" target="_blank" rel="noopener noreferrer" class="">compares llms.txt to the keywords meta tag</a>. For the uninitiated, the keywords meta tag is so discredited that invoking it in SEO circles is equivalent to recommending bloodletting at a medical conference. Mueller's message was clear: <a class="" href="https://southpawriter.com/docs/glossary/library/llms-txt">llms.txt</a> is unnecessary, self-reported data that Google has no intention of using.</p>
<p><strong>July 2025.</strong> Gary Illyes, also from Google's Search team, confirms the position at Search Central Live. <a href="https://searchengineland.com/google-says-normal-seo-works-for-ranking-in-ai-overviews-and-llms-txt-wont-be-used-459422" target="_blank" rel="noopener noreferrer" class="">No support. Won't be used.</a> Normal SEO works fine for AI Overviews. The standard is, officially, not something Google is interested in.</p>
<p><strong>December 3, 2025.</strong> An SEO professional named Lidia Infante discovers an llms.txt file on Google's own Search Central documentation. Mueller's response, posted to Bluesky: "hmmn :-/". The file was removed within hours.</p>
<p>So far, a clean narrative. Google said no, someone at Google accidentally deployed one, it was caught and deleted, and the official position holds. Embarrassing, but coherent.</p>
<p>Then I started pulling at threads.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="five-files-walk-into-a-search-engine">Five Files Walk Into a Search Engine<a href="https://southpawriter.com/blog/google-didnt-get-the-memo#five-files-walk-into-a-search-engine" class="hash-link" aria-label="Direct link to Five Files Walk Into a Search Engine" title="Direct link to Five Files Walk Into a Search Engine" translate="no">​</a></h2>
<p>I'm in the middle of building an evidence inventory for an analytical paper about llms.txt (the <a class="" href="https://southpawriter.com/blog/waf-paradox">Access Paradox research</a> that's been taking over my evenings since mid-February). I work with Claude on this project (<a class="" href="https://southpawriter.com/blog/validator-made-up">introduced properly two posts ago</a>), and earlier this week Claude flagged something I'd been putting off: making local copies of every source in the evidence inventory. PDFs, screenshots, the works. "Sources disappear," or words to that effect. I'd been meaning to do it anyway. So I started going through the inventory, saving things.</p>
<p>One of the sources was an article from Omnius titled "Google Adds LLMs.txt to Docs After Publicly Dismissing It." The article was gone. The domain acknowledged it had existed, the Wayback Machine had a record of the URL, but nobody had indexed the actual content (or the offending llms.txt file) before it vanished. An article about Google quietly adopting llms.txt had itself quietly disappeared.</p>
<p><img decoding="async" loading="lazy" alt="Wayback Machine capture of the Omnius article listing showing &amp;quot;Google Adds LLMs.txt to Docs After Publicly Dismissing It,&amp;quot; dated Dec 3, with a summary describing Google quietly adding and then removing an llms.txt file from Search Central documentation. The article itself was never indexed." src="https://southpawriter.com/assets/images/omnius-wayback-1a8d387e74c85f2f41736bf151e76dba.png" width="1090" height="1092" class="img_ev3q">
So naturally I went looking for the primary evidence myself.</p>
<p>I checked Google.</p>
<p>Not Google Search. Not the property where the December incident happened. Google's <em>developer documentation</em> properties. The teams that write guides for Firebase, Chrome extensions, the AI SDK, Flutter, web performance. The teams whose audience is developers building things, not marketers optimizing rankings.</p>
<p>Here's what I found:</p>
<ul>
<li class=""><strong>ai.google.dev</strong> -- llms.txt at <code>/api/llms.txt</code></li>
<li class=""><strong>developer.chrome.com</strong> -- llms.txt at <code>/docs/llms.txt</code> (including Flutter documentation)</li>
<li class=""><strong>firebase.google.com</strong> -- llms.txt at <code>/docs/llms.txt</code></li>
<li class=""><strong>google.github.io/adk-docs</strong> -- llms.txt at <code>/llms.txt</code></li>
<li class=""><strong>web.dev</strong> -- llms.txt at <code>/articles/llms.txt</code> (the only one not under <code>/docs/</code>)</li>
</ul>
<p>Five properties. All live. All serving llms.txt files while two of Google's most visible search representatives were telling the world the standard was comparable to a dead meta tag.</p>
<p>I saved PDF copies of every single one, because I have been burned by disappearing evidence before and because I am the kind of person who maintains a local archive manifest in his research repository. (This is not a personality trait I recommend developing. It is, however, one I cannot stop.)</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-content-isnt-revolutionary-the-adoption-is">The Content Isn't Revolutionary. The Adoption Is.<a href="https://southpawriter.com/blog/google-didnt-get-the-memo#the-content-isnt-revolutionary-the-adoption-is" class="hash-link" aria-label="Direct link to The Content Isn't Revolutionary. The Adoption Is." title="Direct link to The Content Isn't Revolutionary. The Adoption Is." translate="no">​</a></h2>
<p>Let me manage expectations. These files aren't going to win any documentation awards. They're basic sitemap-style link lists. No rich summaries, no curated context hierarchies, no sophisticated <a class="" href="https://southpawriter.com/docs/glossary/library/inference">inference</a>-time optimizations. If DocStratum were grading them, they'd pass L0 (it parses) and maybe L1 (it has structure), but they wouldn't be winning medals at L3 or L4.</p>
<p>That's not the story.</p>
<p>The story is that these files exist at all. Five separate Google developer documentation teams independently decided that llms.txt was worth implementing, <em>after</em> two of Google's most senior search voices publicly said it wasn't. Nobody made a press release. Nobody updated the corporate talking points. They just shipped it.</p>
<p>If you've ever worked in a large organization, you know exactly what happened here. Policy flows downhill. Implementation flows sideways. The people writing Firebase documentation are solving a different problem than the people briefing journalists about Search ranking signals, and the two groups do not consult each other before deploying a text file.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="agentsmd">AGENTS.md<a href="https://southpawriter.com/blog/google-didnt-get-the-memo#agentsmd" class="hash-link" aria-label="Direct link to AGENTS.md" title="Direct link to AGENTS.md" translate="no">​</a></h2>
<p>Google's Agent Development Kit (ADK) Python repository, <code>google/adk-python</code>, includes a file called <code>AGENTS.md</code>. It's essentially a context document for <a class="" href="https://southpawriter.com/docs/glossary/library/ai-agent">AI agents</a>: Claude Code, Gemini CLI, GitHub Copilot, Cursor, and any other AI coding assistant that might need to understand the project.</p>
<p>At the bottom of <code>AGENTS.md</code>, in the Additional Resources section, there's this line:</p>
<blockquote>
<p><strong>LLM Context:</strong> <code>llms.txt</code> (summarized), <code>llms-full.txt</code> (comprehensive)</p>
</blockquote>
<p>A Google repository is <em>explicitly instructing AI agents to use llms.txt</em> as their entry point for understanding the codebase. Not accidentally hosting one. Not passively allowing one to exist. <em>Directing</em> AI tools to consume it. At <a class="" href="https://southpawriter.com/docs/glossary/library/inference">inference</a> time. For context loading. Which is <em>exactly</em> what the <a href="https://llmstxt.org/" target="_blank" rel="noopener noreferrer" class="">llms.txt specification</a> was designed for.</p>
<p>This is the thing Jeremy Howard proposed in September 2024, that Mueller compared to a dead meta tag seven months later, that Illyes said Google wouldn't use three months after that, that Mueller said "hmmn :-/" about when caught with one in December, and that Google's ADK team is now <em>explicitly telling AI agents to read.</em></p>
<p>I don't know what the organizational chart looks like between Google's Search team and Google's ADK team. But I know what the git history looks like, and the git history says these files shipped.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-github-issue-that-says-the-quiet-part-loud">The GitHub Issue That Says the Quiet Part Loud<a href="https://southpawriter.com/blog/google-didnt-get-the-memo#the-github-issue-that-says-the-quiet-part-loud" class="hash-link" aria-label="Direct link to The GitHub Issue That Says the Quiet Part Loud" title="Direct link to The GitHub Issue That Says the Quiet Part Loud" translate="no">​</a></h2>
<p>There's also <a href="https://github.com/google/adk-docs/issues/726" target="_blank" rel="noopener noreferrer" class="">Issue #726 on google/adk-docs</a>, titled "Update llms.txt to align with the llms.txt standard and act as a sitemap for models." The title alone frames llms.txt as <em>a standard worth aligning with.</em> Not a discredited meta tag. Not an unnecessary format. A standard. Worth conforming to.</p>
<p>This is grassroots adoption pressure from Google's own developer community, showing up in Google's own issue trackers, referencing the same specification Google's search executives publicly dismissed.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-this-actually-means">What This Actually Means<a href="https://southpawriter.com/blog/google-didnt-get-the-memo#what-this-actually-means" class="hash-link" aria-label="Direct link to What This Actually Means" title="Direct link to What This Actually Means" translate="no">​</a></h2>
<p>Google hasn't reversed its position on llms.txt. No announcement, no blog post, no update to Search Central. Mueller and Illyes haven't recanted. (I've <a class="" href="https://southpawriter.com/blog/844k-sites-that-werent">written before</a> about what happens when people get sloppy with adoption narratives, and I'm not about to do it myself.)</p>
<p>But the official line no longer describes what's actually happening inside Google's developer ecosystem. Five documentation teams have llms.txt files. The ADK team has an explicit inference-time directive pointing AI agents at theirs. A community issue is asking for better spec alignment. These aren't accidents; they're not residual files from a confused deployment. They're active implementations across distinct properties.</p>
<p>The gap between "we don't need this" and "hey AI, read our llms.txt" is the gap between policy and practice. And if you've been following <a class="" href="https://southpawriter.com/blog/waf-paradox">this research series</a>, that gap should feel familiar. It's the same pattern.</p>
<p>Your <a class="" href="https://southpawriter.com/docs/glossary/library/web-application-firewall">WAF</a> blocks AI crawlers while your marketing team publishes llms.txt files. Your executives dismiss a standard while your developers implement it. Your <a class="" href="https://southpawriter.com/docs/glossary/library/robots-txt">robots.txt</a> says one thing, your infrastructure does another. The institutions that shape how AI interacts with the web are not, it turns out, internally coherent on the subject. I'm starting to think this is the rule, not the exception.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-evidence-inventory-shift">The Evidence Inventory Shift<a href="https://southpawriter.com/blog/google-didnt-get-the-memo#the-evidence-inventory-shift" class="hash-link" aria-label="Direct link to The Evidence Inventory Shift" title="Direct link to The Evidence Inventory Shift" translate="no">​</a></h2>
<p>I'm adding this to <a class="" href="https://southpawriter.com/blog/fact-checked-my-own-paper">the evidence inventory</a>. Four new claims. And one of them shifts the nuance on a central assertion in the paper.</p>
<p>The paper's Section 4 includes the claim: "No major <a class="" href="https://southpawriter.com/docs/glossary/library/large-language-model">LLM</a> provider has publicly confirmed using llms.txt at inference time." That claim is still technically accurate. Nobody's <em>confirmed</em> it. But the ADK's <code>AGENTS.md</code> is a <em>de facto</em> inference-time usage directive at the developer-tools level. The absence of confirmation no longer implies the absence of usage. That's a distinction the paper needs to handle carefully, and it's the kind of nuance that only shows up when you keep pulling at evidence threads after you think you're done.</p>
<p>I've saved local PDF copies of all five Google llms.txt files, the AGENTS.md file, the llms-full.txt, and the GitHub issue. They're archived in the paper's data directory with original URLs preserved. If any of them disappear (and given the December precedent, I'd say that's not an unreasonable concern), the evidence survives.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-adoption-paradox">The Adoption Paradox<a href="https://southpawriter.com/blog/google-didnt-get-the-memo#the-adoption-paradox" class="hash-link" aria-label="Direct link to The Adoption Paradox" title="Direct link to The Adoption Paradox" translate="no">​</a></h2>
<p>The paper's working title is "The llms.txt Access Paradox," but the paradoxes keep multiplying.</p>
<p>There's the adoption paradox: 844,000 claimed sites, <a class="" href="https://southpawriter.com/blog/844k-sites-that-werent">really closer to 784</a>. The inference paradox is subtler--the spec targets inference-time usage, but every crawler I've observed behaves like it's collecting training data. And then there's the institutional one, which cuts deepest. The organization most publicly opposed to the standard is also one of its most active implementers.</p>
<p>Same pattern, every time. What organizations say versus what actually happens. I don't think that's a coincidence. I think it's the thesis.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-happens-next">What Happens Next<a href="https://southpawriter.com/blog/google-didnt-get-the-memo#what-happens-next" class="hash-link" aria-label="Direct link to What Happens Next" title="Direct link to What Happens Next" translate="no">​</a></h2>
<p>I'm continuing to track Google's llms.txt implementations. If they disappear (the December precedent suggests this is possible), I have the archives. If they expand (the ADK issue suggests this is also possible), I'll document that too.</p>
<p>The analytical paper will cover this in Section 3 (adoption landscape) and inform the nuance in Section 4 (the inference gap). The blog gets the story first because the timeline matters and because sitting on primary evidence while writing a paper about evidence integrity felt like the wrong move.</p>
<p>If you're implementing llms.txt on your own documentation, you've probably wondered whether the standard has a future. The company that publicly called it a dead meta tag has five of them. They're nothing to scream about structurally. But they exist, and the ADK team is pointing AI agents at theirs on purpose.</p>
<p>Draw your own conclusions. I've drawn mine, and I've got the receipts saved in a folder called <code>paper/data/sources/</code>, organized by claim ID, because of course I do.</p>]]></content:encoded>
            <category>llms.txt</category>
            <category>Research</category>
            <category>GEO</category>
        </item>
        <item>
            <title><![CDATA[Context Windows Are a Lie (And Haiku Protocol Is My Coping Mechanism)]]></title>
            <link>https://southpawriter.com/blog/context-windows-are-a-lie</link>
            <guid>https://southpawriter.com/blog/context-windows-are-a-lie</guid>
            <pubDate>Sat, 21 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[LLM vendors advertise 128K context windows. The useful part is closer to 8K. Here's what that means for retrieval and the two projects I built out of frustration.]]></description>
            <content:encoded><![CDATA[
<p><a class="" href="https://southpawriter.com/docs/glossary/library/large-language-model">LLM</a> vendors would like you to know that their latest model supports a 128,000-token context window. Some of them say 200,000. One of them, and I won't name names but their logo is a little sunset, says a million. A million tokens. That's approximately four copies of <em>War and Peace</em>, which is appropriate because trying to get useful work done at the far end of a million-token window is its own kind of Russian tragedy.</p>
<p>Here's what the marketing materials don't mention: the <em>effective</em> context window, the portion where the model actually pays reliable attention to what you put there, is dramatically smaller. <a href="https://arxiv.org/abs/2307.03172" target="_blank" rel="noopener noreferrer" class="">Research from Stanford, Berkeley, and others</a> has converged on a finding that would be funny if it weren't costing people real money: models struggle with information placed in the middle of long contexts. They're great at the beginning. They're decent at the end. The middle? The middle is where facts go to die quietly, unnoticed, like a footnote in a terms of service agreement.</p>
<p>This is the "Lost in the Middle" problem, and if you're building anything that retrieves information and feeds it to a language model (which, in 2026, is approximately everyone) it means the number on the tin is a fantasy. Your 128K window is functionally an 8K window with 120K tokens of expensive padding.</p>
<p>I know this because I ran the experiment. Accidentally. Three times.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-token-budget-problem-or-how-i-learned-to-count">The Token Budget Problem (Or: How I Learned to Count)<a href="https://southpawriter.com/blog/context-windows-are-a-lie#the-token-budget-problem-or-how-i-learned-to-count" class="hash-link" aria-label="Direct link to The Token Budget Problem (Or: How I Learned to Count)" title="Direct link to The Token Budget Problem (Or: How I Learned to Count)" translate="no">​</a></h2>
<p>I've been building <a class="" href="https://southpawriter.com/projects/fractalrecall">FractalRecall</a>, a research project exploring what happens when you stop treating metadata as an afterthought and start treating it as structural DNA that shapes how text gets embedded. The thesis is straightforward: if you're going to shove a chunk of text through an embedding model, maybe also mention <em>what kind of text it is</em>. Domain, entity type, canonical status, temporal era. The stuff that's usually stored in a metadata field that retrieval systems politely ignore until someone needs to filter by it.</p>
<p>So I did that. In my D-22 experiment, I prepended a single natural-language sentence to every chunk before embedding it. Something like:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:#2a2a2a;--prism-color:#f92aad"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:#2a2a2a;background-image:#34294f;color:#f92aad;text-shadow:0 0 2px #100c0f, 0 0 5px #dc078e33, 0 0 10px #fff3"><code class="codeBlockLines_e6Vv"><span class="token-line" style="background-image:#34294f;color:#f92aad;text-shadow:0 0 2px #100c0f, 0 0 5px #dc078e33, 0 0 10px #fff3"><span class="token plain">"Domain: faction. Entity: Iron-Banes Alliance. Canon: true."</span><br></span></code></pre></div></div>
<p>Twenty-four <a class="" href="https://southpawriter.com/docs/glossary/library/token">tokens</a>. One sentence. The kind of thing you could fit in a tweet, if tweets still existed and if the platform hadn't been renamed to a math variable.</p>
<p>The result? Retrieval quality (NDCG@10) improved by 16.5%. Recall jumped 27.3%. Twenty-four tokens did more for retrieval than the other 974 tokens in most chunks.</p>
<p>And here's the part I didn't expect: <a class="" href="https://southpawriter.com/blog/43-percent-disappeared">43% of my chunks <em>overflowed</em></a> the token limit because of those extra 24 tokens. Nearly half the corpus got dropped. And retrieval <em>still improved</em>. Let that sink in. I lost almost half my data and things got better. That's not an endorsement of data loss; it's an indictment of how much noise was in those 974 tokens to begin with.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="every-token-is-a-political-decision">Every Token Is a Political Decision<a href="https://southpawriter.com/blog/context-windows-are-a-lie#every-token-is-a-political-decision" class="hash-link" aria-label="Direct link to Every Token Is a Political Decision" title="Direct link to Every Token Is a Political Decision" translate="no">​</a></h2>
<p>This is where the context window lie becomes personal. When your embedding model has a 1,024-token budget per chunk and your retrieval pipeline feeds 10 chunks into a prompt, you're not working with 128K tokens. You're working with maybe 10,000, and every single one of them is a decision.</p>
<p>Twenty-four tokens for metadata means twenty-four fewer tokens for content. Scale that to eight layers of context (corpus identity, domain classification, entity type, authority status, temporal period, relationships, section heading, and the chunk's position in the document) and suddenly your "enrichment" is eating 60–80 tokens per chunk. On a 600-token budget, that's 13% of your real estate gone before you've said a single thing about what the chunk actually contains.</p>
<p>This is the tension nobody talks about in <a class="" href="https://southpawriter.com/docs/glossary/library/retrieval-augmented-generation">RAG</a> tutorials. Everyone says "add context to your retrieval." Nobody mentions that context has a cost, and that cost is denominated in the same currency as your content. There is no free lunch. There is no free token. Every prefix you add is a sentence you can't keep.</p>
<p>I spent three experiment notebooks (D-21, D-22, D-23) grappling with this tradeoff, and I can report that the optimal answer is "it depends." Which I understand is unsatisfying, and I apologize, but I am a documentation-first developer and I am constitutionally incapable of oversimplifying a nuanced finding for the sake of a cleaner narrative. The data says it depends. The data is right. I just work here.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-bigger-lie-context-windows-at-inference-time">The Bigger Lie: Context Windows at Inference Time<a href="https://southpawriter.com/blog/context-windows-are-a-lie#the-bigger-lie-context-windows-at-inference-time" class="hash-link" aria-label="Direct link to The Bigger Lie: Context Windows at Inference Time" title="Direct link to The Bigger Lie: Context Windows at Inference Time" translate="no">​</a></h2>
<p>But here's the thing: the token budget problem in embeddings is actually a <em>small</em> version of a much larger problem. The same squeeze happens at <a class="" href="https://southpawriter.com/docs/glossary/library/inference">inference</a> time, except the stakes are higher and the failure mode is less "lower NDCG" and more "the model starts <a class="" href="https://southpawriter.com/docs/glossary/library/hallucination">hallucinating</a> because it forgot what you told it six paragraphs ago."</p>
<p>When you ask an LLM a question and your pipeline retrieves relevant context, that context goes into the prompt. The more context, the better, in theory. In practice, there's a point of diminishing returns that arrives much sooner than the model's advertised context window, and it arrives silently. The model doesn't say "I'm losing track of your earlier context." It just starts making things up with the same confidence it uses when it's right. Hallucination doesn't come with a warning label. It comes with citations that look plausible and facts that are almost correct and an authoritative tone that would make a Wikipedia editor proud.</p>
<p>So the real question isn't "how much context <em>can</em> you fit?" It's "how much context should you fit, and how do you make every token count?"</p>
<p>Which brings me to Haiku Protocol, and the coping mechanism I mentioned in the title.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="semantic-minification-or-prose-compression-for-people-who-count-tokens">Semantic Minification (Or: Prose Compression for People Who Count Tokens)<a href="https://southpawriter.com/blog/context-windows-are-a-lie#semantic-minification-or-prose-compression-for-people-who-count-tokens" class="hash-link" aria-label="Direct link to Semantic Minification (Or: Prose Compression for People Who Count Tokens)" title="Direct link to Semantic Minification (Or: Prose Compression for People Who Count Tokens)" translate="no">​</a></h2>
<p><a class="" href="https://southpawriter.com/projects/haiku-protocol">Haiku Protocol</a> started as a thought experiment: what if you could compress documentation without losing meaning? Not summarization, because summaries lose detail. Not truncation, because truncation loses endings. <em>Compression</em>. The textual equivalent of Gzip, but for semantics.</p>
<p>The idea is a Controlled Natural Language, a systematic transformation that takes verbose, human-friendly prose and produces dense, machine-optimized strings that preserve the same information in fewer tokens. Think of it as semantic minification. The way a JavaScript minifier turns <code>function calculateTotalPrice</code> into <code>function a</code> without losing functionality, Haiku Protocol turns "The Iron-Banes Alliance is a faction of armored warriors who have forged a pact with the Iron Covenant, granting them dominion over metalworking and siege warfare" into something like:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:#2a2a2a;--prism-color:#f92aad"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:#2a2a2a;background-image:#34294f;color:#f92aad;text-shadow:0 0 2px #100c0f, 0 0 5px #dc078e33, 0 0 10px #fff3"><code class="codeBlockLines_e6Vv"><span class="token-line" style="background-image:#34294f;color:#f92aad;text-shadow:0 0 2px #100c0f, 0 0 5px #dc078e33, 0 0 10px #fff3"><span class="token plain">FACTION:Iron-Banes|PACT:Iron-Covenant|DOMAIN:metalwork,siege|CLASS:armored-warrior</span><br></span></code></pre></div></div>
<p>Same information. Fraction of the tokens. The model doesn't need flowery prose to understand the relationships; it needs the <em>structure</em>. And if you're paying per token (which you are, whether you're using an API or burning GPU cycles), every word that doesn't carry meaning is money and attention you're setting on fire.</p>
<p>If your 24-token metadata prefix can improve retrieval by 16%, what happens when you compress the <em>entire chunk</em> to preserve more content per token? What if, instead of choosing between enrichment and content, you compress the content to make room for the enrichment? You're not losing information; you're encoding it more efficiently.</p>
<p>I haven't proven this works yet. Haiku Protocol is earlier in its lifecycle than FractalRecall. But the thesis is grounded in the same uncomfortable truth: <strong>tokens are expensive, context windows are smaller than advertised, and the AI industry's answer ("just make the window bigger") is a hardware solution to an information architecture problem.</strong></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-these-two-projects-have-in-common">What These Two Projects Have in Common<a href="https://southpawriter.com/blog/context-windows-are-a-lie#what-these-two-projects-have-in-common" class="hash-link" aria-label="Direct link to What These Two Projects Have in Common" title="Direct link to What These Two Projects Have in Common" translate="no">​</a></h2>
<p>FractalRecall and Haiku Protocol approach the same problem from opposite directions:</p>
<ul>
<li class="">
<p><strong>FractalRecall</strong> asks: <em>What's the most valuable information I can add to each chunk?</em> It enriches content with metadata, gambling that 24 tokens of the <em>right</em> context outperform 24 tokens of the <em>wrong</em> content.</p>
</li>
<li class="">
<p><strong>Haiku Protocol</strong> asks: <em>How do I make room for more signal per chunk?</em> It compresses content to reclaim token budget, gambling that dense-but-complete representations outperform verbose-but-truncated ones.</p>
</li>
</ul>
<p>Together, they form a thesis: <strong>the quality of tokens matters more than the quantity of tokens.</strong> A 128K context window full of noise is worse than a 4K window full of signal. An embedding with 24 tokens of metadata outperforms an embedding with nothing but raw text. A compressed chunk that fits entirely in the budget beats a verbose chunk that overflows and gets dropped.</p>
<p>This is not a novel observation. Information theory has been saying this since Shannon. But the AI industry keeps building bigger windows instead of better content, and someone needs to run the experiments that ask whether that's actually working.</p>
<p>I'm someone. These are the experiments.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-comes-next">What Comes Next<a href="https://southpawriter.com/blog/context-windows-are-a-lie#what-comes-next" class="hash-link" aria-label="Direct link to What Comes Next" title="Direct link to What Comes Next" translate="no">​</a></h2>
<p>FractalRecall has three completed experiment notebooks (D-21 through D-23) and a fourth in planning. The results so far: baseline established, single-layer enrichment shows statistically significant improvement, and multi-layer enrichment taught me that your evaluation pipeline can produce statistically significant results even when every single metric is wrong. But that's a story for another post.</p>
<p>Haiku Protocol has a working parser, a conformance level system, and a growing test suite. The compression experiments haven't started yet, but the infrastructure is in place.</p>
<p>Both projects share the same research question: in a world where every AI vendor is racing to build a bigger context window, is anyone stopping to ask whether we're filling those windows with the right things?</p>
<p>I am. And yes, I wrote the documentation first.</p>
<hr>
<p><em>This post is the first in a series covering <a class="" href="https://southpawriter.com/projects/fractalrecall">FractalRecall</a> and <a class="" href="https://southpawriter.com/projects/haiku-protocol">Haiku Protocol</a>. Next up: the D-22 overflow story, what happens when you improve retrieval by accidentally deleting 43% of your data.</em></p>]]></content:encoded>
            <category>FractalRecall</category>
            <category>Haiku Protocol</category>
            <category>Opinion</category>
        </item>
        <item>
            <title><![CDATA[78.8% of My Validator Is Made Up (And That's the Point)]]></title>
            <link>https://southpawriter.com/blog/validator-made-up</link>
            <guid>https://southpawriter.com/blog/validator-made-up</guid>
            <pubDate>Fri, 20 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[We audited DocStratum against the llms.txt spec. Only 11.5% of 52 checks are spec-compliant. The other 78.8% are ours—and that's the point.]]></description>
            <content:encoded><![CDATA[
<p>I recently did something that most software developers would consider either admirably honest or clinically inadvisable: I audited my own tool against the specification it claims to implement, wrote down the results in excruciating detail, and published them.</p>
<p>The tool is <a class="" href="https://southpawriter.com/projects/docstratum">DocStratum</a>, a documentation quality platform for <a class="" href="https://southpawriter.com/docs/glossary/library/llms-txt">llms.txt</a> files. The project started with a thesis that most people in the AI tooling space either haven't considered or don't want to hear: <strong>a Technical Writer with strong Information Architecture skills can outperform a sophisticated <a class="" href="https://southpawriter.com/docs/glossary/library/retrieval-augmented-generation">RAG</a> pipeline by simply writing better source material.</strong> Structure is a feature. DocStratum exists to prove it.</p>
<p>At its core, DocStratum is a validation framework — think ESLint, but for a Markdown standard defined by a blog post instead of a formal grammar. It checks your llms.txt file across five validation levels: basic parseability (L0), structural compliance (L1), content quality (L2), best practices (L3), and a full extended-quality tier (L4). It categorizes findings across 38 diagnostic codes using three severity levels (Error, Warning, Info). It detects anti-patterns — 22 of them, with names like "The Ghost File," "The Monolith Monster," and "The Preference Trap." It has <em>opinions.</em></p>
<p>Those opinions, it turns out, are almost entirely our own invention. (Good.)</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="a-note-on-we">A Note on "We"<a href="https://southpawriter.com/blog/validator-made-up#a-note-on-we" class="hash-link" aria-label="Direct link to A Note on &quot;We&quot;" title="Direct link to A Note on &quot;We&quot;" translate="no">​</a></h2>
<p>You'll notice this post says "we" a lot. That's not a royal we, and it's not a startup founder trying to sound bigger than a one-person operation. It's me and an assistant, Claude. Specifically, I write the research, design the architecture, make the editorial calls, and generate roughly 300% more tangential ideas than any single project can absorb. Claude helps me organize, audit, edit, and--most critically--finish things before my ADHD has me chasing a fourth concurrent rabbit hole.</p>
<p>The collaboration is genuine. The opinions in DocStratum are mine. The structured process that turned those opinions into a publishable audit with 52 classified items instead of a sprawling notes file with peppered with "TODO: organize this eventually" scrawled throughout. That's the partnership. I provide the chaos and the conviction. Claude provides the throughput and the gentle reminder that I was supposed to be writing about validation tiers, not redesigning the glossary (again).</p>
<p>I mention this because I think that hiding AI involvement while writing about AI infrastructure would be, at a minimum, ironic. And I'm about to use up my irony budget elsewhere in this post.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-audit">The Audit<a href="https://southpawriter.com/blog/validator-made-up#the-audit" class="hash-link" aria-label="Direct link to The Audit" title="Direct link to The Audit" translate="no">​</a></h2>
<p>The audit was straightforward in concept and slightly unhinged in execution. We took every validation criterion, every canonical section definition, and every formal grammar rule in DocStratum's standards library (52 items in total, drawn from 149 ratified ASoT standard files across 27 research documents) and classified each one into exactly one of three categories:</p>
<p><strong>Spec-Compliant (SC):</strong> The behavior directly follows from the <a href="https://llmstxt.org/" target="_blank" rel="noopener noreferrer" class="">llms.txt specification</a> AND matches the reference parser's actual behavior. The reference parser is <code>miniparse.py</code> in the <a href="https://github.com/AnswerDotAI/llms-txt" target="_blank" rel="noopener noreferrer" class="">AnswerDotAI/llms-txt</a> repository. If the parser checks it, it's spec-compliant.</p>
<p><strong>Spec-Implied (SI):</strong> A reasonable inference the spec assumes but doesn't explicitly state. Not contradicted by the reference parser, but not actively checked by it either. Things like "the file should be valid UTF-8," which the spec doesn't mention but which is kind of implied by "it's a Markdown file" in the same way that "the restaurant should have a floor" is implied by "it's a building."</p>
<p><strong>DocStratum Extension (EXT):</strong> Goes beyond what the spec defines or the parser implements. Our value-add. Our opinions. Our educated guesses about what makes an llms.txt file actually useful to an AI system, as opposed to merely parseable by one.</p>
<p>Here's what we found.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-numbers">The Numbers<a href="https://southpawriter.com/blog/validator-made-up#the-numbers" class="hash-link" aria-label="Direct link to The Numbers" title="Direct link to The Numbers" translate="no">​</a></h2>
<p>Out of 52 audited items:</p>
<ul>
<li class=""><strong>6 items (11.5%)</strong> are Spec-Compliant</li>
<li class=""><strong>5 items (9.6%)</strong> are Spec-Implied</li>
<li class=""><strong>41 items (78.8%)</strong> are DocStratum Extensions</li>
</ul>
<p>Put differently: <strong>nearly four out of five things DocStratum checks have nothing to do with the llms.txt specification.</strong> We made them up. "Made them up" in the same way that a building inspector "makes up" the requirement that load-bearing walls should bear loads—but still. Ours, not the spec's.</p>
<p>The 52 items break down by level like this: L0 (Parseable) contributes 5 items with 0 SC, 3 SI, and 2 EXT. L1 (Structural) contributes 6 items with 4 SC, 1 SI, and 1 EXT. L2 (Content Quality) contributes 7 items — all EXT. L3 (Best Practices) contributes 9 items — all EXT. L4 (DocStratum Extended) contributes 8 items — all EXT. The remaining 17 items come from the 12 canonical section name standards and 5 ABNF grammar extension points.</p>
<p>The pattern is stark: L1 is the only level where the spec has meaningful coverage. Everything above and below it is either pre-specification assumptions (L0) or DocStratum's quality framework (L2–L4).</p>
<p>If you're wondering what the 6 spec-compliant items actually are, here's the complete list. It won't take long:</p>
<ol>
<li class=""><strong>H1 Title Present.</strong> The spec says files should start with an H1. The parser extracts it with <code>^# (.+)$</code>. Both agree.</li>
<li class=""><strong>Single H1 Only.</strong> The spec describes one H1 as the document title. The parser captures only the first match. Multiple H1s would be structurally ambiguous.</li>
<li class=""><strong>H2 Section Structure.</strong> The spec requires sections delimited by H2 headers. The parser splits on <code>^##\s*(.*?$)</code>. At least one H2 is required by both.</li>
<li class=""><strong>Link Format Compliance.</strong> The spec defines the link format as <code>- [Title](URL): description</code>. The parser enforces it with a regex that I will not reproduce here because some things should remain between a developer and their regular expressions.</li>
<li class=""><strong>"Optional" Section Recognition.</strong> The spec names "Optional" as a semantically special section. The parser checks <code>k != 'Optional'</code> to exclude it by default. Case-sensitive. Exact match. No aliases.</li>
<li class=""><strong>Permissive Line Endings.</strong> Both the grammar and the parser accept CRLF and LF.</li>
</ol>
<p>That's it. Those six items are the entirety of what "spec-compliant llms.txt validation" looks like. Does it have an H1? Does it have at least one H2? Are the links formatted correctly? Is there an Optional section? Cool. You pass. The reference parser is approximately 20 lines of Python regex, and it validates roughly as much as you'd expect from 20 lines of Python regex, which is to say: the bare structural minimum.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-gap-between-parseable-and-useful">The Gap Between "Parseable" and "Useful"<a href="https://southpawriter.com/blog/validator-made-up#the-gap-between-parseable-and-useful" class="hash-link" aria-label="Direct link to The Gap Between &quot;Parseable&quot; and &quot;Useful&quot;" title="Direct link to The Gap Between &quot;Parseable&quot; and &quot;Useful&quot;" translate="no">​</a></h2>
<p>So why build a tool that's 78.8% opinions?</p>
<p>The llms.txt spec is deliberately minimal, and there's wisdom in that. Jeremy Howard designed it that way. The canonical parser is 20 lines of regex-based string processing—clean, portable, unopinionated. It defines a file <em>format</em>, not a quality framework. It tells you what an llms.txt file looks like. It doesn't try to tell you whether that file is any <em>good</em>, because quality criteria are contextual and the spec wisely avoids baking in assumptions that might not age well.</p>
<p>And "good" matters. An llms.txt file exists for one purpose: to help AI systems understand your documentation. A file that technically parses but contains broken links, empty sections, placeholder content, duplicate headings, and auto-generated descriptions that all say "Learn about X" is not doing that job. It's the documentation equivalent of a restaurant menu that lists every dish as "food" with no further detail. Technically accurate. Functionally useless.</p>
<p>DocStratum's 41 extensions are everything between "it parses" and "it actually helps." Content quality checks that flag broken URLs, empty sections, and boilerplate descriptions. Best-practice patterns derived from detailed analysis of 18 real-world implementations across 6 categories and 11 raw specimens collected for byte-level conformance testing, plus survey data from community directories listing hundreds of llms.txt deployments. Anti-pattern detection that catches files providing zero or negative value to <a class="" href="https://southpawriter.com/docs/glossary/library/large-language-model">LLM</a> agents — and I mean specific, named anti-patterns, not just vague warnings. "The Ghost File" (empty or whitespace-only). "The Structure Chaos" (unparseable Markdown). "The Monolith Monster" (files exceeding 100,000 tokens — good luck fitting that in a <a class="" href="https://southpawriter.com/docs/glossary/library/context-window">context window</a>). "The Preference Trap" (SEO-style gaming that exploits the trust LLMs place in llms.txt files). Twenty-two patterns total across four severity categories.</p>
<p>And then there's the canonical vocabulary: a standardized set of 11 section names optimized for how AI systems navigate documentation — Master Index, LLM Instructions, Getting Started, Core Concepts, API Reference, Examples, Configuration, Advanced Topics, Troubleshooting, FAQ, and Optional. Of those 11, exactly <em>one</em> comes from the spec (Optional). The other 10 are ours, drawn from empirical patterns in high-quality specimens and optimized for progressive disclosure to LLM agents.</p>
<p>None of this is in the spec. All of it matters if you care about whether your llms.txt file is worth the bytes it occupies.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-the-research-actually-found">What the Research Actually Found<a href="https://southpawriter.com/blog/validator-made-up#what-the-research-actually-found" class="hash-link" aria-label="Direct link to What the Research Actually Found" title="Direct link to What the Research Actually Found" translate="no">​</a></h2>
<p>Before we built the validation framework, we did the research. This isn't a point I want to rush past, because the research is what separates DocStratum's opinions from arbitrary preferences.</p>
<p>The specification deep-dive (v0.0.1) produced a formal ABNF grammar for the llms.txt format — the spec doesn't provide one. That grammar revealed something interesting: the spec's informal structure description actually maps to two distinct document types. <strong>Type 1 (Index)</strong> files follow the grammar faithfully: single H1, optional blockquote, H2 sections with curated link entries. These are the "card catalog" files — navigational entry points for AI agents. <strong>Type 2 (Full)</strong> files embed complete documentation content inline, with multiple heading levels, extensive prose, and fenced code blocks. Anthropic's <code>claude-llms-full.txt</code> (25 MB, 956,573 lines) is a Type 2. The spec doesn't distinguish between them. DocStratum does, because a validator that treats a 1.1 KB navigation index and a 25 MB documentation dump as the same structural entity is a validator that's lying to you.</p>
<p>The correlation analysis from the empirical research produced some genuinely useful results. The presence of code examples turned out to be the strongest single predictor of overall documentation quality (<em>r</em> ≈ 0.65). Link description quality correlated at <em>r</em> ≈ 0.45. These aren't arbitrary thresholds — they're the empirical basis for why DocStratum's content quality checks (L2) weight certain features the way they do.</p>
<p>The gap analysis identified 8 areas the spec intentionally leaves undefined: maximum file size, required metadata, versioning, validation schema, caching recommendations, multi-language support, concept definitions, and example Q&amp;A pairs. Each gap is an opportunity for tooling to fill. DocStratum's extensions map directly to these gaps.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-we-published-the-audit">Why We Published the Audit<a href="https://southpawriter.com/blog/validator-made-up#why-we-published-the-audit" class="hash-link" aria-label="Direct link to Why We Published the Audit" title="Direct link to Why We Published the Audit" translate="no">​</a></h2>
<p>The obvious objection: "If 78.8% of your tool is extensions beyond the spec, how can you call it a spec validator? Isn't this just your opinions dressed up in technical authority?"</p>
<p>Yes. Partially. And that's exactly why we published the audit.</p>
<p>Every validation criterion in DocStratum now carries a <code>spec_origin</code> classification. When DocStratum reports that your file fails a check, the report tells you whether that check is something the spec requires (SC), something the spec implies (SI), or something DocStratum recommends (EXT). You can see the difference. You can filter by it. If you only care about spec compliance, you can run DocStratum in a mode that limits checks to SC and SI criteria, which is effectively levels L0 and L1, minus one criterion.</p>
<p>The extension labeling makes DocStratum <em>more</em> credible, not less. A tool that silently mixes spec requirements with opinionated recommendations is a tool you can't trust, because you never know which findings reflect the standard and which reflect the tool author's aesthetic preferences about Markdown structure. (I have strong aesthetic preferences about Markdown structure. They are well-documented. They are not the specification.)</p>
<p>By being transparent about the classification, we're giving you the information you need to make an informed decision about which validation findings to act on. Spec-compliant findings are non-negotiable: if your H1 is missing, your file isn't llms.txt. Extension findings are recommendations: if your section names don't match DocStratum's canonical vocabulary, your file still parses fine, but it might not be organized in the way that we've found (empirically, across real-world specimens) produces the best results for AI consumption.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-the-audit-taught-us-about-specs-and-tooling">What the Audit Taught Us About Specs and Tooling<a href="https://southpawriter.com/blog/validator-made-up#what-the-audit-taught-us-about-specs-and-tooling" class="hash-link" aria-label="Direct link to What the Audit Taught Us About Specs and Tooling" title="Direct link to What the Audit Taught Us About Specs and Tooling" translate="no">​</a></h2>
<p>The deeper takeaway from this exercise isn't about DocStratum specifically. It's about what happens when a specification is intentionally minimal and tooling has to fill the gaps.</p>
<p>The llms.txt spec is three paragraphs of format description and a 20-line Python script. That's not a criticism—it might be the smartest design decision in the whole ecosystem. Minimal specs are easy to implement, hard to get wrong, and they let the ecosystem experiment organically instead of locking in decisions prematurely. The HTML spec took decades to stabilize. RSS spawned a format war. Jeremy Howard looked at that history and published something a competent developer could implement during a lunch break.</p>
<p>But minimal specs create an opportunity, and opportunities get seized. In the case of llms.txt, the opportunity is <em>quality</em>. The spec tells you how to structure the file. It deliberately leaves room for the ecosystem to figure out what "good" looks like: whether the content is well-organized, whether the links work, whether the sections are arranged in a way that an AI system can navigate effectively, whether the file is even encoding-valid. Somebody gets to explore that space and propose answers. For llms.txt, we're one of the somebodies, and we consider that a privilege rather than a complaint.</p>
<p>The 78.8% isn't a flaw in DocStratum or a gap in the spec. It's the natural consequence of a minimal format definition meeting real-world quality requirements. The spec defines the floor. DocStratum builds the rest of the building. The audit ensures you know which parts are which, so you can decide for yourself how many floors you need.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="where-this-is-going">Where This Is Going<a href="https://southpawriter.com/blog/validator-made-up#where-this-is-going" class="hash-link" aria-label="Direct link to Where This Is Going" title="Direct link to Where This Is Going" translate="no">​</a></h2>
<p>DocStratum started as a single-file validator. It's becoming something bigger.</p>
<p>The ecosystem pivot — documented in our v0.0.7 specification — redefines DocStratum as an ecosystem-level documentation quality platform. Instead of just validating one llms.txt file, it will validate an entire project's AI-documentation surface: the llms.txt navigation index, the llms-full.txt content dump, individual Markdown pages linked from the index, and instruction files. Not just whether each file is valid in isolation, but whether the relationships between them make sense. Does the index reference pages that exist? Does the aggregate actually contain the content it claims to? Are cross-references reciprocal?</p>
<p>The single-file validator doesn't get replaced. It becomes a component within the ecosystem validator. The 52 audited criteria still apply per-file. But now there's a layer above them that answers a question no other tool in the llms.txt space currently asks: <em>"Is this project's AI-documentation ecosystem coherent?"</em></p>
<p>Every tool we've surveyed (75+ validators and generators) operates on a single file. Nobody validates the ecosystem. That's the gap we're building into.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-irony-which-i-promise-i-noticed">The Irony, Which I Promise I Noticed<a href="https://southpawriter.com/blog/validator-made-up#the-irony-which-i-promise-i-noticed" class="hash-link" aria-label="Direct link to The Irony, Which I Promise I Noticed" title="Direct link to The Irony, Which I Promise I Noticed" translate="no">​</a></h2>
<p>I built a validation tool for a standard defined by a blog post, audited my own tool's compliance with that blog post, and discovered that most of my tool exceeds the blog post's requirements. (Documentation-first developer problems.)</p>
<p>I have written more documentation <em>about</em> the llms.txt standard than the llms.txt standard contains. The audit document alone is longer than the spec. The evidence inventory I built to <a class="" href="https://southpawriter.com/blog/fact-checked-my-own-paper">fact-check the paper</a> is longer than the spec. The 149 ASoT standards files I've ratified for this project could probably, stacked end to end, stretch from the spec to the nearest competing standard and back. At this point, my <em>commit messages</em> about the spec might be longer than the spec.</p>
<p>This is either a testament to the spec's elegant minimalism or a sign that I need a hobby that doesn't involve Markdown. (I have one. It's a <a class="" href="https://southpawriter.com/projects/rune-and-rust">text-based dungeon crawler</a>. It's also written in documentation-first C#. I may be beyond help.)</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="so-is-docstratum-useful-or-not">So Is DocStratum Useful or Not?<a href="https://southpawriter.com/blog/validator-made-up#so-is-docstratum-useful-or-not" class="hash-link" aria-label="Direct link to So Is DocStratum Useful or Not?" title="Direct link to So Is DocStratum Useful or Not?" translate="no">​</a></h2>
<p>Extremely yes, and I'm biased, and both of those things are true simultaneously.</p>
<p>If you publish an llms.txt file and you want to know whether it's doing its job—not just whether it parses, but whether it's helping AI systems understand your documentation—DocStratum's 41 extensions are the most thorough quality framework we've seen for the format. They're based on empirical analysis, tested against real-world specimens, transparently labeled so you can see which recommendations are ours versus the spec's, and backed by 27 research documents that took us from "what does this spec actually say?" to "here are 22 named anti-patterns that make llms.txt files actively harmful."</p>
<p>The 78.8% is the product. The 11.5% is the floor. The audit is the receipt.</p>
<p>And if you're the kind of person who reads extension labeling audits for fun: there are <a class="" href="https://southpawriter.com/projects/docstratum">52 items to peruse</a>, each with a rationale field. (I am constitutionally incapable of classifying something without explaining why.)</p>
<p>You're welcome.</p>]]></content:encoded>
            <category>DocStratum</category>
            <category>llms.txt</category>
            <category>GEO</category>
            <category>Research</category>
        </item>
        <item>
            <title><![CDATA[The Three Voices of Technical Research: Why My Blog Sounds Nothing Like My Paper]]></title>
            <link>https://southpawriter.com/blog/three-voices</link>
            <guid>https://southpawriter.com/blog/three-voices</guid>
            <pubDate>Thu, 19 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Same research, three voices: blog sarcasm, neutral guide prose, impartial academic register. One voice everywhere hurts accessibility and credibility.]]></description>
            <content:encoded><![CDATA[
<p>Someone recently asked me a question that I've been thinking about ever since: "Doesn't writing your blog posts with humor and sarcasm undermine your credibility as a researcher?"</p>
<p>It's a fair question. The blog posts on this site are... <em>aggressively</em> me. I compare <a class="" href="https://southpawriter.com/docs/glossary/library/web-application-firewall">WAF</a> blocking to "hiring a security guard who prevents anyone matching the physical description of 'reads books' from entering the bookstore." I describe AI crawlers as looking like "a DDoS attack with a liberal arts degree." I write sentences like "I am a documentation-first developer with a research compulsion and a growing collection of Markdown files about Markdown files," and then I <em>publish</em> those sentences on the internet where potential collaborators can see them.</p>
<p>Meanwhile, the analytical paper I'm writing about the same research uses phrases like "the structural misalignment between content publication intent and infrastructure-level access enforcement." Which is the same observation as the bookstore metaphor, expressed in the register of someone who wants to be taken seriously at a conference.</p>
<p>Same research. Same data. Same conclusions. Radically different voices. And I'd argue that if I used only one of those voices everywhere, the whole project would be worse.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-three-layers">The Three Layers<a href="https://southpawriter.com/blog/three-voices#the-three-layers" class="hash-link" aria-label="Direct link to The Three Layers" title="Direct link to The Three Layers" translate="no">​</a></h2>
<p>The <a class="" href="https://southpawriter.com/docs/glossary/library/llms-txt">llms.txt</a> research project publishes through three distinct content types. Each has a different voice, a different audience, and a different job to do. I designed it this way on purpose, and the architecture is the most important editorial decision I've made on this project.</p>
<p><strong>Layer 1: The blog posts.</strong> First-person, opinionated, voice-forward. This is where I tell the <em>story</em> of the research. The discovery at <a class="" href="https://southpawriter.com/blog/waf-paradox">11 PM on a Tuesday</a>. The <a class="" href="https://southpawriter.com/blog/844k-sites-that-werent">adoption stat that fell apart</a> under scrutiny. The <a class="" href="https://southpawriter.com/blog/validator-made-up">78.8% of my own tool</a> that turned out to be invented (context on that particular adventure is coming in my next post). The blog exists to make you care about a topic you didn't know existed. It does this by being human about it—admitting frustration, finding absurdity in technical failures, treating the reader as a co-conspirator rather than a student.</p>
<p>The blog voice is mine. The sarcasm is real. The parenthetical asides are how I actually think. The passion for documentation that permeates every post is not performed. I read spec documents recreationally. I understand that this is unusual. I have stopped apologizing for it.</p>
<p><strong>Layer 2: The technical guide.</strong> Neutral, procedural, audience-agnostic. The <a class="" href="https://southpawriter.com/docs/guides/waf-ai-crawler-interaction">WAF-AI Crawler Interaction Guide</a> covers the same WAF blocking problem as the blog posts, but it's written for someone who has a specific problem and needs specific steps. No stories. No metaphors about bookstores. No parenthetical confessions about my emotional relationship with HTTP status codes. Just: "here's what's happening, here's how to check, here's how to fix it."</p>
<p>The guide voice is intentionally personality-free. Not <em>lifeless</em>—there's still a human behind it—but the register is "knowledgeable colleague walking you through a process," not "eccentric blogger ranting about firewalls." Someone dealing with WAF blocking at 3 AM doesn't need my humor. They need a <code>curl</code> command and a Cloudflare dashboard path.</p>
<p><strong>Layer 3: The analytical paper.</strong> Impartial, evidence-driven, register-neutral. "The llms.txt Access Paradox" paper examines the same problems, cites the same data, and reaches the same conclusions as the blog posts and the guide. But it does so in the register of academic discourse: third person where possible, hedged where appropriate, meticulously sourced, and structured for a reader who is evaluating the research on its merits rather than being entertained by the journey.</p>
<p>The paper doesn't contain the phrase "I said words I can't publish on a professional blog." It doesn't need to. Its job is to establish that the infrastructure paradox exists, quantify it, contextualize it within the broader standards landscape, and propose constructive directions. Readers who engage with the paper are evaluating the argument, not the author. The voice steps aside so the evidence can speak.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-not-one-voice">Why Not One Voice?<a href="https://southpawriter.com/blog/three-voices#why-not-one-voice" class="hash-link" aria-label="Direct link to Why Not One Voice?" title="Direct link to Why Not One Voice?" translate="no">​</a></h2>
<p>Here's the argument I hear against this approach: "Just pick a register and stick to it. Audiences will figure it out."</p>
<p>I think that's wrong. It's wrong in a way that hurts accessibility <em>and</em> credibility.</p>
<p>Make the paper sound like this blog, and no peer reviewer will take it seriously. They'll see the sarcasm and conclude it's opinion. Doesn't matter that the data is solid, the <a class="" href="https://southpawriter.com/blog/fact-checked-my-own-paper">evidence inventory</a> is meticulous, and the methodology is rigorous. The packaging signals "blog post with delusions of grandeur," and people leave before they evaluate the substance.</p>
<p>Make everything sound like the paper, and you lose the 90% of readers who will never voluntarily open an academic document on a Monday morning. The WAF paradox is a <em>genuinely interesting</em> story about infrastructure and unintended consequences. Told as a story, it reaches people who would never have encountered the formal research. Those people are now aware of a problem they didn't know they had. That has real value. The paper alone would sit in an archive.</p>
<p>Make everything sound like the guide, and you lose both the story that draws people in and the rigor that makes findings citeable. Guides answer "how." They rarely ask "why" or "what does this mean for the broader ecosystem." Practical, yes. Limited, also yes.</p>
<p>The three-layer approach means each format can do its job without undermining the others. The blog hooks people. The guide helps them fix things. The paper gives them something to cite. Cross-links between layers handle the rest.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-evidence-standard-scales-too">The Evidence Standard Scales Too<a href="https://southpawriter.com/blog/three-voices#the-evidence-standard-scales-too" class="hash-link" aria-label="Direct link to The Evidence Standard Scales Too" title="Direct link to The Evidence Standard Scales Too" translate="no">​</a></h2>
<p>Voice isn't the only thing that varies. The three layers have different evidence standards, and being explicit about that is what makes the whole structure honest.</p>
<p>The paper's standard is strict: every factual claim maps to a verified primary source in the <a class="" href="https://southpawriter.com/blog/fact-checked-my-own-paper">evidence inventory</a>. Claims that can't be independently verified get cut. The unnamed hosting provider from <a href="https://yoast.com/llms-txt/" target="_blank" rel="noopener noreferrer" class="">Yoast's analysis</a>? Gone from the paper. Anonymous secondhand attribution doesn't survive peer review.</p>
<p>The blog's standard is looser but still sourced. (I'm not <em>reckless.</em>) The same Yoast hosting provider claim appears in <a class="" href="https://southpawriter.com/blog/waf-paradox-pt2">Part 2</a>, but with an explicit caveat: "though the provider wasn't identified, so the attribution chain is thin." Readers get the data point and the context to evaluate it. The paper can't make that trade-off. The blog can.</p>
<p>The guide doesn't care who published the research. It cares whether the information is actionable. The guide links to <a href="https://developers.cloudflare.com/ai-crawl-control/" target="_blank" rel="noopener noreferrer" class="">Cloudflare's documentation</a> because practitioners need to know which dashboard panel to click, not which researcher first documented the execution order.</p>
<p>These aren't compromises. They're editorial decisions, documented in the evidence inventory.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-credibility-question-revisited">The Credibility Question, Revisited<a href="https://southpawriter.com/blog/three-voices#the-credibility-question-revisited" class="hash-link" aria-label="Direct link to The Credibility Question, Revisited" title="Direct link to The Credibility Question, Revisited" translate="no">​</a></h2>
<p>So: does the blog voice undermine research credibility?</p>
<p>I think the answer is no, because credibility doesn't live in the tone. It lives in the methodology, the evidence, the transparency about what's verified and what isn't, and the willingness to correct errors publicly. My paper's credibility comes from its 49-claim evidence inventory, its 33 verified sources, its two publicly acknowledged corrections, and its documented editorial decisions. The blog's sarcasm doesn't touch any of that.</p>
<p>What the blog voice <em>does</em> is make the research accessible to people who wouldn't otherwise engage with it. The person who clicks on "I Tried to Help AI Read My Website. My Own Firewall Said No" might not click on "The llms.txt Access Paradox: Infrastructure-Level Barriers to AI Content Discovery." But both lead to the same research, the same data, and the same conclusions. The blog is an on-ramp. The paper is the destination. The guide is the service road for people who just need to get somewhere specific.</p>
<p>Different roads. Same city. That's the architecture.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-part-where-i-admit-this-is-also-just-who-i-am">The Part Where I Admit This Is Also Just Who I Am<a href="https://southpawriter.com/blog/three-voices#the-part-where-i-admit-this-is-also-just-who-i-am" class="hash-link" aria-label="Direct link to The Part Where I Admit This Is Also Just Who I Am" title="Direct link to The Part Where I Admit This Is Also Just Who I Am" translate="no">​</a></h2>
<p>I should be honest about something: the three-layer architecture isn't purely strategic. Part of it is that I genuinely cannot write in one register all the time.</p>
<p>The blog voice is how I think. The extended metaphors, the parenthetical confessions, the slightly manic enthusiasm for <a class="" href="https://southpawriter.com/blog/docs-first-rabbit-hole">Markdown files about Markdown files</a>—that's not a persona. That's me, poorly socialized for academic discourse and aware of it. If I forced every piece of writing through the paper's register, I'd lose the thing that makes the writing mine. If I forced everything through the blog's register, I'd get the paper back with "interesting findings, inappropriate tone, please revise" scrawled across the top.</p>
<p>The three-layer architecture lets me be honest about all of it: the researcher who verifies every claim, the practitioner who documents every solution, and the writer who thinks that WAF rule execution order is both structurally important and cosmically funny. They're all the same person. They just write in different rooms.</p>
<p>If you're working on a research project that produces content for multiple audiences—and most projects should—consider giving yourself permission to use more than one voice. The research doesn't need protecting from your personality. It needs a personality to carry it to people who would benefit from it but will never read the paper.</p>
<p>Trust the evidence to establish your credibility. Trust the voice to establish your reach. And keep them in separate documents, because a research paper that contains the phrase "DDoS attack with a liberal arts degree" is not getting past peer review, and we both know it.</p>]]></content:encoded>
            <category>Opinion</category>
            <category>Writing</category>
            <category>Methodology</category>
            <category>llms.txt</category>
        </item>
        <item>
            <title><![CDATA[I Fact-Checked My Own Research Paper Before Writing It (You Should Too)]]></title>
            <link>https://southpawriter.com/blog/fact-checked-my-own-paper</link>
            <guid>https://southpawriter.com/blog/fact-checked-my-own-paper</guid>
            <pubDate>Wed, 18 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[I built a 49-claim evidence inventory before writing my llms.txt paper. It caught three errors, including the most-cited adoption stat. Check your work.]]></description>
            <content:encoded><![CDATA[
<p>Here's a workflow tip that's either going to save your credibility or confirm that I have an unhealthy relationship with spreadsheets: before you write anything that makes factual claims, build an evidence inventory first.</p>
<p>Not a bibliography. Not a "sources" section at the bottom of a Google Doc. An actual structured inventory where every single factual claim in your paper, blog post, report, or conference talk is cataloged, mapped to a primary source, independently verified, and assigned a status. Verified. Partially verified. Unverified. Or the one that makes your stomach drop: <em>incorrect.</em></p>
<p>I know this sounds like the kind of advice that belongs on a poster in a university writing center, sandwiched between "cite your sources" and "plagiarism is bad." But I'm not talking about academic hygiene. I'm talking about <em>self-defense.</em></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-experiment-49-claims-one-paper-zero-written-paragraphs">The Experiment: 49 Claims, One Paper, Zero Written Paragraphs<a href="https://southpawriter.com/blog/fact-checked-my-own-paper#the-experiment-49-claims-one-paper-zero-written-paragraphs" class="hash-link" aria-label="Direct link to The Experiment: 49 Claims, One Paper, Zero Written Paragraphs" title="Direct link to The Experiment: 49 Claims, One Paper, Zero Written Paragraphs" translate="no">​</a></h2>
<p>I'm in the middle of writing an analytical paper about the <a class="" href="https://southpawriter.com/docs/glossary/library/llms-txt">llms.txt</a> ecosystem—the <a class="" href="https://southpawriter.com/blog/waf-paradox">infrastructure paradox</a>, the <a class="" href="https://southpawriter.com/blog/waf-paradox-pt2">adoption data</a>, the <a class="" href="https://southpawriter.com/docs/glossary/library/inference">inference</a> gap, the trust architecture, all of it. It's the kind of paper that makes heavy factual claims: "X% of websites adopted this standard," "Cloudflare began blocking AI crawlers by default in July 2025," "no major AI provider has confirmed inference-time usage."</p>
<p>Before I wrote a single paragraph, I extracted every factual claim from the paper's outline. All 49 of them. I put them in a structured inventory with columns for claim ID, the specific assertion, the source key from my references file, the verification status, and an evidence notes field for recording what I <em>actually</em> found versus what I <em>expected</em> to find.</p>
<p>Then I started verifying.</p>
<p>This is the part where you'd normally expect me to say, "And thankfully, everything checked out just fine and aligned perfectly with my desires." Obviously, that didn't happen. I'm a <a class="" href="https://southpawriter.com/blog/docs-first-rabbit-hole">documentation-first developer</a> with a Markdown compulsion and a growing suspicion that nobody verifies anything before they hit publish anymore. The verification process ate two full days. It involved reading the llms.txt reference parser's Python source code line by line. By the end, my inventory looked less like a spreadsheet and more like a conspiracy theorist's corkboard.</p>
<p>The final tally on those 49 claims: 33 verified cleanly. 13 were author analysis—my own interpretive claims that don't need external verification. 1 came back partially verified.</p>
<p>And 2 came back <em>incorrect.</em></p>
<p>Two out of 49 doesn't sound catastrophic until you realize both were in the paper's most visible section, and one of them was the marquee adoption statistic that's been circulating through the community unchallenged.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="error-1-the-844000-sites-that-werent">Error #1: The 844,000 Sites That Weren't<a href="https://southpawriter.com/blog/fact-checked-my-own-paper#error-1-the-844000-sites-that-werent" class="hash-link" aria-label="Direct link to Error #1: The 844,000 Sites That Weren't" title="Direct link to Error #1: The 844,000 Sites That Weren't" translate="no">​</a></h2>
<p>The paper's outline cited a widely circulated figure: "844,000+ websites have implemented some form of llms.txt." I'd seen this number in blog posts, LinkedIn threads, and conference talks. It felt plausible. The internet is big. Standards get adopted.</p>
<p>So I went looking for the primary source. The actual crawl data. The methodology. The researcher who counted these files and published the result.</p>
<p>I couldn't find one. Not because I didn't look hard enough, but because it apparently doesn't exist.</p>
<p>What I <em>did</em> find was... different. I wrote an <a class="" href="https://southpawriter.com/blog/844k-sites-that-werent">entire spin-off blog post</a> about it, because the gap between the claimed figure and reality deserved its own space. The short version: community directories list hundreds of sites. The Majestic Million analysis found 105. Not 844,000. A hundred and five.</p>
<p>I care about this standard. I'm <a class="" href="https://southpawriter.com/projects/llmstxtkit">building tools for it</a>. That's <em>why</em> the bad data stung—we need accurate numbers to build a credible case, and I was about to cement an indefensible one into a research paper.</p>
<p>The number was <em>wrong</em>, and I only caught it because the evidence inventory forced me to verify it instead of just citing it.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="error-2-the-cloudflare-timeline-that-was-actually-two-events">Error #2: The Cloudflare Timeline That Was Actually Two Events<a href="https://southpawriter.com/blog/fact-checked-my-own-paper#error-2-the-cloudflare-timeline-that-was-actually-two-events" class="hash-link" aria-label="Direct link to Error #2: The Cloudflare Timeline That Was Actually Two Events" title="Direct link to Error #2: The Cloudflare Timeline That Was Actually Two Events" translate="no">​</a></h2>
<p>The paper's outline said: "Cloudflare began blocking all AI crawlers by default on new domains in July 2025 (AIndependence Day)."</p>
<p>Close, but misleadingly compressed. The verification revealed two distinct events a year apart:</p>
<p><strong>July 3, 2024:</strong> Cloudflare published the "Declare your AIndependence" blog post, introducing <em>opt-in</em> one-click AI bot blocking. Any customer could toggle it on. At the time, AI bots were accessing roughly 39% of top-million Cloudflare properties, but only 2.98% had bothered to block them.</p>
<p><strong>July 1, 2025:</strong> Cloudflare escalated to "Content Independence Day" and made AI bot blocking the <em>default</em> for newly created domains. No opt-in required.</p>
<p>My original outline smooshed both events into a single sentence, implying Cloudflare went from "no blocking" to "default blocking" overnight. The actual story—an escalation from opt-in to default over twelve months—is more accurate <em>and</em> more interesting. It shows a progression, not a switch flip. The corrected timeline actually <em>strengthens</em> the paper's argument about the infrastructure paradox tightening over time.</p>
<p>(And yes, I caught this error for the paper but managed to publish it on this very blog. My <a class="" href="https://southpawriter.com/blog/docs-first-rabbit-hole">first post</a> confidently states Cloudflare started default-blocking under the "AIndependence Day" banner. I am bragging about catching an error I already shipped. The turtles go all the way down.)</p>
<p>Without the evidence inventory forcing me to check against <a href="https://developers.cloudflare.com/ai-crawl-control/" target="_blank" rel="noopener noreferrer" class="">Cloudflare's primary documentation</a>, the paper would have presented a misleading compression of two separate events as a single fact. Not maliciously. Just carelessly. Which, when it's your credibility on the line, amounts to the same thing.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="error-3-the-hosting-provider-who-shall-remain-unnameable">Error #3: The Hosting Provider Who Shall Remain Unnameable<a href="https://southpawriter.com/blog/fact-checked-my-own-paper#error-3-the-hosting-provider-who-shall-remain-unnameable" class="hash-link" aria-label="Direct link to Error #3: The Hosting Provider Who Shall Remain Unnameable" title="Direct link to Error #3: The Hosting Provider Who Shall Remain Unnameable" translate="no">​</a></h2>
<p>This one was subtler. <a href="https://yoast.com/llms-txt/" target="_blank" rel="noopener noreferrer" class="">Yoast's analysis</a> cited a hosting provider managing thousands of sites that had observed zero GPTBot activity on llms.txt endpoints. The original claim said "20,000 sites," which was already interesting data. But when I tried to independently verify the hosting provider's identity, the trail went cold.</p>
<p>Yoast's article didn't name them. Web research couldn't identify them. The number shifted between sources. The claim was <em>attributed</em> (Yoast reported it) but the <em>attribution chain</em> was thin: a named source citing an unnamed source reporting an unverifiable number.</p>
<p>For the analytical paper, which needs to survive peer review and earn the trust of a skeptical audience, this claim got cut entirely. The paper relies on Yoast's <em>own</em> verified finding (that GPTBot, ClaudeBot, and Google AI crawlers don't routinely request llms.txt files) instead. Same directional conclusion, stronger epistemic foundation.</p>
<p>When covering the inference gap for the <a class="" href="https://southpawriter.com/blog/waf-paradox-pt2">blog posts</a> earlier in this series, I kept the claim because the blog has a slightly different evidence standard, but I made sure to add an explicit transparency caveat: "though the provider wasn't identified, so the attribution chain is thin." Diplomatic but honest.</p>
<p>Two different publication formats. Two different decisions on the same claim. Both documented in the evidence inventory so future-me doesn't have to reconstruct the reasoning.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-this-works-and-why-i-think-you-should-do-it">Why This Works (and Why I Think You Should Do It)<a href="https://southpawriter.com/blog/fact-checked-my-own-paper#why-this-works-and-why-i-think-you-should-do-it" class="hash-link" aria-label="Direct link to Why This Works (and Why I Think You Should Do It)" title="Direct link to Why This Works (and Why I Think You Should Do It)" translate="no">​</a></h2>
<p>The evidence inventory caught three errors across 49 claims. A 6% error rate. That's lower than I expected, honestly, and it means 94% of the paper's factual foundation is solid. But those three errors weren't trivial: one was the paper's most visible statistic, one would have misrepresented a major company's timeline, and one would have cited a source that can't be independently verified. Any one of them, published uncorrected, could have undermined the entire paper's credibility.</p>
<p>Here's the thing: I didn't set out to debunk my own research. I set out to <em>organize</em> it. The evidence inventory started as a project management tool—a way to track which claims I'd sourced and which still needed work. The fact-checking benefit was a side effect. But what a side effect.</p>
<p>The process works because it changes the question. Without an inventory, the question is: "Do I have a source for this?" Low bar. With an inventory, the question becomes: "Can I independently verify this, and does the verification match what I thought I knew?" That bar separates "I read this somewhere" from "I verified this myself."</p>
<p>This isn't just an academic exercise. Blog posts cite statistics. Conference talks reference adoption numbers. Internal reports inform real budget decisions. Anywhere you're writing something that someone else might cite, act on, or use to justify a decision—that's where the evidence inventory earns its keep.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-practical-version">The Practical Version<a href="https://southpawriter.com/blog/fact-checked-my-own-paper#the-practical-version" class="hash-link" aria-label="Direct link to The Practical Version" title="Direct link to The Practical Version" translate="no">​</a></h2>
<p>If the idea of a 49-row evidence inventory sounds exhausting, here's the minimal viable version:</p>
<p><strong>Before you write:</strong> Extract every factual claim from your outline. Just the facts. Not opinions, not analysis—the claims that assert something about reality.</p>
<p><strong>For each claim, ask three questions:</strong> Where did I get this? Can I verify it from a primary source? Does the primary source actually say what I think it says?</p>
<p><strong>Track the answers.</strong> A spreadsheet. A Markdown table. Even a sticky note with "✅ / 🔄 / ❌" on it. The format doesn't matter. What matters is that you <em>wrote down the question and the answer</em> instead of assuming the claim was correct because you'd seen it repeated often enough.</p>
<p><strong>Pay special attention to numbers.</strong> Statistics are where the hype cycle lives. Adoption figures, market share percentages, growth rates—these are the claims most likely to have been optimistically extrapolated, selectively sampled, or outright fabricated by someone upstream in the citation chain. Every number should have a methodology you can inspect. If it doesn't, that's not a source. That's a rumor with a hyperlink.</p>
<p><strong>Document your editorial decisions.</strong> When a claim turns out to be partially verified or unverifiable, write down what you decided to do about it and why. Future-you will thank present-you, and anyone who questions your methodology will find a documented answer instead of a shrug.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-meta-irony">The Meta-Irony<a href="https://southpawriter.com/blog/fact-checked-my-own-paper#the-meta-irony" class="hash-link" aria-label="Direct link to The Meta-Irony" title="Direct link to The Meta-Irony" translate="no">​</a></h2>
<p>Yes, I'm aware that writing a blog post about the importance of evidence inventories is itself a documentation-first compulsion taken to its logical extreme. I documented my documentation process and then published documentation about the documentation of the documentation process. There are probably turtles under here somewhere.</p>
<p>But those three errors would have gone to print. Other researchers would have cited them. The 844K figure would have continued circulating through AI discourse with my paper as another link in the chain. The Cloudflare timeline would have been wrong in a way that distorts how quickly the infrastructure landscape is tightening. And the unnamed hosting provider would have been presented as verified data in a document that claims to be rigorous.</p>
<p>Docs first isn't just a workflow. It's epistemic hygiene. Check your facts before you publish them. Check them <em>especially</em> when they confirm what you already believe. And when they come back wrong, publish the corrections with the same enthusiasm you'd publish the findings.</p>
<p>The corrections <em>are</em> the findings. That's the whole point.</p>]]></content:encoded>
            <category>Research</category>
            <category>Methodology</category>
            <category>Opinion</category>
            <category>llms.txt</category>
        </item>
        <item>
            <title><![CDATA[The 844,000 Sites That Weren't: How an AI Adoption Stat Fell Apart Under Scrutiny]]></title>
            <link>https://southpawriter.com/blog/844k-sites-that-werent</link>
            <guid>https://southpawriter.com/blog/844k-sites-that-werent</guid>
            <pubDate>Tue, 17 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[The llms.txt standard was supposedly adopted by 844,000+ websites. I tried to verify that number. What I found instead was a cautionary tale about AI hype, citation chains, and the difference between a directory listing and actual adoption.]]></description>
            <content:encoded><![CDATA[
<p>I need to tell you about a number. It's a number that shows up in blog posts and LinkedIn threads and conference talks and those AI trend reports that get passed around Slack channels like contraband. The number is <strong>844,000</strong>, and it refers to the number of websites that have supposedly adopted the <a class="" href="https://southpawriter.com/docs/glossary/library/llms-txt">llms.txt</a> standard.</p>
<p>I encountered this number while building the evidence inventory for an analytical paper about llms.txt (the Markdown-based content discovery format proposed by Jeremy Howard in September 2024). Because I am the kind of person who builds evidence inventories before writing papers, the kind of person who catalogs every factual claim and traces it back to a primary source before committing a single sentence to a draft, I decided to verify it.</p>
<p>I should not have done this on a weeknight. The verification process involved what I can only describe as the five stages of grief, but for statistics.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="stage-1-acceptance-premature">Stage 1: Acceptance (Premature)<a href="https://southpawriter.com/blog/844k-sites-that-werent#stage-1-acceptance-premature" class="hash-link" aria-label="Direct link to Stage 1: Acceptance (Premature)" title="Direct link to Stage 1: Acceptance (Premature)" translate="no">​</a></h2>
<p>The claim seemed straightforward enough. The llms.txt ecosystem has been growing since the spec launched in late 2024. Community directories exist. Big-name companies like Anthropic, Cloudflare, Stripe, and Vercel have published llms.txt files. There's momentum. There's energy. There's a community of people who genuinely believe in the standard's potential — and I count myself among them, given that I'm <a class="" href="https://southpawriter.com/projects/llmstxtkit">building tools for it</a>, <a href="https://southpawriter.com/assets/files/llms-e1b3b9c6b41a465f74f25eeb5cd3b014.txt" target="_blank" class="">publishing an llms.txt file on this site</a>, and writing an analytical paper about its ecosystem. Eight hundred forty-four thousand sounded high, sure, but maybe not <em>impossibly</em> high. The internet is a big place. Maybe I just underestimated how quickly a good idea spreads.</p>
<p>So I went looking for the primary source. You know, the actual methodology, the crawl data, the published study that counted these files and arrived at precisely "844,000+."</p>
<p>I'm still looking.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="stage-2-bargaining">Stage 2: Bargaining<a href="https://southpawriter.com/blog/844k-sites-that-werent#stage-2-bargaining" class="hash-link" aria-label="Direct link to Stage 2: Bargaining" title="Direct link to Stage 2: Bargaining" translate="no">​</a></h2>
<p>Let me be precise about what I found, because precision matters when you're about to tell an entire community that one of their most-cited statistics might not mean what they think it means.</p>
<p>The community directories linked from the official llms.txt specification? One lists approximately <strong>784 entries</strong>. The other lists about <strong>684</strong>. Not 784,000. Seven hundred and eighty-four. There's overlap between the two, so even the most generous union wouldn't break into four digits.</p>
<p>"Okay," I told myself, because I am nothing if not willing to give a number the benefit of the doubt before I start writing its obituary, "maybe the 844K figure comes from a broader crawl. Maybe someone actually scanned the internet."</p>
<p>Someone did. <a href="https://www.rankability.com/data/llms-txt-adoption/" target="_blank" rel="noopener noreferrer" class="">Rankability's 2025 adoption report</a> crawled approximately 300,000 domains and found about 10% had llms.txt files. That's 30,000 — genuinely impressive, and also roughly 814,000 fewer than the claim. But as I covered in <a class="" href="https://southpawriter.com/blog/waf-paradox-pt2">Part 2</a>, that 10% depends entirely on <em>which</em> domains you crawl. A sample skewed toward developer documentation and tech companies will find higher adoption rates. Samples are not populations, and extrapolating from a self-selecting sample to "the internet" is the statistical equivalent of surveying a Star Trek convention and concluding that 94% of humans own Vulcan ears.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="stage-3-depression">Stage 3: Depression<a href="https://southpawriter.com/blog/844k-sites-that-werent#stage-3-depression" class="hash-link" aria-label="Direct link to Stage 3: Depression" title="Direct link to Stage 3: Depression" translate="no">​</a></h2>
<p>Then I pulled up the data that made me close my laptop and go make coffee.</p>
<p><a href="https://www.chris-green.net/post/million-websites-in-search-of-llms-txt" target="_blank" rel="noopener noreferrer" class="">Chris Green's Majestic Million analysis</a> — the same dataset I examined in <a class="" href="https://southpawriter.com/blog/waf-paradox-pt2">Part 2</a> — found exactly <strong>105 sites</strong> with llms.txt files out of the top million websites by traffic. Among the top 1,000? Zero. Not "nearly zero." Not "a few early adopters." <em>Zero.</em></p>
<p>Let me put those numbers side by side, because I think the contrast is worth sitting with. The claim floating through AI discourse is 844,000+ sites. The verifiable reality: 784 directory listings, 105 files in the top million, and zero in the top thousand. That's a gap between narrative and evidence that you could park several data centers inside.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="stage-4-anger-constructive">Stage 4: Anger (Constructive)<a href="https://southpawriter.com/blog/844k-sites-that-werent#stage-4-anger-constructive" class="hash-link" aria-label="Direct link to Stage 4: Anger (Constructive)" title="Direct link to Stage 4: Anger (Constructive)" translate="no">​</a></h2>
<p>I should note the irony: I'm spending my evenings building a .NET library, a validation framework, and an analytical paper for this standard. If I didn't believe in llms.txt's potential, I would not be deep enough in its ecosystem to notice that a number was wrong. The anger here isn't directed at the community. It's directed at the gap between what the community deserves — honest, defensible data — and what it's been working with.</p>
<p>Now, I want to be clear: I'm not accusing anyone of deliberately fabricating a number. I wasn't able to trace the 844K figure to a single origin point with a clear methodology, which means it may have emerged from some combination of misinterpretation, aggregation errors, optimistic extrapolation, and the telephone game that happens when secondary sources cite tertiary sources citing a blog post that cited a tweet that cited a conference slide that cited vibes.</p>
<p>This happens in every fast-moving technology space. A number gets attached to a narrative, the narrative gets repeated, and eventually the number achieves a kind of undead immortality where it shambles through LinkedIn posts and analyst reports long after anyone remembers where it came from or whether it was ever verified. I've seen it happen with cloud adoption rates, container usage figures, and the perennial claim that "X% of digital transformations fail" (the origin of that statistic is a <em>Dilbert</em> comic, and I am not joking).</p>
<p>What makes the 844K figure specifically worth examining is that it reinforces a particular narrative: that llms.txt has already achieved significant mainstream adoption, that the standard is gaining real traction across the web, and that the ecosystem is mature enough to justify building tools and workflows around it.</p>
<p>The real data tells a different story, and it's actually a <em>more interesting</em> story.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="stage-5-acceptance-actual">Stage 5: Acceptance (Actual)<a href="https://southpawriter.com/blog/844k-sites-that-werent#stage-5-acceptance-actual" class="hash-link" aria-label="Direct link to Stage 5: Acceptance (Actual)" title="Direct link to Stage 5: Acceptance (Actual)" translate="no">​</a></h2>
<p>Here's what the verified adoption data actually shows, and why I think it's more compelling than the inflated number ever was.</p>
<p>As I documented in <a class="" href="https://southpawriter.com/blog/waf-paradox-pt2">Part 2</a>, llms.txt adoption is real but extremely concentrated. The biggest names — Anthropic, Cloudflare, Stripe, Vercel, Coinbase — have impressive implementations. Anthropic's llms.txt alone is 8,364 <a class="" href="https://southpawriter.com/docs/glossary/library/token">tokens</a>. Their llms-full.txt is 481,349. These are serious, curated documents maintained by organizations that understand the standard's value proposition.</p>
<p>But outside the developer-tools bubble? The standard hasn't crossed over. Healthcare, education, government, news, e-commerce — the sectors that represent the vast majority of the web haven't adopted llms.txt, and even if they did today, the <a class="" href="https://southpawriter.com/blog/waf-paradox">infrastructure paradox</a> means their <a class="" href="https://southpawriter.com/docs/glossary/library/web-application-firewall">WAF</a> would probably block the AI crawlers anyway.</p>
<p>This is the story the honest numbers tell: <strong>llms.txt is a young standard in its earliest adoption phase, concentrated among the developer-documentation community that understands its value proposition best.</strong> That's not a failure. That's how standards work. RSS, Markdown itself, <a class="" href="https://southpawriter.com/docs/glossary/library/robots-txt">robots.txt</a> — they all started as niche tools adopted by the people closest to the problem before reaching broader audiences. But it's a very different narrative from "844,000 sites and growing," and the decisions people make about whether to invest in llms.txt should be based on verified reality, not on a number that dissolves under scrutiny like sugar in hot coffee.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-im-publishing-this">Why I'm Publishing This<a href="https://southpawriter.com/blog/844k-sites-that-werent#why-im-publishing-this" class="hash-link" aria-label="Direct link to Why I'm Publishing This" title="Direct link to Why I'm Publishing This" translate="no">​</a></h2>
<p>I write about llms.txt because I think the standard has genuine potential. I'm <a class="" href="https://southpawriter.com/projects/llmstxtkit">building tools for it</a>. I've spent weeks researching its ecosystem. I published an <a href="https://southpawriter.com/assets/files/llms-e1b3b9c6b41a465f74f25eeb5cd3b014.txt" target="_blank" class="">llms.txt file on this very site</a>. I am not a detractor. I am, if anything, an overly enthusiastic supporter who happens to also be the kind of person who fact-checks his own enthusiasm before publishing it.</p>
<p>And that's exactly why the number matters. If the llms.txt community wants to be taken seriously by the broader web ecosystem (by CMS platforms, by hosting providers, by the CDN companies whose WAFs are blocking their standard), they need to lead with honest data. The corrected numbers, roughly 784 directory entries, 105 in the Majestic Million, zero in the top 1,000, tell a story of a young standard in its earliest adoption phase. That story is honest, defensible, and a perfectly fine foundation for an argument that the standard deserves investment and attention.</p>
<p>The 844K number, by contrast, tells a story that crumbles under the first hard question anyone asks. And in a space where trust and credibility are already in short supply (Google's John Mueller compared llms.txt to the discredited <a href="https://www.searchenginejournal.com/google-says-llms-txt-comparable-to-keywords-meta-tag/544804/" target="_blank" rel="noopener noreferrer" class="">keywords meta tag</a>, remember), the last thing the community deserves is to have its credibility anchored to a number that can't survive peer review.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-i-learned-about-verification">What I Learned About Verification<a href="https://southpawriter.com/blog/844k-sites-that-werent#what-i-learned-about-verification" class="hash-link" aria-label="Direct link to What I Learned About Verification" title="Direct link to What I Learned About Verification" translate="no">​</a></h2>
<p>The 844K number wasn't the only claim that fell apart when I checked it. Before I wrote a single paragraph of the analytical paper behind this series, I built an evidence inventory — every factual assertion mapped to a primary source, every source independently verified. The 844K figure was the most visible casualty, but it wasn't the only one.</p>
<p>The full story of what that inventory caught, and why I think every technical writer should build one before they publish anything with factual claims, deserves its own post. That's next.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-real-adoption-picture">The Real Adoption Picture<a href="https://southpawriter.com/blog/844k-sites-that-werent#the-real-adoption-picture" class="hash-link" aria-label="Direct link to The Real Adoption Picture" title="Direct link to The Real Adoption Picture" translate="no">​</a></h2>
<p>So where does llms.txt adoption actually stand? I covered the full dataset in <a class="" href="https://southpawriter.com/blog/waf-paradox-pt2">Part 2</a>. The short version: the scale is measured in hundreds of directory listings and a fraction of a percent of top websites, not hundreds of thousands. The trajectory is promising — Anthropic, Cloudflare, Stripe, and Vercel don't adopt formats on a whim — but the distribution is concentrated among early believers rather than spanning the broader web. That's where every good standard starts. It's just not where the 844K number claims it already is.</p>
<p>I'll keep watching the numbers. I'll keep verifying them. And if the day comes when 844,000 sites genuinely have llms.txt files, I'll be the first to write the correction to this correction.</p>
<p>I just want to see the primary source first.</p>]]></content:encoded>
            <category>llms.txt</category>
            <category>Research</category>
            <category>GEO</category>
        </item>
        <item>
            <title><![CDATA[The llms.txt Access Paradox: The Data Nobody Wants to Hear]]></title>
            <link>https://southpawriter.com/blog/waf-paradox-pt2</link>
            <guid>https://southpawriter.com/blog/waf-paradox-pt2</guid>
            <pubDate>Mon, 16 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[llms.txt adoption is thin, no AI provider confirms inference-time usage, and Cloudflare's panels make WAF configuration a nightmare. It's systemic.]]></description>
            <content:encoded><![CDATA[
<p>In <a class="" href="https://southpawriter.com/blog/waf-paradox">Part 1</a>, I told the story of discovering that my own hosting infrastructure was blocking AI crawlers from reading the llms.txt file I'd specifically published for them. A Web Application Firewall (WAF), the security layer that inspects every inbound HTTP request, can't tell the difference between "AI system reading curated content as intended" and "malicious bot probing endpoints for vulnerabilities," and the result is a paradox that would be hilarious if it weren't also my actual production environment.</p>
<p>That was the personal version, the "I discovered this at 11 PM and said words I can't publish on a professional blog" version. This is the systemic version. The one where I pull at the thread and the whole sweater starts to unravel.</p>
<p>Because once I started asking "how widespread is this?", the answers didn't just confirm the WAF problem. They complicated the entire premise of what llms.txt is supposed to do. And I mean <em>the entire premise.</em></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="llmstxt-adoption-what-the-numbers-actually-show">llms.txt Adoption: What the Numbers Actually Show<a href="https://southpawriter.com/blog/waf-paradox-pt2#llmstxt-adoption-what-the-numbers-actually-show" class="hash-link" aria-label="Direct link to llms.txt Adoption: What the Numbers Actually Show" title="Direct link to llms.txt Adoption: What the Numbers Actually Show" translate="no">​</a></h2>
<p>The adoption picture for llms.txt (the Markdown-based content discovery format proposed by Jeremy Howard in September 2024) is more complicated than the community directories suggest.</p>
<p>Community directories list hundreds of implementations. <a href="https://llmstxt.site/" target="_blank" rel="noopener noreferrer" class="">llmstxt.site</a> has roughly 784 entries, <a href="https://directory.llmstxt.cloud/" target="_blank" rel="noopener noreferrer" class="">directory.llmstxt.cloud</a> has around 684. A <a href="https://www.rankability.com/data/llms-txt-adoption/" target="_blank" rel="noopener noreferrer" class="">Rankability crawl</a> of approximately 300,000 domains found that about 10% had llms.txt files. Ten percent! That sounds like a real standard gaining real traction, until you look at <em>where</em> those files are.</p>
<p>Cross-reference against the Majestic Million (the top one million websites ranked by traffic) and the picture sharpens dramatically. <a href="https://www.chris-green.net/post/million-websites-in-search-of-llms-txt" target="_blank" rel="noopener noreferrer" class="">Chris Green's analysis</a> found only 105 sites with llms.txt files as of May 2025. That's <strong>0.011% of the top million</strong>. Among the top 1,000? Zero. Not "a few." Not "we're getting there." <em>Zero.</em></p>
<p>The adoption is real, but it's overwhelmingly concentrated in developer documentation and tech companies. Stripe has one. Anthropic has one. If you're building AI tools for developers, you'll trip over llms.txt files. If you're building AI tools for literally anything else (healthcare, education, government, news, e-commerce) the standard hasn't crossed over yet. That's not unusual for an eighteen-month-old spec — it's the same trajectory that every developer-born standard follows. The early adopters are the people closest to the problem. Broader adoption follows if (and only if) the tooling and evidence catch up.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-inference-gap-the-data-that-surprised-me">The Inference Gap: The Data That Surprised Me<a href="https://southpawriter.com/blog/waf-paradox-pt2#the-inference-gap-the-data-that-surprised-me" class="hash-link" aria-label="Direct link to The Inference Gap: The Data That Surprised Me" title="Direct link to The Inference Gap: The Data That Surprised Me" translate="no">​</a></h2>
<p>The inference gap is the disconnect between how llms.txt was <em>designed</em> to be used and how AI systems <em>actually</em> interact with it, and it genuinely surprised me, despite the fact that I spend most of my free time reading specification documents for fun.</p>
<p>The llms.txt spec explicitly states that "our expectation is that llms.txt will mainly be useful for <a class="" href="https://southpawriter.com/docs/glossary/library/inference">inference</a>." Inference time. Real-time. When a user asks Claude or ChatGPT a question and the AI goes to fetch information to answer it. That's the entire value proposition: you curate your content so the AI can give better answers <em>right now</em>, not six months from now when the training data catches up.</p>
<p>But <strong>no major AI provider has publicly confirmed using llms.txt at inference time</strong>. Not one.</p>
<p>Google has explicitly rejected the standard. John Mueller of Google compared it to the discredited keywords meta tag in April 2025 (<a href="https://www.searchenginejournal.com/google-says-llms-txt-comparable-to-keywords-meta-tag/544804/" target="_blank" rel="noopener noreferrer" class="">Search Engine Journal</a>), and Gary Illyes stated there's no support at Search Central Live in July 2025 (<a href="https://searchengineland.com/google-says-normal-seo-works-for-ranking-in-ai-overviews-and-llms-txt-wont-be-used-459422" target="_blank" rel="noopener noreferrer" class="">Search Engine Land</a>). <a href="https://yoast.com/llms-txt/" target="_blank" rel="noopener noreferrer" class="">Yoast's analysis</a> found that GPTBot, ClaudeBot, and Google's AI crawlers don't routinely request llms.txt files. Their report also cited an unnamed hosting provider managing thousands of sites that observed zero GPTBot activity on llms.txt endpoints, though the provider wasn't identified, so the attribution chain is thin. Still: even taking that datapoint with appropriate caution, the pattern is consistent. The crawlers aren't showing up for the content that was written for them. That's not evidence of the inference-time usage the spec was designed for. It's a gap between the standard's vision and the ecosystem's current behavior — and it's worth understanding honestly, because the people investing effort in their llms.txt files deserve to know what the data actually shows.</p>
<p>There <em>is</em> evidence of AI crawlers visiting llms.txt files. <a href="https://mintlify.com/blog/llms-txt" target="_blank" rel="noopener noreferrer" class="">Mintlify's analysis</a> shows Microsoft and OpenAI crawlers accessing them, and developer Ray Martinez of Archer Education <a href="https://www.archeredu.com/hemj/are-llms-txt-files-being-implemented-across-the-web/" target="_blank" rel="noopener noreferrer" class="">documented GPTBot pinging their llms.txt every 15 minutes</a> with server log evidence. So <em>someone</em> is knocking. But crawling and inference-time usage are different things. Training-time data collection looks the same in server logs as real-time retrieval, and the evidence is far more consistent with "hoovering up data for the next model" than "using this to answer a user's question right now."</p>
<p>This means the site operators who are carefully crafting llms.txt files and fighting their WAF configurations to make those files accessible may be doing it for an audience that isn't actually showing up at the door. They're leaving the porch light on for guests who may have already eaten elsewhere.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="cloudflare-ai-bot-settings-a-configuration-labyrinth">Cloudflare AI Bot Settings: A Configuration Labyrinth<a href="https://southpawriter.com/blog/waf-paradox-pt2#cloudflare-ai-bot-settings-a-configuration-labyrinth" class="hash-link" aria-label="Direct link to Cloudflare AI Bot Settings: A Configuration Labyrinth" title="Direct link to Cloudflare AI Bot Settings: A Configuration Labyrinth" translate="no">​</a></h2>
<p>Cloudflare (which according to <a href="https://w3techs.com/technologies/details/cn-cloudflare" target="_blank" rel="noopener noreferrer" class="">W3Techs</a> sits in front of roughly 20% of all public websites) provides the tools to configure AI crawler access. They're just scattered across multiple dashboard panels with overlapping authority and an execution order that I can only describe as "designed by someone who wanted job security for the support team."</p>
<p>For the record, I did try to fix the WAF problem on my own site. I am nothing if not persistent, and I had already written a 386-line guide about WAF-AI interactions, so I felt confident. Reader, I was not prepared for what awaited me.</p>
<p>There's the <strong>AI Audit dashboard</strong> (Security &gt; Bots &gt; AI Scrapers and Crawlers), where you can toggle individual AI crawlers between "Blocked" and "Allowed." Cloudflare introduced the one-click opt-in AI blocking controls in July 2024 under the "AIndependence Day" banner. A year later, in July 2025, they escalated: AI bot blocking became the <em>default</em> on newly created domains. No opt-in required.</p>
<p>There's <strong>AI Crawl Control</strong>, which offers more granular categories: <code>ai-train</code> (training data collection), <code>search</code> (search engine use), and <code>ai-input</code> (inference-time access). You can set different policies for each category per crawler.</p>
<p>And then there's the <strong>WAF Custom Rules</strong> layer (Security &gt; WAF &gt; Custom Rules), where you can create path-based exceptions for specific user agents.</p>
<p>Here's the punchline: according to <a href="https://developers.cloudflare.com/ai-crawl-control/configuration/ai-crawl-control-with-waf/" target="_blank" rel="noopener noreferrer" class="">Cloudflare's own documentation</a>, WAF custom rules execute <em>before</em> AI Crawl Control settings. So a security rule you wrote months ago (back when you were a younger, more innocent version of yourself who thought "block non-browser traffic" was a reasonable blanket policy) will intercept an AI crawler before Cloudflare's own AI-specific controls get a chance to say "actually, this one's allowed."</p>
<p>I know this because I sat in my Cloudflare dashboard at 11 PM on a Tuesday, testing and retesting with curl commands, watching the blocks persist, and wondering which of the overlapping security controls was winning the argument:</p>
<div class="container_eudI"><div class="titleBar_n9j0"><div class="trafficLights_qTd9"><span class="dot_siZ2 dotRed_ktsf"></span><span class="dot_siZ2 dotYellow_YbTJ"></span><span class="dot_siZ2 dotGreen_WTxv"></span></div><span class="titleText_RUAw">11 PM on a Tuesday, debugging WAF vs. AI Crawl Control</span><div class="controls_BJva"></div></div><div class="body_qMjP" style="--terminal-line-count:24"></div></div>
<p>It was not my finest hour. But it was educational in the way that touching a hot stove is educational: you only need to learn it once, and you will remember it forever.</p>
<p>If you're in this situation yourself, I've documented the complete diagnostic and mitigation workflow in <a class="" href="https://southpawriter.com/docs/guides/waf-ai-crawler-interaction">Why AI Crawlers Get Blocked (and What You Can Do About It)</a>. That guide covers everything from curl-based diagnosis to WAF allowlist rules to edge-served workarounds. It's the practical document I wish I'd had when I started down this path.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-systemic-problem-why-fixing-one-site-doesnt-fix-llmstxt">The Systemic Problem: Why Fixing One Site Doesn't Fix llms.txt<a href="https://southpawriter.com/blog/waf-paradox-pt2#the-systemic-problem-why-fixing-one-site-doesnt-fix-llmstxt" class="hash-link" aria-label="Direct link to The Systemic Problem: Why Fixing One Site Doesn't Fix llms.txt" title="Direct link to The Systemic Problem: Why Fixing One Site Doesn't Fix llms.txt" translate="no">​</a></h2>
<p>The llms.txt Access Paradox isn't a configuration error on one site. It's a structural conflict between how the standard was designed and how the modern web's security infrastructure operates.</p>
<p>The llms.txt standard assumes a frictionless path between "file published" and "file consumed by AI." That assumption doesn't survive contact with the modern web's security infrastructure. Cloudflare alone sits in front of roughly 20% of all public websites. Akamai, AWS WAF, Fastly, Sucuri, Vercel: they all operate bot-detection systems that treat AI crawlers as threats by default. The burden of configuring exceptions falls entirely on the site operator, who has to understand WAF rule ordering, know that <a class="" href="https://southpawriter.com/docs/glossary/library/robots-txt">robots.txt</a> and WAF policies operate independently, and realize that a <code>200 OK</code> response might be lying to them.</p>
<p>That's not a scalable solution. That's hoping that every site operator who publishes an llms.txt file also happens to be the kind of person who reads Cloudflare documentation recreationally and finds joy in the phrase "WAF rule execution order." Some of us are. (Hi.) Most people have better things to do with their Tuesday nights.</p>
<p>The infrastructure needs to catch up to the intent. If the web wants AI systems to read curated content (and the existence of llms.txt, <a class="" href="https://southpawriter.com/docs/glossary/library/content-signals">Content Signals</a>, <a class="" href="https://southpawriter.com/docs/glossary/library/cc-signals">CC Signals</a>, and a half-dozen other emerging standards suggests it does) then there needs to be a mechanism that's easier than "create a WAF custom rule that skips bot detection for specific paths and specific user agents and make sure it's ordered above all your other security rules." Something more like: <em>if a site publishes an llms.txt file, that file should be accessible to the systems it was written for, without requiring the site operator to become a WAF expert.</em></p>
<p>We're not there yet.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-im-doing-about-it">What I'm Doing About It<a href="https://southpawriter.com/blog/waf-paradox-pt2#what-im-doing-about-it" class="hash-link" aria-label="Direct link to What I'm Doing About It" title="Direct link to What I'm Doing About It" translate="no">​</a></h2>
<p>This experience is why the WAF paradox has become the central finding of a longer analytical paper I'm writing about the state of the llms.txt ecosystem. The paper covers the inference gap (nobody's using it at inference time), the infrastructure paradox (your own security blocks it even when you want AI access), the trust architecture problem (no validation, no freshness guarantees, no consistency checks), and the standards fragmentation issue (too many overlapping standards, none of them complete).</p>
<p>It's also why <a class="" href="https://southpawriter.com/projects/llmstxtkit">LlmsTxtKit</a> treats WAF blocking as a first-class engineering concern rather than an edge case. When the library's fetcher gets a 403 or a JavaScript challenge page, it doesn't just throw an <code>HttpRequestException</code> and leave you to figure out what happened. It classifies the failure, logs diagnostic context, checks for cached content, and returns a structured error that distinguishes "this resource doesn't exist" from "this resource exists but the infrastructure won't let me have it." Because that distinction matters, and because a meaningful percentage of the llms.txt files on the internet are in exactly this state. I wrote the error taxonomy at 2 AM after the Cloudflare incident. It has <em>specificity.</em></p>
<p>And it's why I spent three hours on a Tuesday night arguing with a Cloudflare dashboard instead of writing C# code, which is what I actually enjoy doing, because sometimes the documentation-first developer has to do fieldwork and the field is a web dashboard with more tabs than a browser window at finals week.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="key-takeaways">Key Takeaways<a href="https://southpawriter.com/blog/waf-paradox-pt2#key-takeaways" class="hash-link" aria-label="Direct link to Key Takeaways" title="Direct link to Key Takeaways" translate="no">​</a></h2>
<ul>
<li class=""><strong>llms.txt adoption is real but narrow.</strong> Community directories list hundreds of sites, but Chris Green's Majestic Million analysis found only 105 out of the top million by traffic (0.011%), and zero in the top 1,000. Adoption is concentrated almost entirely in developer documentation.</li>
<li class=""><strong>No major AI provider has confirmed inference-time usage.</strong> Google explicitly rejected the standard (Mueller, April 2025; Illyes, July 2025). Server log evidence is contradictory and more consistent with training-time crawling than real-time retrieval. The spec was designed for inference; the data doesn't show that happening.</li>
<li class=""><strong>Cloudflare's AI controls conflict with each other.</strong> WAF custom rules execute <em>before</em> AI Crawl Control settings, meaning a security rule can override Cloudflare's own AI-specific toggles. Configuring this correctly requires understanding rule execution order that most site operators never encounter.</li>
<li class=""><strong>This is a structural problem, not a configuration problem.</strong> The llms.txt standard assumes frictionless access. The modern web's security infrastructure provides anything but. The burden of making llms.txt work falls entirely on site operators who have to become WAF experts to serve a Markdown file.</li>
<li class=""><strong>The most impactful action isn't adding llms.txt; it's reviewing your WAF configuration.</strong> The infrastructure paradox affects whether <em>any</em> AI system can access <em>any</em> content on your site, llms.txt or otherwise.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="where-this-goes-from-here">Where This Goes From Here<a href="https://southpawriter.com/blog/waf-paradox-pt2#where-this-goes-from-here" class="hash-link" aria-label="Direct link to Where This Goes From Here" title="Direct link to Where This Goes From Here" translate="no">​</a></h2>
<p>If you're a site operator who has published an llms.txt file, I'd encourage you to actually test whether AI crawlers can reach it. The <a class="" href="https://southpawriter.com/docs/guides/waf-ai-crawler-interaction">diagnostic guide</a> walks you through it. It takes about five minutes with curl. You might be surprised by what you find. I certainly was.</p>
<p>While researching the adoption numbers for this post, I kept running into one figure that didn't add up: a claim that 844,000+ websites had adopted llms.txt. It shows up in blog posts, LinkedIn threads, and conference slides. I tried to trace it back to a primary source. What I found instead was a cautionary tale about how numbers achieve a kind of undead immortality in AI discourse — and why the real adoption data, though smaller, tells a more honest and more interesting story. That's <a class="" href="https://southpawriter.com/blog/844k-sites-that-werent">Part 3</a>.</p>
<p>And that verification process? It taught me something about the difference between having sources and having <em>verified</em> sources that changed how I approach every factual claim I write. But that's a methodology story, not a firewall story, and it deserves its own space.</p>
<p>The firewall story is where the Access Paradox starts, because this is where theory meets infrastructure and infrastructure wins. Every time.</p>]]></content:encoded>
            <category>llms.txt</category>
            <category>GEO</category>
            <category>LlmsTxtKit</category>
        </item>
        <item>
            <title><![CDATA[I Tried to Help AI Read My Website. My Own Firewall Said No.]]></title>
            <link>https://southpawriter.com/blog/waf-paradox</link>
            <guid>https://southpawriter.com/blog/waf-paradox</guid>
            <pubDate>Sun, 15 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[I published an llms.txt file for AI crawlers. My Cloudflare WAF blocked every one. The llms.txt Access Paradox, your own infrastructure vs. your content strategy.]]></description>
            <content:encoded><![CDATA[
<p>I did everything right. I wrote the file. I followed the spec. I deployed it to production. I even tested it in my browser: clean Markdown rendering, proper H2 sections, curated links with useful descriptions. My llms.txt file was, and I say this without hyperbole, the best piece of structured content I had ever placed at a root URL. I was <em>proud</em> of that file, in the way that only a documentation-first developer can be proud of a Markdown file that nobody has read.</p>
<p>Then an AI system tried to read it, and my own infrastructure said <em>no.</em></p>
<p>Not a polite "no, sorry, you don't have permission." Not even a helpful "no, that file doesn't exist." The kind of <em>no</em> where Cloudflare intercepts the request before it touches my server, decides the visitor looks suspicious on the basis of (and I love this) being exactly the kind of visitor the file was created for, and serves a JavaScript challenge page instead. To the AI crawler, my lovingly curated Markdown might as well not exist. In its place: a blob of obfuscated HTML designed to prove the visitor is human. Which, by definition, the AI crawler is not. Nor does it aspire to be. That's the <em>entire point.</em></p>
<p>Welcome to what I've started calling the <strong>llms.txt Access Paradox</strong>: the structural conflict between publishing content for AI systems and running the security infrastructure that blocks them. It's the kind of problem that makes you close your laptop, open it again, and start writing a research paper instead of just a blog post.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-setup-building-an-llmstxt-tool-and-finding-a-wall">The Setup: Building an llms.txt Tool and Finding a Wall<a href="https://southpawriter.com/blog/waf-paradox#the-setup-building-an-llmstxt-tool-and-finding-a-wall" class="hash-link" aria-label="Direct link to The Setup: Building an llms.txt Tool and Finding a Wall" title="Direct link to The Setup: Building an llms.txt Tool and Finding a Wall" translate="no">​</a></h2>
<p>Some context. I'd been building <a class="" href="https://southpawriter.com/projects/llmstxtkit">LlmsTxtKit</a>, a C#/.NET library for parsing, fetching, validating, and caching llms.txt files, and I needed to test it against real-world sites. Not test files I'd written myself. Real sites, with real infrastructure, behind real CDNs and firewalls.</p>
<p>The <a href="https://llmstxt.org/" target="_blank" rel="noopener noreferrer" class="">llms.txt standard</a> is a content discovery format proposed by Jeremy Howard of Answer.AI in September 2024. The premise is elegant: put a Markdown file at <code>/llms.txt</code> on your site that gives AI systems a curated, structured summary of your content. Instead of forcing a language model to parse your entire HTML soup at <a class="" href="https://southpawriter.com/docs/glossary/library/inference">inference</a> time, you hand it the good stuff on a silver platter. The spec is deliberately minimal: an H1 title, an optional blockquote summary, H2 sections with link lists, and an Optional section for lower-priority content. The canonical Python parser is roughly 20 lines of regex. It's the kind of elegant minimalism that makes a documentation nerd's heart sing. (Mine sang. I'll admit it.)</p>
<p>So I started fetching llms.txt files from sites I knew had them. Anthropic. Cloudflare. Stripe. Vercel. The big names that show up in the community directories.</p>
<p>And the responses started coming back wrong.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-a-waf-block-looks-like-when-youre-expecting-markdown">What a WAF Block Looks Like When You're Expecting Markdown<a href="https://southpawriter.com/blog/waf-paradox#what-a-waf-block-looks-like-when-youre-expecting-markdown" class="hash-link" aria-label="Direct link to What a WAF Block Looks Like When You're Expecting Markdown" title="Direct link to What a WAF Block Looks Like When You're Expecting Markdown" translate="no">​</a></h2>
<p>When you fetch a well-formed llms.txt file from a site that isn't blocking you, the HTTP response is clean and predictable: a <code>200 OK</code> with <code>text/markdown</code> content. Here's what that <em>should</em> look like:</p>
<div class="container_eudI"><div class="titleBar_n9j0"><div class="trafficLights_qTd9"><span class="dot_siZ2 dotRed_ktsf"></span><span class="dot_siZ2 dotYellow_YbTJ"></span><span class="dot_siZ2 dotGreen_WTxv"></span></div><span class="titleText_RUAw">expected llms.txt response</span><div class="controls_BJva"></div></div><div class="body_qMjP" style="--terminal-line-count:9"></div></div>
<p>And here's what I actually got from a meaningful percentage of sites: a <code>403 Forbidden</code>, or worse, a <code>200</code> that delivers a JavaScript challenge page instead of Markdown:</p>
<div class="container_eudI"><div class="titleBar_n9j0"><div class="trafficLights_qTd9"><span class="dot_siZ2 dotRed_ktsf"></span><span class="dot_siZ2 dotYellow_YbTJ"></span><span class="dot_siZ2 dotGreen_WTxv"></span></div><span class="titleText_RUAw">actual response, WAF block</span><div class="controls_BJva"></div></div><div class="body_qMjP" style="--terminal-line-count:13"></div></div>
<p>A <code>403 Forbidden</code>. And when it wasn't a 403, it was worse: a <code>200 OK</code> that looked fine at the status code level but delivered a JavaScript challenge page instead of my Markdown. The response body was a wall of obfuscated <code>&lt;script&gt;</code> tags, a <code>&lt;noscript&gt;</code> block telling me to enable JavaScript, and a Cloudflare-branded interstitial designed for a web browser that the AI crawler fundamentally is not.</p>
<p>My tool dutifully parsed this HTML as if it were an llms.txt file and returned garbage. Because from the HTTP layer, everything looked fine. The 200 status code said "here's your content." The content said "prove you're a human first."</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-wafs-block-ai-crawlers-the-short-version">How WAFs Block AI Crawlers (The Short Version)<a href="https://southpawriter.com/blog/waf-paradox#how-wafs-block-ai-crawlers-the-short-version" class="hash-link" aria-label="Direct link to How WAFs Block AI Crawlers (The Short Version)" title="Direct link to How WAFs Block AI Crawlers (The Short Version)" translate="no">​</a></h2>
<p>A <a class="" href="https://southpawriter.com/docs/glossary/library/web-application-firewall">Web Application Firewall (WAF)</a> is a security layer that sits between the public internet and your web server, inspecting every inbound request to distinguish legitimate visitors from threats. I've since written a <a class="" href="https://southpawriter.com/docs/guides/waf-ai-crawler-interaction">comprehensive guide</a> covering the full diagnostic and mitigation workflow for WAF-AI conflicts. If you're dealing with this yourself, start there. But the short version is instructive.</p>
<p>WAFs block real attacks: SQL injection, cross-site scripting, credential stuffing, DDoS floods, and the daily background radiation of bots probing every exposed endpoint on the internet. They are very, very good at their jobs. They are also completely indifferent to yours.</p>
<p>The problem is that AI crawlers look <em>exactly</em> like the threats WAFs are designed to block:</p>
<ul>
<li class="">They don't execute JavaScript. (Neither do SQL injection bots.)</li>
<li class="">They don't maintain cookies or session state. (Neither do scraping bots.)</li>
<li class="">They originate from data center IP ranges. (So does virtually all malicious traffic.)</li>
<li class="">They use non-browser <a class="" href="https://southpawriter.com/docs/glossary/library/user-agent">user agents</a> like <code>GPTBot/1.0</code> or <code>ClaudeBot/1.0</code>. (At least they're honest about it. The WAF punishes the honesty.)</li>
<li class="">They make a single stateless GET request and vanish. (To the WAF, this looks like a surgical strike, not a browsing session.)</li>
</ul>
<p>Each signal individually might not trigger a block. Together, they're a neon sign that reads "I AM NOT A HUMAN AND I'M NOT EVEN TRYING TO PRETEND." Which, to be fair, is refreshingly honest in an age of sophisticated impersonation. The WAF does not reward honesty.</p>
<p>And here's the paradox: the site owner who put that llms.txt file at their root <em>wants</em> these visitors. They created the file <em>for</em> these visitors. But their own infrastructure (infrastructure they're running because they have to, because the internet is a hostile environment) can't distinguish between "AI system reading my curated content as intended" and "malicious bot probing my endpoints for vulnerabilities."</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-invisible-failure-why-most-site-operators-never-notice">The Invisible Failure: Why Most Site Operators Never Notice<a href="https://southpawriter.com/blog/waf-paradox#the-invisible-failure-why-most-site-operators-never-notice" class="hash-link" aria-label="Direct link to The Invisible Failure: Why Most Site Operators Never Notice" title="Direct link to The Invisible Failure: Why Most Site Operators Never Notice" translate="no">​</a></h2>
<p>The most consequential aspect of the llms.txt Access Paradox is that the failure mode is invisible to the people who could fix it.</p>
<p>When Cloudflare blocks an AI crawler from reading your llms.txt, it doesn't send you an email. It doesn't generate an alert. It doesn't pop up a notification saying "Hey, we prevented ClaudeBot from reading the file you specifically published for ClaudeBot." The block shows up in your Security Events dashboard, if you know where to look and if you're the kind of person who regularly reviews WAF logs for blocked requests to a Markdown file. Most site operators aren't. Why would they be?</p>
<p>So the file sits there. It's technically online. You can read it in your browser. If someone asks "do you have an llms.txt?" you can say yes, and you'd be right in the most literal and least useful sense. It's the digital equivalent of publishing a book and then hiring a security guard who prevents anyone matching the physical description of "reads books" from entering the bookstore.</p>
<p>I only discovered the blocking because I'm the kind of person who builds a .NET library to fetch Markdown files from strangers' web servers and then gets emotionally invested in the HTTP status codes. Most site operators are not this person. (Congratulations to them.) Their llms.txt files are functionally inaccessible to their entire target audience, and nobody (not the site operator, not the WAF provider, not the AI system) has any reason to notice.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="key-takeaways">Key Takeaways<a href="https://southpawriter.com/blog/waf-paradox#key-takeaways" class="hash-link" aria-label="Direct link to Key Takeaways" title="Direct link to Key Takeaways" translate="no">​</a></h2>
<ul>
<li class=""><strong>The llms.txt Access Paradox is real.</strong> A site can publish an llms.txt file and have it be functionally inaccessible to every AI system it was designed for, because the WAF between the internet and the server treats AI crawlers as threats.</li>
<li class=""><strong>AI crawlers fail every layer of bot detection.</strong> No JavaScript execution, no cookies, data center IPs, non-browser user agents, and single stateless requests. To a WAF, an AI crawler is indistinguishable from a malicious bot.</li>
<li class=""><strong>The failure is invisible to site operators.</strong> WAFs don't alert you when they block ClaudeBot or GPTBot. Your llms.txt file looks fine in a browser. The only way to discover the block is to test with AI-representative user agents, which most site operators never do.</li>
<li class=""><strong>If you have an llms.txt file, test it.</strong> A five-minute <code>curl</code> test with a GPTBot user agent string will tell you whether your WAF is blocking the audience your file was written for. The <a class="" href="https://southpawriter.com/docs/guides/waf-ai-crawler-interaction">diagnostic guide</a> walks you through it.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-comes-next">What Comes Next<a href="https://southpawriter.com/blog/waf-paradox#what-comes-next" class="hash-link" aria-label="Direct link to What Comes Next" title="Direct link to What Comes Next" translate="no">​</a></h2>
<p>At this point I had a choice. I could fix the problem on my own site, update my Cloudflare settings, pat myself on the back, maybe write a quick "TIL" post, and move on with my life like a normal person.</p>
<p>Or I could start asking harder questions. How widespread is this? Is the blocking the only problem, or is it a symptom of something deeper? Does anyone actually <em>use</em> llms.txt at inference time, or is the whole standard running on good intentions and an absence of anyone checking?</p>
<p>If you've read <a class="" href="https://southpawriter.com/blog/docs-first-rabbit-hole">my first blog post</a>, you already know which option I picked. I am not a normal person. I am a documentation-first developer with a research compulsion and a growing collection of Markdown files about Markdown files. I chose the harder questions.</p>
<p>In <a class="" href="https://southpawriter.com/blog/waf-paradox-pt2">Part 2</a>, I dig into the adoption numbers, the inference gap that nobody's talking about, the Cloudflare configuration labyrinth I navigated at 11 PM on a Tuesday, and why fixing this on one site doesn't fix the systemic problem. The firewall story is where the Access Paradox starts, but the data is where it gets genuinely uncomfortable.</p>]]></content:encoded>
            <category>llms.txt</category>
            <category>GEO</category>
            <category>LlmsTxtKit</category>
        </item>
        <item>
            <title><![CDATA[I Write the Docs Before the Code, and Yes, I Know That's Weird]]></title>
            <link>https://southpawriter.com/blog/docs-first-rabbit-hole</link>
            <guid>https://southpawriter.com/blog/docs-first-rabbit-hole</guid>
            <pubDate>Fri, 13 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[A technical writer walks into an AI ecosystem, starts documenting everything before writing a single line of code, and accidentally stumbles into an under-researched rabbit hole involving llms.txt, hostile firewalls, and the existential question of whether AI can even read the internet it was trained on.]]></description>
            <content:encoded><![CDATA[<p>I have a confession to make. When I start a new project, any project, doesn't matter what it is, the first thing I do is open a Markdown file and start writing documentation for something that doesn't exist yet.</p>
<p>Not code. Not a prototype. Not even a to-do list. <em>Documentation.</em></p>
<p>I realize this makes me sound like the kind of person who reads the terms of service before clicking "I Agree." I promise I'm not. (I absolutely am.)</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-docs-first-affliction">The Docs-First Affliction<a href="https://southpawriter.com/blog/docs-first-rabbit-hole#the-docs-first-affliction" class="hash-link" aria-label="Direct link to The Docs-First Affliction" title="Direct link to The Docs-First Affliction" translate="no">​</a></h2>
<p>There's a term for this in software circles: "documentation-driven development." It sounds very official and intentional, like something a thoughtful engineering team adopts after a retrospective. In my case, it's more of a compulsion. I'm a technical writer by trade, and somewhere along the way my brain got wired to believe that if something isn't documented, it doesn't really exist.</p>
<p>This means that before I write a class, I've already written the spec that describes the class. Before I write a test, I've already written the test plan that describes what the test should cover and why. Before I write <em>this blog post</em>, I wrote a content outline, a publication schedule, and a front matter template. It's documentation all the way down.</p>
<p>You might think this slows me down. You'd be right. You might also think it produces better outcomes when I eventually do write code. You'd also be right about that.</p>
<p>The tradeoff is that I start every project with a pile of Markdown files and zero functioning software, which looks, to the outside observer, like I've accomplished nothing. But those Markdown files? They've saved me from building the wrong thing more times than I can count. Turns out, it's really hard to write a clear product requirements spec for a bad idea. The spec fights back. It asks questions you didn't want to answer. It exposes assumptions you were hoping to sneak past yourself.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-menagerie">The Menagerie<a href="https://southpawriter.com/blog/docs-first-rabbit-hole#the-menagerie" class="hash-link" aria-label="Direct link to The Menagerie" title="Direct link to The Menagerie" translate="no">​</a></h2>
<p>This docs-first habit has produced, over time, a slightly eclectic collection of projects. Each one started as a Markdown file that got increasingly detailed until I had no choice but to actually build the thing.</p>
<p><strong>Lexichord</strong> started as a frustrated manifesto. I was watching every AI writing tool on the market treat "style guide compliance" as a system prompt afterthought, just paste your brand voice into a text box and hope for the best. Anyone who's ever maintained a real enterprise style guide (with terminology databases, audience-specific complexity rules, document-type-specific formatting, and the seventeen exceptions to every rule) knows that's not how any of this works. So I wrote a spec for what an AI <a class="" href="https://southpawriter.com/docs/glossary/library/orchestration">orchestration</a> tool <em>should</em> look like for technical writers. Not a chatbot with a style guide bolted on, but a proper orchestration layer that coordinates multiple AI capabilities around a writer's actual workflow. Then I had to go build it, because the spec was too compelling to ignore.</p>
<p><strong>FractalRecall</strong> started as a question scribbled in a notebook: "What if metadata isn't the thing you attach to <a class="" href="https://southpawriter.com/docs/glossary/library/embedding">embeddings</a> <em>after</em> you create them, but the thing that should shape <em>how</em> you create them in the first place?" That question turned into a design spec. The design spec turned into a system that treats metadata as structural DNA, something that fundamentally informs how data gets vectorized, rather than just tagging it after the fact. Most <a class="" href="https://southpawriter.com/docs/glossary/library/retrieval-augmented-generation">RAG</a> pipelines treat metadata as a post-retrieval filter. FractalRecall argues that's backwards. The jury is still out on whether I'm onto something or just being stubborn. (These are not mutually exclusive.)</p>
<p><strong>DocStratum</strong> is exactly what it sounds like if you're the kind of person who names validation tools after geological layering processes. It validates llms.txt files against the spec (think ESLint, but for a Markdown standard that's defined by a blog post rather than a formal grammar). More on llms.txt in a moment, because that particular rabbit hole deserves its own section.</p>
<p><strong>Rune &amp; Rust</strong> is the odd one out and I love it for that. It's a text-based dungeon crawler written in C#, which in 2026 makes me either charmingly retro or deeply out of touch, depending on who you ask. No rendering engine, no sprite sheets, just state machines and narrative branching and the quiet satisfaction of modeling combat mechanics with pattern matching expressions. I built it partly because game logic is an incredible workout for design patterns, and partly because sometimes you just need to write code that stabs a goblin instead of parsing a compliance document.</p>
<p>And then there's <strong>LlmsTxtKit</strong>, which brings us to the rabbit hole.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-i-accidentally-became-an-llmstxt-researcher">How I Accidentally Became an llms.txt Researcher<a href="https://southpawriter.com/blog/docs-first-rabbit-hole#how-i-accidentally-became-an-llmstxt-researcher" class="hash-link" aria-label="Direct link to How I Accidentally Became an llms.txt Researcher" title="Direct link to How I Accidentally Became an llms.txt Researcher" translate="no">​</a></h2>
<p>It started, as these things often do, with a problem that seemed small.</p>
<p>I was poking around the AI tooling landscape (as one does when one has a docs-first compulsion and a GitHub account) and I noticed something odd. The llms.txt standard, proposed by Jeremy Howard in late 2024, was gaining real traction — community directories listing hundreds of implementations, big names like Anthropic and Cloudflare publishing files, a growing ecosystem of tools. The idea is elegant: put a Markdown file at <code>/llms.txt</code> on your site that gives AI systems a curated, structured summary of your content. Instead of forcing a language model to parse your entire HTML soup at <a class="" href="https://southpawriter.com/docs/glossary/library/inference">inference</a> time, you hand it the good stuff on a silver platter.</p>
<p>Great concept. I started looking for tools to work with it. Python module? Yep. JavaScript implementation? Sure. VitePress plugin, Docusaurus plugin, PHP library, Drupal module? All present. C#/.NET implementation? <em>Crickets.</em> Not a single one.</p>
<p>Now, I'm a C# developer. I like C#. It has <a class="" href="https://southpawriter.com/docs/glossary/library/nullable-reference-types">nullable reference types</a> and <a class="" href="https://southpawriter.com/docs/glossary/library/pattern-matching">pattern matching</a> and <a class="" href="https://southpawriter.com/docs/glossary/library/file-scoped-namespaces">file-scoped namespaces</a> and it doesn't make me manage my own memory like some kind of 1990s survivalist. So naturally, I decided to build the .NET llms.txt library myself. Docs first, obviously. I wrote the product requirements spec. I wrote the design spec. I started sketching out the architecture: parsing, fetching, validation, caching, context generation, the whole pipeline.</p>
<p>And that's when I tripped over the rabbit hole.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-rabbit-hole-has-a-firewall">The Rabbit Hole Has a Firewall<a href="https://southpawriter.com/blog/docs-first-rabbit-hole#the-rabbit-hole-has-a-firewall" class="hash-link" aria-label="Direct link to The Rabbit Hole Has a Firewall" title="Direct link to The Rabbit Hole Has a Firewall" translate="no">​</a></h2>
<p>See, when you build a tool that fetches llms.txt files from the open web, you quickly discover something that almost nobody is talking about: <em>the web doesn't want to let you.</em></p>
<p>Cloudflare sits in front of roughly 20% of all public websites. In July 2025, they started blocking AI crawlers by default on new domains. They called it "AIndependence Day," which is the kind of branding that makes you wonder if there's a pun committee at Cloudflare and whether they're hiring.</p>
<p>The blocking happens because AI crawlers (programs that fetch web pages for language models) trip every bot-detection heuristic in the book. They don't execute JavaScript. They don't maintain cookies. They come from data center IP ranges. They use non-browser <a class="" href="https://southpawriter.com/docs/glossary/library/user-agent">user agents</a>. To a <a class="" href="https://southpawriter.com/docs/glossary/library/web-application-firewall">Web Application Firewall</a>, they look indistinguishable from a DDoS attack with a liberal arts degree.</p>
<p>Here's the paradox: a site owner creates an llms.txt file specifically to help AI systems read their content. Their hosting infrastructure then blocks the AI systems that try to read it. The AI systems, unable to fetch the curated content, fall back to search APIs and get whatever Google decides to surface instead. Everyone involved is doing exactly what they're supposed to do. The system still doesn't work.</p>
<p>I found this <em>fascinating.</em> Not "mildly interesting" fascinating. "I need to write a 10,000-word analytical paper about this" fascinating.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="docs-first-research-second-code-eventually">Docs-First, Research Second, Code... Eventually<a href="https://southpawriter.com/blog/docs-first-rabbit-hole#docs-first-research-second-code-eventually" class="hash-link" aria-label="Direct link to Docs-First, Research Second, Code... Eventually" title="Direct link to Docs-First, Research Second, Code... Eventually" translate="no">​</a></h2>
<p>So that's what I'm doing. The llms.txt rabbit hole has expanded into a full research initiative spanning three interconnected projects:</p>
<p><strong>"The llms.txt Access Paradox"</strong> is an analytical paper that documents this whole mess: the gap between how the standard was <em>designed</em> to work, how it <em>actually</em> works in practice, and why the infrastructure itself is the biggest obstacle. It covers the inference gap (no major AI provider has confirmed they use llms.txt at inference time), the WAF paradox (your own security tools block the very systems you're trying to help), the trust problem (Google's John Mueller compared llms.txt to the discredited keywords meta tag, and he's not entirely wrong), and the fragmentation issue (llms.txt, <a class="" href="https://southpawriter.com/docs/glossary/library/content-signals">Content Signals</a>, <a class="" href="https://southpawriter.com/docs/glossary/library/cc-signals">CC Signals</a>, and <a class="" href="https://southpawriter.com/docs/glossary/library/ietf-aipref">IETF aipref</a> are all trying to solve overlapping problems in incompatible ways).</p>
<p><strong>LlmsTxtKit</strong> is the C# library that started this whole journey. It parses, fetches, validates, caches, and generates context from llms.txt files. It also ships as an <a class="" href="https://southpawriter.com/docs/glossary/library/mcp-server">MCP server</a> (that's the <a class="" href="https://southpawriter.com/docs/glossary/library/model-context-protocol">Model Context Protocol</a> from Anthropic) so AI agents can use it as a tool directly. The design is documentation-first (obviously), and the library handles the WAF-blocking problem gracefully instead of just throwing an exception and calling it a day.</p>
<p><strong>The Context Collapse Mitigation Benchmark</strong> is the empirical study asking the question that actually matters: does any of this make a difference? If you feed an LLM curated Markdown via llms.txt instead of raw HTML scraped from the web, does it give better answers? I'm running a controlled experiment with paired conditions across 30–50 websites, scoring factual accuracy, <a class="" href="https://southpawriter.com/docs/glossary/library/hallucination">hallucination</a> rates, and citation fidelity. If the answer is "no, it doesn't matter," that's a valid and publishable finding. Science doesn't care about your hypothesis.</p>
<p>And yes, I wrote the research proposal, the roadmap, the paper outline, and the benchmark methodology spec before writing any experiment code. The docs-first affliction is real.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-this-blog-is-now">What This Blog Is Now<a href="https://southpawriter.com/blog/docs-first-rabbit-hole#what-this-blog-is-now" class="hash-link" aria-label="Direct link to What This Blog Is Now" title="Direct link to What This Blog Is Now" translate="no">​</a></h2>
<p>If you read the earlier posts on this site, you'll notice they were... let's call them "warming up." A welcome post. A Docusaurus setup walkthrough. Perfectly fine content, but not exactly what keeps someone's RSS feed interesting.</p>
<p>Going forward, this blog is where all of this work gets published. The llms.txt research has a dedicated eight-post series planned, starting with the WAF story (because it's a genuinely good story and I have opinions), moving through the research findings, the .NET ecosystem gap, GEO practices, and eventually the benchmark results.</p>
<p>But it's not going to be all llms.txt, all the time. I have posts planned about FractalRecall's metadata-as-DNA approach to embeddings, about what building a dungeon crawler teaches you about state machines (more than you'd expect), about why enterprise AI tools keep failing technical writers, and about the general experience of being a C# developer in a Python-dominated AI ecosystem. It's lonely out here. We have great pattern matching and nobody to share it with.</p>
<p>The common thread across all of it is the docs-first philosophy: understand the problem deeply, write it down clearly, and then build something worth building. Every project I work on starts with documentation that could stand on its own, because if you can't explain what you're building and why, you probably shouldn't be building it.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="before-you-go">Before You Go<a href="https://southpawriter.com/blog/docs-first-rabbit-hole#before-you-go" class="hash-link" aria-label="Direct link to Before You Go" title="Direct link to Before You Go" translate="no">​</a></h2>
<p>If any of this resonates (the documentation-first approach, the C# advocacy, the "wait, firewalls block the thing they're supposed to serve?" reaction) stick around. Subscribe to the <a href="https://southpawriter.com/blog/rss.xml">RSS feed</a> if that's your thing. The llms.txt research is ongoing, the tools are taking shape, and I guarantee the benchmark results are going to make at least one person on LinkedIn very uncomfortable.</p>
<p>And if you're a .NET developer who's been quietly wondering why the entire AI ecosystem seems to pretend C# doesn't exist: you're not alone. Pull up a chair. We have work to do.</p>]]></content:encoded>
            <category>Opinion</category>
            <category>Career</category>
        </item>
    </channel>
</rss>