<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Ai on Dexome</title>
        <link>https://blog.dexome.com/tags/ai/</link>
        <description>Recent content in Ai on Dexome</description>
        <generator>Hugo -- gohugo.io</generator>
        <language>en</language>
        <lastBuildDate>Wed, 10 Jun 2026 00:00:00 +0530</lastBuildDate><atom:link href="https://blog.dexome.com/tags/ai/index.xml" rel="self" type="application/rss+xml" /><item>
        <title>Running Local LLMs on an Intel Arc A310</title>
        <link>https://blog.dexome.com/post/self-hosting-llms-consumer-gpus/</link>
        <pubDate>Wed, 10 Jun 2026 00:00:00 +0530</pubDate>
        
        <guid>https://blog.dexome.com/post/self-hosting-llms-consumer-gpus/</guid>
        <description>&lt;p&gt;I use an Intel Arc A310 with 4 GB VRAM for Frigate, Immich and media transcoding.
I wanted to use the same GPU for a private LibreChat setup, with Ollama as the
model server and Netdata exposed through MCP.&lt;/p&gt;
&lt;p&gt;Small models worked for normal chat. Netdata tool calling did not work reliably.
The first problem was Ollama&amp;rsquo;s default 4096-token context, which removed the MCP
tool definitions from the prompt. After increasing the context, the remaining
problem was the capability of models which could fit on this GPU.&lt;/p&gt;
&lt;h2 id=&#34;my-setup&#34;&gt;My setup
&lt;/h2&gt;&lt;p&gt;Ollama 0.12.11 and later include an experimental Vulkan backend that can use
Intel Arc. I passed only &lt;code&gt;/dev/dri/renderD128&lt;/code&gt;, joined the render and video
groups, and kept the API on private Podman networks. LibreChat reached Ollama by
container DNS; Ollama used a separate egress bridge to pull models.&lt;/p&gt;
&lt;figure class=&#34;article-diagram&#34;&gt;
    &lt;pre class=&#34;mermaid&#34;&gt;
flowchart TB
	USER[Browser] --&gt; CHAT[LibreChat]
	CHAT --&gt;|private bridge| O[Ollama and Vulkan]
	CHAT --&gt;|13 tool schemas| MCP[Netdata MCP]
	O --&gt; ARC[Intel Arc A310, 4 GB]
	VIDEO[Frigate, Immich, Plex, Jellyfin] --&gt; ARC
	O --&gt;|model pulls only| NET[Internet bridge]
&lt;/pre&gt;
    &lt;figcaption&gt;The model API stays private while one render node is shared with the host&amp;#39;s media and vision workloads.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The Vulkan backend worked for small local models. The first issue appeared before
the model generated any response.&lt;/p&gt;
&lt;h2 id=&#34;mcp-tools-were-removed-by-context-truncation&#34;&gt;MCP tools were removed by context truncation
&lt;/h2&gt;&lt;p&gt;LibreChat sent the system message, Netdata instructions, 13 MCP tool schemas,
and the user turn. The resulting first prompt was about 11,511 tokens. Ollama&amp;rsquo;s
default context was 4,096.&lt;/p&gt;
&lt;p&gt;The log made the failure explicit:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;div class=&#34;chroma&#34;&gt;
&lt;table class=&#34;lntable&#34;&gt;&lt;tr&gt;&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;lnt&#34;&gt;1
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;
&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;truncating input prompt limit=4096 prompt=11511 keep=4 new=4095
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;
&lt;/div&gt;
&lt;/div&gt;&lt;p&gt;The tool definitions were near the discarded end. A model that never receives
the schemas cannot issue a tool call, regardless of its instruction following.
It answered in prose because prose was the only action left.&lt;/p&gt;
&lt;figure class=&#34;article-diagram&#34;&gt;
    &lt;pre class=&#34;mermaid&#34;&gt;
flowchart TB
	P[11.5k-token prompt] --&gt; CUT{4,096-token context}
	CUT --&gt; KEEP[Small retained slice]
	CUT -. discarded .-&gt; TOOLS[Netdata instructions and 13 tool schemas]
	KEEP --&gt; MODEL[Model sees no callable tools]
	MODEL --&gt; TEXT[Plain-text answer]
&lt;/pre&gt;
    &lt;figcaption&gt;At the default 4k context, truncation removed the MCP schemas before the model evaluated the request.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Raising &lt;code&gt;OLLAMA_CONTEXT_LENGTH&lt;/code&gt; made structured tool attempts appear. That
proved truncation was causal. It did not yet produce a usable system.&lt;/p&gt;
&lt;h2 id=&#34;larger-context-used-more-vram&#34;&gt;Larger context used more VRAM
&lt;/h2&gt;&lt;p&gt;At 16,384 tokens, the KV cache occupied about 1.8 GB. Only 18 of 29 model layers
fit on the GPU; the other 11 spilled to CPU. Processing the first 11.5k-token
prompt took 131 seconds, during which nothing streamed. LibreChat aborted before
the first token.&lt;/p&gt;
&lt;p&gt;Turning off the 6 KB server-instructions block helped less than expected. The
prompt still measured 10,189 tokens because the 13 tool schemas themselves were
the dominant cost, and LibreChat could not expose only a subset of MCP tools.
An 8,192-token context still truncated them.&lt;/p&gt;
&lt;p&gt;The practical setting was 12,288: enough for the tool prompt plus answer
headroom, with fewer layers displaced than at 16k. The first turn remained slow,
roughly 80 to 100 seconds for a 3B model. Follow-up turns were fast because
Ollama cached the prompt prefix; one 67-token follow-up returned in 0.8 seconds.&lt;/p&gt;
&lt;figure class=&#34;article-diagram&#34;&gt;
    &lt;pre class=&#34;mermaid&#34;&gt;
flowchart TB
	C4[4k context] --&gt;|small KV cache| FAST[More GPU residency]
	C4 --&gt;|but| TRUNC[Tool schemas truncated]
	TRUNC -. next test .-&gt; C12[12k context]
	C12 --&gt;|schemas fit| VISIBLE[Tools visible]
	C12 --&gt;|larger KV cache| SPLIT[Some layers spill to CPU]
	SPLIT -. next test .-&gt; C16[16k context]
	C16 --&gt;|about 1.8 GB KV| SLOW[18 of 29 layers on GPU, 131s prompt evaluation]
&lt;/pre&gt;
    &lt;figcaption&gt;A larger context fixed schema visibility but enlarged the KV cache and forced model layers onto the CPU.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The model files fitted, but the model, KV cache and the other GPU workloads did
not always fit at the same time.&lt;/p&gt;
&lt;h2 id=&#34;models-i-tested&#34;&gt;Models I tested
&lt;/h2&gt;&lt;p&gt;With logs confirming no truncation, I tested the models against simple Netdata
questions.&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Model&lt;/th&gt;
          &lt;th style=&#34;text-align: right&#34;&gt;GPU placement&lt;/th&gt;
          &lt;th style=&#34;text-align: right&#34;&gt;First prompt&lt;/th&gt;
          &lt;th&gt;Result&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;Qwen2.5 1.5B&lt;/td&gt;
          &lt;td style=&#34;text-align: right&#34;&gt;29/29 layers&lt;/td&gt;
          &lt;td style=&#34;text-align: right&#34;&gt;about 55s&lt;/td&gt;
          &lt;td&gt;Printed a tool call as JSON text; LibreChat could not execute it&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Llama 3.2 3B&lt;/td&gt;
          &lt;td style=&#34;text-align: right&#34;&gt;20-21/29 layers&lt;/td&gt;
          &lt;td style=&#34;text-align: right&#34;&gt;105-131s&lt;/td&gt;
          &lt;td&gt;Emitted a real call with schema-invalid arguments&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Qwen2.5 7B&lt;/td&gt;
          &lt;td style=&#34;text-align: right&#34;&gt;13/29 layers&lt;/td&gt;
          &lt;td style=&#34;text-align: right&#34;&gt;about 218s&lt;/td&gt;
          &lt;td&gt;Too slow and returned an unusable result&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The 3B model even failed the no-argument &lt;code&gt;list_raised_alerts&lt;/code&gt; schema. This was
not limited to one complicated metrics query. It could choose a tool and emit a
tool-call shape, but not produce arguments the MCP client accepted reliably.&lt;/p&gt;
&lt;p&gt;LibreChat&amp;rsquo;s &amp;ldquo;Ran tool&amp;rdquo; pill was not proof of success. It appeared when dispatch
began; the execution log later showed &lt;code&gt;Received tool input did not match expected schema&lt;/code&gt;. Without checking that log, I would have mistaken an attempted
call followed by hallucinated prose for real monitoring data.&lt;/p&gt;
&lt;p&gt;Increasing the timeout only waited longer for the same result. It did not fix
the invalid tool arguments or make the 7B model fit better.&lt;/p&gt;
&lt;h2 id=&#34;testing-the-intel-sycl-backend&#34;&gt;Testing the Intel SYCL backend
&lt;/h2&gt;&lt;p&gt;Intel&amp;rsquo;s old IPEX-LLM path looked attractive because it promised optimized SYCL
inference. By the time I evaluated it, the repository was archived, its bundled
Ollama was old, the images used rolling tags, and the project was flagged with
known security issues. I rejected it.&lt;/p&gt;
&lt;p&gt;The maintained high-throughput option is upstream llama.cpp&amp;rsquo;s Intel SYCL image.
I staged it beside Ollama rather than replacing the working service. It needed
both &lt;code&gt;renderD128&lt;/code&gt; and this host&amp;rsquo;s actual card node, &lt;code&gt;card1&lt;/code&gt;; assuming &lt;code&gt;card0&lt;/code&gt;
prevented container creation.&lt;/p&gt;
&lt;p&gt;llama-server introduced its own constraints:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;the default four parallel slots divided the usable context per request;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--parallel 1&lt;/code&gt; was necessary for the large single-user prompt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--jinja&lt;/code&gt; was required for structured tool calls;&lt;/li&gt;
&lt;li&gt;oversized prompts hard-failed unless context shifting was configured;&lt;/li&gt;
&lt;li&gt;one server process loaded one GGUF model, unlike Ollama&amp;rsquo;s model manager;&lt;/li&gt;
&lt;li&gt;the first SYCL request paid a JIT compilation cost.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I kept Ollama plus Vulkan as the default. SYCL can improve throughput, but it
does not make a 3B model format better arguments, nor does it make a 7B model fit
in 4 GB.&lt;/p&gt;
&lt;h2 id=&#34;what-works-well-on-the-a310&#34;&gt;What works well on the A310
&lt;/h2&gt;&lt;p&gt;The A310 is useful for private chat, summarization and small experiments. Llama
3.2 3B was the best local default from my tests. It fitted well enough and could
emit a real tool-call structure. Qwen2.5 1.5B was faster but printed the tool call
as text.&lt;/p&gt;
&lt;p&gt;It was not reliable for the 13 Netdata MCP tools. Their schemas required a large
context and the models which remained usable on 4 GB VRAM could not consistently
produce valid tool arguments.&lt;/p&gt;
&lt;p&gt;The debugging order I use now is:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Confirm the accelerator backend actually loaded; do not infer GPU use from
container access to &lt;code&gt;/dev/dri&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Read the prompt token count and truncation log.&lt;/li&gt;
&lt;li&gt;Measure KV cache size and GPU layer placement at the chosen context.&lt;/li&gt;
&lt;li&gt;Separate prompt-evaluation latency from generation speed.&lt;/li&gt;
&lt;li&gt;Verify tool execution success in logs, not in UI decoration.&lt;/li&gt;
&lt;li&gt;Check whether the remaining problem is model capability rather than another
timeout or context setting.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I kept Ollama with Vulkan as the default because it supports model management and
normal local chat worked. llama.cpp with SYCL is available for comparison, but a
faster backend does not make the small model better at producing schema-valid
tool calls.&lt;/p&gt;
</description>
        </item>
        <item>
        <title>Keeping AI Agent Notes in My Homelab Repository</title>
        <link>https://blog.dexome.com/post/durable-memory-for-ai-coding-agent/</link>
        <pubDate>Fri, 05 Jun 2026 00:00:00 +0530</pubDate>
        
        <guid>https://blog.dexome.com/post/durable-memory-for-ai-coding-agent/</guid>
        <description>&lt;p&gt;I use AI coding agents regularly while working on my homelab repository. The
chat history is useful during a task, but a new session does not automatically
know what was proved in an older one.&lt;/p&gt;
&lt;p&gt;This became a problem for operational details. For example, one session found
the safe way to evaluate my Nix flake without copying 20 GiB of ignored model
files. Another found the exact SSH and TTY sequence required for privileged
diagnostics. I did not want to investigate those details again.&lt;/p&gt;
&lt;p&gt;I now store these notes as Markdown under &lt;code&gt;docs/agent-memory/&lt;/code&gt;. They are part of
the repository, so I can review and update them along with the configuration.&lt;/p&gt;
&lt;h2 id=&#34;what-i-wanted-from-these-notes&#34;&gt;What I wanted from these notes
&lt;/h2&gt;&lt;p&gt;The system needed to be:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;durable:&lt;/strong&gt; survive chat sessions, machines, and editor profiles;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;reviewable:&lt;/strong&gt; change through the same review process as code;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;routable:&lt;/strong&gt; let an agent find one relevant note without reading hundreds;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;historical:&lt;/strong&gt; preserve old decisions without presenting them as current;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;safe:&lt;/strong&gt; record where secrets live, never their values;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;cheap:&lt;/strong&gt; make adding a lesson easier than rediscovering it.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I did not want one large file which every agent had to read at the start of a
task. I also wanted the notes to remain visible outside one editor or agent.&lt;/p&gt;
&lt;h2 id=&#34;one-markdown-file-for-each-topic&#34;&gt;One Markdown file for each topic
&lt;/h2&gt;&lt;p&gt;Every memory is a dated Markdown file under &lt;code&gt;docs/agent-memory/&lt;/code&gt;:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;div class=&#34;chroma&#34;&gt;
&lt;table class=&#34;lntable&#34;&gt;&lt;tr&gt;&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;lnt&#34;&gt;1
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;2
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;3
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;
&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;2026-06-12-traefik-macvlan-asymmetric-return-route.md
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;2026-07-23-copyparty-rename-strips-group-write-arrs-import-denied.md
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;2026-08-10-netdata-2.10.3-collector-fixes.md
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;
&lt;/div&gt;
&lt;/div&gt;&lt;p&gt;The ISO date gives a useful filesystem order. The descriptive suffix makes the
file discoverable with ordinary text search. One topic per file lets a later
note supersede one conclusion without invalidating an unrelated section of a
large document.&lt;/p&gt;
&lt;p&gt;New notes carry queryable frontmatter:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;div class=&#34;chroma&#34;&gt;
&lt;table class=&#34;lntable&#34;&gt;&lt;tr&gt;&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;lnt&#34;&gt;1
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;2
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;3
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;4
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;5
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;6
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;7
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;8
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;9
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;
&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-yaml&#34; data-lang=&#34;yaml&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nn&#34;&gt;---&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;type&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;incident&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;title&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;Podman network drift recreate needs rm -f&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;description&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;Stopped containers remain associated with their networks.&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;timestamp&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;ld&#34;&gt;2026-06-12&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;tags&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;p&#34;&gt;[&lt;/span&gt;&lt;span class=&#34;l&#34;&gt;podman, networking]&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;resource&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;hosts/heavymetal/containers.nix&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;status&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;resolved&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nn&#34;&gt;---&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;
&lt;/div&gt;
&lt;/div&gt;&lt;p&gt;I normally write the body in this order:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;What was the symptom?&lt;/li&gt;
&lt;li&gt;What evidence separated it from similar failures?&lt;/li&gt;
&lt;li&gt;What was the root cause?&lt;/li&gt;
&lt;li&gt;What exact change fixed it?&lt;/li&gt;
&lt;li&gt;How was the fix verified?&lt;/li&gt;
&lt;li&gt;How can it be undone or superseded?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I include a failed attempt only when it is likely to be tried again. Normal
terminal exploration and command typos do not need to be stored.&lt;/p&gt;
&lt;h2 id=&#34;using-a-small-index-to-find-the-correct-note&#34;&gt;Using a small index to find the correct note
&lt;/h2&gt;&lt;p&gt;Opening every note at the start of every request would replace forgetting with
context overload. The repository instead has a hand-curated &lt;code&gt;index.md&lt;/code&gt;, grouped
by service or topic. Each entry is one complete problem-to-outcome sentence:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;div class=&#34;chroma&#34;&gt;
&lt;table class=&#34;lntable&#34;&gt;&lt;tr&gt;&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;lnt&#34;&gt;1
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;2
&lt;/span&gt;&lt;span class=&#34;lnt&#34;&gt;3
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;
&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-markdown&#34; data-lang=&#34;markdown&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;-&lt;/span&gt; [&lt;span class=&#34;nt&#34;&gt;2026-06-12-podman-network-drift-recreate-needs-force.md&lt;/span&gt;](&lt;span class=&#34;na&#34;&gt;...&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  — Recreating a drifted Podman network requires rm -f because stopped
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  containers remain associated with it. (_resolved_)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;
&lt;/div&gt;
&lt;/div&gt;&lt;p&gt;The agent&amp;rsquo;s startup rule is simple: read the index, select the few summaries
that match the current task, then open only those notes.&lt;/p&gt;
&lt;figure class=&#34;article-diagram&#34;&gt;
    &lt;pre class=&#34;mermaid&#34;&gt;
flowchart TB
	Q[New engineering request] --&gt; I[Read compact routing index]
	I --&gt; S{Select matching summaries}
	S --&gt; N1[Open relevant incident note]
	S --&gt; N2[Open relevant reference note]
	N1 --&gt; W[Work with prior evidence]
	N2 --&gt; W
	ALL[Hundreds of unrelated notes] -. not loaded .-&gt; W
&lt;/pre&gt;
    &lt;figcaption&gt;The routing index keeps startup context small while preserving access to detailed operational evidence.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I write these summaries by hand because a filename or heading usually does not
say which fix actually worked. A script checks that every note is included in
the index, but it does not generate the summary.&lt;/p&gt;
&lt;h2 id=&#34;not-every-note-needs-a-dated-file&#34;&gt;Not every note needs a dated file
&lt;/h2&gt;&lt;p&gt;Not every fact belongs in the dated-note catalog. I separate three classes:&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Memory class&lt;/th&gt;
          &lt;th&gt;Example&lt;/th&gt;
          &lt;th&gt;Lifetime&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;Incident or migration&lt;/td&gt;
          &lt;td&gt;Why macvlan replies used the wrong route&lt;/td&gt;
          &lt;td&gt;Historical, may be superseded&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Living reference&lt;/td&gt;
          &lt;td&gt;Port allocation inventory&lt;/td&gt;
          &lt;td&gt;Updated in place&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Durable service guidance&lt;/td&gt;
          &lt;td&gt;Stable operating constraint&lt;/td&gt;
          &lt;td&gt;Promoted into reference documentation&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;A port inventory would become noisy as a sequence of dated files. A completed
incident should not be silently rewritten whenever understanding changes.
For example, I update the port allocation inventory in place. An incident stays
as a dated file because a later note may supersede it without removing the old
evidence.&lt;/p&gt;
&lt;h2 id=&#34;marking-an-old-note-as-superseded&#34;&gt;Marking an old note as superseded
&lt;/h2&gt;&lt;p&gt;Deleting an old note destroys the path that explains why a decision existed.
Leaving it unmarked lets an agent follow obsolete instructions. The compromise
is bidirectional supersession.&lt;/p&gt;
&lt;p&gt;The old note says:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;div class=&#34;chroma&#34;&gt;
&lt;table class=&#34;lntable&#34;&gt;&lt;tr&gt;&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;lnt&#34;&gt;1
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;
&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-markdown&#34; data-lang=&#34;markdown&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;&amp;gt; &lt;/span&gt;&lt;span class=&#34;ge&#34;&gt;**Superseded by:** [2026-08-17-new-design.md](...) — reason.
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;
&lt;/div&gt;
&lt;/div&gt;&lt;p&gt;The new note says:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;div class=&#34;chroma&#34;&gt;
&lt;table class=&#34;lntable&#34;&gt;&lt;tr&gt;&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;lnt&#34;&gt;1
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;
&lt;td class=&#34;lntd&#34;&gt;
&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-markdown&#34; data-lang=&#34;markdown&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;&amp;gt; &lt;/span&gt;&lt;span class=&#34;ge&#34;&gt;**Supersedes:** [2026-06-05-old-design.md](...).
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;
&lt;/div&gt;
&lt;/div&gt;&lt;p&gt;Both frontmatter records and index entries carry the same state. The old
evidence remains available, but every retrieval path points toward current
guidance.&lt;/p&gt;
&lt;figure class=&#34;article-diagram&#34;&gt;
    &lt;pre class=&#34;mermaid&#34;&gt;
flowchart TB
	OLD[Old note: status superseded] --&gt;|superseded by| NEW[New note: current guidance]
	NEW --&gt;|supersedes| OLD
	INDEX[Routing index] --&gt;|current entry| NEW
	INDEX -. historical entry marked superseded .-&gt; OLD
&lt;/pre&gt;
    &lt;figcaption&gt;Supersession preserves the audit trail while making the current instruction unambiguous.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I used this when a temporary design was retired and when a shell test showed
that &lt;code&gt;umask 0002&lt;/code&gt; worked but the real container supervisor reset it. The old
test was still useful, but the new note needed to be the current instruction.&lt;/p&gt;
&lt;h2 id=&#34;telling-the-agent-when-to-read-the-notes&#34;&gt;Telling the agent when to read the notes
&lt;/h2&gt;&lt;p&gt;Notes do not help if the agent does not know when to read them. Repository
instructions establish a retrieval protocol:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;At the start of a request, read the routing index.&lt;/li&gt;
&lt;li&gt;Open only relevant notes.&lt;/li&gt;
&lt;li&gt;Consult living inventories before changing shared resources such as ports.&lt;/li&gt;
&lt;li&gt;Record a newly proven operational lesson in the repository.&lt;/li&gt;
&lt;li&gt;Update the index in the same change.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Without these repository instructions, the notes would only be documentation
which an agent might or might not find.&lt;/p&gt;
&lt;figure class=&#34;article-diagram&#34;&gt;
    &lt;pre class=&#34;mermaid&#34;&gt;
flowchart TB
	TASK[Task begins] --&gt; ROUTE[Index and instruction routing]
	ROUTE --&gt; PRIOR[Relevant prior knowledge]
	PRIOR --&gt; ACT[Implement and validate]
	ACT --&gt; LESSON{New durable lesson?}
	LESSON --&gt;|yes| NOTE[Write dated note and index summary]
	NOTE --&gt; CHECK[Run coverage and consistency checks]
	CHECK --&gt; TASK
	LESSON --&gt;|no| DONE[Finish]
&lt;/pre&gt;
    &lt;figcaption&gt;Work produces evidence, evidence becomes a note, and future instructions route the next task back through it.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id=&#34;checking-the-index&#34;&gt;Checking the index
&lt;/h2&gt;&lt;p&gt;The repository includes a small dependency-free checker. It compares dated
files on disk with links in the index and reports:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;notes missing from the index;&lt;/li&gt;
&lt;li&gt;index links whose files no longer exist;&lt;/li&gt;
&lt;li&gt;notes without frontmatter.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It exits nonzero for missing coverage or broken links. It does not generate
summaries, decide which topic heading is best, or resolve contradictory advice.
Those are semantic tasks.&lt;/p&gt;
&lt;p&gt;A periodic manual pass checks what code cannot reliably infer:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;stale &lt;code&gt;pending&lt;/code&gt; or &lt;code&gt;parked&lt;/code&gt; statuses;&lt;/li&gt;
&lt;li&gt;one-directional supersession links;&lt;/li&gt;
&lt;li&gt;newer notes that contradict older guidance;&lt;/li&gt;
&lt;li&gt;duplicate incidents without a relationship;&lt;/li&gt;
&lt;li&gt;ports or service references that drifted;&lt;/li&gt;
&lt;li&gt;lessons mature enough to move into stable documentation.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The script checks the structure. I still review the summaries and decide whether
one note replaces another.&lt;/p&gt;
&lt;h2 id=&#34;do-not-store-secrets-or-duplicate-the-code&#34;&gt;Do not store secrets or duplicate the code
&lt;/h2&gt;&lt;p&gt;Repository memory must never contain secret values. A note may say that an MQTT
password comes from a named SOPS secret and which service consumes it. It should
not contain the password, a token copied from a log, or an unredacted credential
example.&lt;/p&gt;
&lt;p&gt;I also avoid storing temporary chat details, large summaries of the repository
and facts already clear from the code. I add a note when it preserves a useful
check, an operational problem or a decision which would take time to reconstruct.&lt;/p&gt;
&lt;h2 id=&#34;current-workflow&#34;&gt;Current workflow
&lt;/h2&gt;&lt;p&gt;I do not use a vector database or an embedding pipeline for this. Markdown,
links, frontmatter, repository instructions and a small checker are enough for
my repository.&lt;/p&gt;
&lt;p&gt;The workflow is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;capture only proven lessons;&lt;/li&gt;
&lt;li&gt;compress each lesson into a routable sentence;&lt;/li&gt;
&lt;li&gt;load details only when relevant;&lt;/li&gt;
&lt;li&gt;preserve history through supersession;&lt;/li&gt;
&lt;li&gt;keep every memory visible to the people responsible for the system.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The agent still starts a new chat without the old conversation. It first reads
the index, opens the notes related to the current task and continues with the
checks and fixes already recorded there.&lt;/p&gt;
</description>
        </item>
        
    </channel>
</rss>
