<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>Agensphere Journal</title>
  <link>https://agensphere.com/journal</link>
  <atom:link href="https://agensphere.com/journal/feed.xml" rel="self" type="application/rss+xml" />
  <description>Engineering notes, teardowns and decision records from Agensphere: how we build AI systems that hold up in production, layer by layer.</description>
  <language>en</language>
  <lastBuildDate>Tue, 06 Oct 2026 09:00:00 GMT</lastBuildDate>
  <item>
    <title>The owned intelligence layer: what we build instead of another enterprise AI plan</title>
    <link>https://agensphere.com/journal/the-owned-intelligence-layer</link>
    <guid isPermaLink="true">https://agensphere.com/journal/the-owned-intelligence-layer</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>An enterprise plan gives every employee a powerful model. It does not route work to the right model, know your business, enforce your data rules, record what happened or stay yours when the market moves. That needs an enterprise AI platform you own, and it is what Agensphere builds.</description>
  </item>
  <item>
    <title>No audit trail: when nobody can explain what the AI did</title>
    <link>https://agensphere.com/journal/ai-audit-trail</link>
    <guid isPermaLink="true">https://agensphere.com/journal/ai-audit-trail</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>A customer disputes a decision, a regulator asks how it was made, or an agent did something unexpected overnight. Without an AI audit trail showing which model, prompt, context and tools produced the outcome, you cannot defend it, fix it or learn from it.</description>
  </item>
  <item>
    <title>Pilot purgatory: why enterprise AI pilots never reach production</title>
    <link>https://agensphere.com/journal/why-ai-pilots-never-reach-production</link>
    <guid isPermaLink="true">https://agensphere.com/journal/why-ai-pilots-never-reach-production</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>The demo impressed everyone, the pilot ran for three months, and then nothing shipped. AI pilots stall on the way to production for predictable reasons: no measure of success, no owner, no integration with real systems, and no answer for risk. Each one can be designed out from day one.</description>
  </item>
  <item>
    <title>Avoid LLM vendor lock-in with a model-agnostic seam</title>
    <link>https://agensphere.com/journal/avoid-llm-vendor-lock-in</link>
    <guid isPermaLink="true">https://agensphere.com/journal/avoid-llm-vendor-lock-in</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>DECISION</category>
    <description>Models are deprecated, repriced and overtaken every few months. When provider-specific calls are scattered through your code, every change is a migration: that is LLM vendor lock-in. One thin seam, with your own prompts and evals behind it, turns switching models into a configuration change.</description>
  </item>
  <item>
    <title>Shadow AI: when every team buys its own model</title>
    <link>https://agensphere.com/journal/shadow-ai-every-team-buys-its-own-model</link>
    <guid isPermaLink="true">https://agensphere.com/journal/shadow-ai-every-team-buys-its-own-model</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>Shadow AI is what happens when marketing has one AI subscription, support another and engineering three. Each team solved its own problem, and the company now pays several times for the same capability with no shared context, no shared controls and no view of the total.</description>
  </item>
  <item>
    <title>Data compliance with LLMs: know where every prompt goes</title>
    <link>https://agensphere.com/journal/llm-data-compliance-where-your-data-goes</link>
    <guid isPermaLink="true">https://agensphere.com/journal/llm-data-compliance-where-your-data-goes</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>Every prompt is a data transfer. Customer records, contracts and employee details leave your systems the moment they are pasted into a model. LLM data privacy starts with a gateway that sees, redacts, routes and records every call.</description>
  </item>
  <item>
    <title>Unowned intelligence: when your AI capability lives in someone else's product</title>
    <link>https://agensphere.com/journal/unowned-intelligence</link>
    <guid isPermaLink="true">https://agensphere.com/journal/unowned-intelligence</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>If your prompts, evaluation data, usage logs and AI workflows live inside a vendor's interface, you are renting a capability you think you are building. AI ownership means the models can be rented while the intelligence around them stays yours.</description>
  </item>
  <item>
    <title>Out of context: why your enterprise AI doesn't know your business</title>
    <link>https://agensphere.com/journal/out-of-context-enterprise-ai</link>
    <guid isPermaLink="true">https://agensphere.com/journal/out-of-context-enterprise-ai</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>A general model knows the internet, not your pricing rules, your customers or last week's policy change. Without a deliberate context layer, enterprise AI gives fluent, generic answers that are wrong for your company.</description>
  </item>
  <item>
    <title>Paying frontier prices for routine work: the enterprise LLM routing problem</title>
    <link>https://agensphere.com/journal/enterprise-ai-paying-frontier-prices-for-routine-work</link>
    <guid isPermaLink="true">https://agensphere.com/journal/enterprise-ai-paying-frontier-prices-for-routine-work</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>Most enterprise AI traffic is classification, extraction and short summaries. Sending all of it to the most expensive model is the most common way companies overpay for AI. LLM routing, matching each task to the cheapest model that passes your eval, fixes it.</description>
  </item>
  <item>
    <title>Agent memory, planning and stopping: what turns a loop into a system</title>
    <link>https://agensphere.com/journal/agent-memory-planning-and-stopping</link>
    <guid isPermaLink="true">https://agensphere.com/journal/agent-memory-planning-and-stopping</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>GUIDE</category>
    <description>A basic agent remembers only its message list, plans one step at a time and stops at an arbitrary limit. Real agents manage their memory, plan when the task is long, and stop on purpose. Here is how each part works, and where to go from here.</description>
  </item>
  <item>
    <title>How AI agents work: a model, tools and a loop</title>
    <link>https://agensphere.com/journal/how-ai-agents-work</link>
    <guid isPermaLink="true">https://agensphere.com/journal/how-ai-agents-work</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>GUIDE</category>
    <description>An AI agent is a model that can call tools, running in a loop until a task is done. That is the whole idea. This walkthrough builds one from scratch, then shows when a fixed workflow is the better choice.</description>
  </item>
  <item>
    <title>How tool calling works: a model never runs a function, it asks you to</title>
    <link>https://agensphere.com/journal/how-tool-calling-works</link>
    <guid isPermaLink="true">https://agensphere.com/journal/how-tool-calling-works</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>GUIDE</category>
    <description>Tool calling, also called function calling, lets a model request an action by returning a structured call that your code executes. Understanding that division of labour is the key to understanding agents, and to keeping them safe.</description>
  </item>
  <item>
    <title>How RAG works, step by step</title>
    <link>https://agensphere.com/journal/how-rag-works-step-by-step</link>
    <guid isPermaLink="true">https://agensphere.com/journal/how-rag-works-step-by-step</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>GUIDE</category>
    <description>Retrieval-augmented generation (RAG) answers questions from your own documents by finding the relevant passages first and handing them to the model. Two pipelines, one for indexing and one for answering, and every quality problem lives in one of their steps.</description>
  </item>
  <item>
    <title>Why LLMs hallucinate, and what actually reduces it</title>
    <link>https://agensphere.com/journal/why-llms-hallucinate</link>
    <guid isPermaLink="true">https://agensphere.com/journal/why-llms-hallucinate</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>GUIDE</category>
    <description>LLMs hallucinate because a model generates the most plausible continuation, not the true one. When the facts are missing or blurry in its weights, plausible and true come apart. Grounding, verification and permission to say “I don't know” reduce it far more than prompts asking for accuracy.</description>
  </item>
  <item>
    <title>Sampling explained: temperature, top-p and why the same prompt gives different answers</title>
    <link>https://agensphere.com/journal/sampling-temperature-top-p-explained</link>
    <guid isPermaLink="true">https://agensphere.com/journal/sampling-temperature-top-p-explained</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>GUIDE</category>
    <description>A model does not output an answer. It outputs a probability for every possible next token, and a sampling step picks one. LLM temperature and top-p control that pick, and choosing them well is part of engineering a reliable system.</description>
  </item>
  <item>
    <title>How LLMs are trained: pretraining, fine-tuning and preference tuning</title>
    <link>https://agensphere.com/journal/how-llms-are-trained</link>
    <guid isPermaLink="true">https://agensphere.com/journal/how-llms-are-trained</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>GUIDE</category>
    <description>LLMs are trained in three stages: they learn language and facts by predicting trillions of next tokens, learn to follow instructions from curated examples, and learn which answers people prefer from comparisons. Each stage explains a different behaviour you see in production.</description>
  </item>
  <item>
    <title>Attention and the transformer, explained without the hand-waving</title>
    <link>https://agensphere.com/journal/attention-and-the-transformer-explained</link>
    <guid isPermaLink="true">https://agensphere.com/journal/attention-and-the-transformer-explained</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>GUIDE</category>
    <description>Attention is how each token decides which earlier tokens matter to it. Stack attention with a small feed-forward network, repeat it dozens of times, and you have the transformer that powers every modern LLM.</description>
  </item>
  <item>
    <title>Tokens and embeddings: how text becomes numbers</title>
    <link>https://agensphere.com/journal/tokens-and-embeddings-how-text-becomes-numbers</link>
    <guid isPermaLink="true">https://agensphere.com/journal/tokens-and-embeddings-how-text-becomes-numbers</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>GUIDE</category>
    <description>A model never sees letters. It sees token IDs, and then vectors. Knowing what a token is in an LLM explains pricing, context limits and odd failures with spelling and numbers; embeddings explain how a model represents meaning.</description>
  </item>
  <item>
    <title>How a large language model works, end to end</title>
    <link>https://agensphere.com/journal/how-a-large-language-model-works</link>
    <guid isPermaLink="true">https://agensphere.com/journal/how-a-large-language-model-works</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>GUIDE</category>
    <description>A large language model does one thing: given some text, it predicts the next token. Everything else, from chat to code to agents, is built on repeating that step. This walkthrough shows how large language models work by following one request from your keyboard to the reply.</description>
  </item>
  <item>
    <title>Put the model inside the workflow, not in a chat box beside it</title>
    <link>https://agensphere.com/journal/put-the-model-inside-the-workflow</link>
    <guid isPermaLink="true">https://agensphere.com/journal/put-the-model-inside-the-workflow</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>Most AI features fail on adoption, not accuracy. Good AI product design puts intelligence at the exact step where a decision is made, instead of in a chat panel beside the real work.</description>
  </item>
  <item>
    <title>Cut LLM cost in the right order: cache, route, batch, then trim</title>
    <link>https://agensphere.com/journal/cut-llm-cost-in-the-right-order</link>
    <guid isPermaLink="true">https://agensphere.com/journal/cut-llm-cost-in-the-right-order</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>Most attempts to reduce LLM cost start with prompt golf and end with a quality regression. Measure first, then take the structural wins: prompt caching, model routing and batch processing, before rewriting a single prompt.</description>
  </item>
  <item>
    <title>Human approval is a durable wait, not a blocking call</title>
    <link>https://agensphere.com/journal/human-approval-is-a-durable-wait</link>
    <guid isPermaLink="true">https://agensphere.com/journal/human-approval-is-a-durable-wait</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>DECISION</category>
    <description>Human-in-the-loop approval should not hold a thread, a socket or a model session while a manager decides. Persist the run, send the request, and resume from the log when the answer arrives, hours or days later.</description>
  </item>
  <item>
    <title>MCP turns agent tools into a protocol, so treat each server as a service</title>
    <link>https://agensphere.com/journal/mcp-turns-tools-into-a-protocol</link>
    <guid isPermaLink="true">https://agensphere.com/journal/mcp-turns-tools-into-a-protocol</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>The Model Context Protocol standardises how models discover and call tools, read resources and use prompt templates. That removes integration glue, and moves the hard questions to permissions, versioning and trust.</description>
  </item>
  <item>
    <title>Prompt injection is a trust problem, not a prompting problem</title>
    <link>https://agensphere.com/journal/prompt-injection-is-an-input-trust-problem</link>
    <guid isPermaLink="true">https://agensphere.com/journal/prompt-injection-is-an-input-trust-problem</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>No system prompt reliably stops prompt injection: a model can follow instructions hidden in an email or a web page. Limit what untrusted text can cause instead: separate data from instructions, scope tools to the user, and gate every risky action.</description>
  </item>
  <item>
    <title>Validate structured outputs at the boundary, then retry with the error</title>
    <link>https://agensphere.com/journal/validate-structured-outputs-at-the-boundary</link>
    <guid isPermaLink="true">https://agensphere.com/journal/validate-structured-outputs-at-the-boundary</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>DECISION</category>
    <description>JSON mode guarantees the shape of LLM structured outputs, not whether they are true. Validate the schema and the business rules where model output enters your system, and feed the exact error back for one bounded retry.</description>
  </item>
  <item>
    <title>Hybrid search beats pure vector search for most RAG systems</title>
    <link>https://agensphere.com/journal/hybrid-search-beats-pure-vector-search</link>
    <guid isPermaLink="true">https://agensphere.com/journal/hybrid-search-beats-pure-vector-search</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>Embeddings are good at meaning and bad at exact strings: order numbers, error codes and product names need keyword search. Hybrid search runs both, merges the results with reciprocal rank fusion, and re-ranks the top of the list.</description>
  </item>
  <item>
    <title>Chunk by document structure, not by token count</title>
    <link>https://agensphere.com/journal/chunk-by-structure-not-token-count</link>
    <guid isPermaLink="true">https://agensphere.com/journal/chunk-by-structure-not-token-count</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>Fixed 512-token chunks split tables mid-row and separate claims from the context that makes them true. A RAG chunking strategy that follows headings, sections and lists is the cheapest retrieval improvement most systems never make.</description>
  </item>
  <item>
    <title>A context window is not memory</title>
    <link>https://agensphere.com/journal/a-context-window-is-not-memory</link>
    <guid isPermaLink="true">https://agensphere.com/journal/a-context-window-is-not-memory</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>A bigger context window lets a model read more on each request. It does not give an LLM memory. Treating the two as the same thing is how agents get slow, expensive and forgetful at once.</description>
  </item>
  <item>
    <title>Trace and cost every model call from day one</title>
    <link>https://agensphere.com/journal/trace-and-cost-every-model-call</link>
    <guid isPermaLink="true">https://agensphere.com/journal/trace-and-cost-every-model-call</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>DECISION</category>
    <description>If a request gets slow or expensive and you cannot say which step did it, you cannot fix it. LLM observability starts with one span per model and tool call, with tokens and cost attached: cheap to add early, painful to add late.</description>
  </item>
  <item>
    <title>Treat a prompt change like a code change: ship it with an eval</title>
    <link>https://agensphere.com/journal/ship-a-prompt-change-with-an-eval</link>
    <guid isPermaLink="true">https://agensphere.com/journal/ship-a-prompt-change-with-an-eval</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>A prompt is code that runs on someone else's computer and changes behaviour without a compiler error. LLM evaluation gives it a test set, a score and a review gate, so model upgrades stop being surprises.</description>
  </item>
  <item>
    <title>Filter by permission before you rank, not after</title>
    <link>https://agensphere.com/journal/filter-by-permission-before-you-rank</link>
    <guid isPermaLink="true">https://agensphere.com/journal/filter-by-permission-before-you-rank</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>NOTE</category>
    <description>Retrieval that ranks first and checks access later is a RAG security hole: it leaks data through snippets, scores and empty answers. Put the permission filter inside the search query and the whole class of bug goes away.</description>
  </item>
  <item>
    <title>Every agent tool needs an effect class before it ships</title>
    <link>https://agensphere.com/journal/every-agent-tool-needs-an-effect-class</link>
    <guid isPermaLink="true">https://agensphere.com/journal/every-agent-tool-needs-an-effect-class</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>DECISION</category>
    <description>If you cannot say what happens when a tool call runs twice, the tool is not ready to give to a model. Four questions, asked before an AI agent tool ships, decide how the runtime is allowed to retry it.</description>
  </item>
  <item>
    <title>KEEL: make the run deterministic, not the model</title>
    <link>https://agensphere.com/journal/keel-make-the-run-deterministic</link>
    <guid isPermaLink="true">https://agensphere.com/journal/keel-make-the-run-deterministic</guid>
    <pubDate>Tue, 06 Oct 2026 09:00:00 GMT</pubDate>
    <category>TEARDOWN</category>
    <description>A refund agent crashed after Stripe moved the money. KEEL is the durable execution runtime we are building so the answer to “how many refunds?” is always one, and so any run can be replayed or forked onto another model.</description>
  </item>
</channel>
</rss>
