JOURNALENGINEERING NOTES, TEARDOWNS, DECISION RECORDS33 ENTRIES
Working
notes.
What we write down while building: teardowns of our own prototypes, the decisions behind them, and short notes on the engineering that keeps AI systems standing. Public, because the reasoning is the proof.
TYPE
LAYER
START HERE · FOUNDATIONSNew to LLMs and agents? Ten walkthroughs, in order, from tokens to production.
- 01How a large language model works, end to end4 MIN
- 02Tokens and embeddings: how text becomes numbers4 MIN
- 03Attention and the transformer, explained without the hand-waving4 MIN
- 04How LLMs are trained: pretraining, fine-tuning and preference tuning5 MIN
- 05Sampling explained: temperature, top-p and why the same prompt gives different answers4 MIN
- 06Why LLMs hallucinate, and what actually reduces it4 MIN
- 07How RAG works, step by step4 MIN
- 08How tool calling works: a model never runs a function, it asks you to4 MIN
- 09How AI agents work: a model, tools and a loop4 MIN
- 10Agent memory, planning and stopping: what turns a loop into a system4 MIN
FOR TEAMS RUNNING AI · ENTERPRISE AIWhere enterprise AI goes wrong: cost, context, compliance and ownership. And what to build instead.
- 01Paying frontier prices for routine work: the enterprise LLM routing problem3 MIN
- 02Out of context: why your enterprise AI doesn't know your business4 MIN
- 03Unowned intelligence: when your AI capability lives in someone else's product3 MIN
- 04Data compliance with LLMs: know where every prompt goes4 MIN
- 05Shadow AI: when every team buys its own model3 MIN
- 06Avoid LLM vendor lock-in with a model-agnostic seam3 MIN
- 07Pilot purgatory: why enterprise AI pilots never reach production3 MIN
- 08No audit trail: when nobody can explain what the AI did3 MIN
- 09The owned intelligence layer: what we build instead of another enterprise AI plan5 MIN

KEEL: make the run deterministic, not the model
A refund agent crashed after Stripe moved the money. KEEL is the durable execution runtime we are building so the answer to “how many refunds?” is always one, and so any run can be replayed or forked onto another model.
AGENSPHERE EDITORIAL · 4 MINREAD →
J-033The owned intelligence layer: what we build instead of another enterprise AI planNOTE5 MINJ-032No audit trail: when nobody can explain what the AI didNOTE3 MINJ-031Pilot purgatory: why enterprise AI pilots never reach productionNOTE3 MINJ-030Avoid LLM vendor lock-in with a model-agnostic seamDECISION3 MINJ-029Shadow AI: when every team buys its own modelNOTE3 MINJ-028Data compliance with LLMs: know where every prompt goesNOTE4 MINJ-027Unowned intelligence: when your AI capability lives in someone else's productNOTE3 MINJ-026Out of context: why your enterprise AI doesn't know your businessNOTE4 MINJ-025Paying frontier prices for routine work: the enterprise LLM routing problemNOTE3 MINJ-024Agent memory, planning and stopping: what turns a loop into a systemGUIDE4 MINJ-023How AI agents work: a model, tools and a loopGUIDE4 MINJ-022How tool calling works: a model never runs a function, it asks you toGUIDE4 MINJ-021How RAG works, step by stepGUIDE4 MINJ-020Why LLMs hallucinate, and what actually reduces itGUIDE4 MINJ-019Sampling explained: temperature, top-p and why the same prompt gives different answersGUIDE4 MINJ-018How LLMs are trained: pretraining, fine-tuning and preference tuningGUIDE5 MINJ-017Attention and the transformer, explained without the hand-wavingGUIDE4 MINJ-016Tokens and embeddings: how text becomes numbersGUIDE4 MINJ-015How a large language model works, end to endGUIDE4 MINJ-014Put the model inside the workflow, not in a chat box beside itNOTE3 MINJ-013Cut LLM cost in the right order: cache, route, batch, then trimNOTE4 MINJ-012Human approval is a durable wait, not a blocking callDECISION3 MINJ-011MCP turns agent tools into a protocol, so treat each server as a serviceNOTE3 MINJ-010Prompt injection is a trust problem, not a prompting problemNOTE3 MINJ-009Validate structured outputs at the boundary, then retry with the errorDECISION2 MINJ-008Hybrid search beats pure vector search for most RAG systemsNOTE3 MINJ-007Chunk by document structure, not by token countNOTE4 MINJ-006A context window is not memoryNOTE3 MINJ-005Trace and cost every model call from day oneDECISION3 MINJ-004Treat a prompt change like a code change: ship it with an evalNOTE3 MINJ-003Filter by permission before you rank, not afterNOTE3 MINJ-002Every agent tool needs an effect class before it shipsDECISION3 MIN