Why I'm building Antumbra
I spend most of my time making AI agents useful, and almost none of that time is inside the model. The useful part lives in the harness around it: lifecycle hooks, MCP tools, re-injected context and markdown, a memory system bolted on so the agent behaves with some continuity from one session to the next. I built one of those harnesses. It works. What kept nagging me is that all of that competence lives outside the model, and I re-pay for it on every single call.
Capability you rent is capability you re-pay for on every call. Capability you own is capability that compounds.
That is the lens I judge these systems by now: not how impressive the demo is, but the cost of one useful, correct output, measured over time. By that measure most AI systems are a bad trade. They bolt a rented flagship onto everything, spend frontier prices on work they have already done a thousand times, and still hand back confident, wrong answers at full freight.
Antumbra is what I'm building instead. It's a private substrate you plug into a coding agent you already use (Claude Code, Cursor, any MCP client). Instead of merely remembering, it gets better: it trains your own verified outcomes into a population of small, frozen specialists over one shared base model, and it learns where each specialist's competence ends. The whole thing runs on my own hardware, and the data that teaches it never leaves the building.
The pattern I got tired of
A normal AI assistant is one enormous, general model in someone else's data center. Three things follow from that, and all three are about value, not ideology. You send your data out to get an answer. The model is exactly as good at your work tomorrow as it is today, because it never learns it. And the thousandth time it does a routine task costs the same as the first.
The industry's answer is the harness: pile more context into the prompt, bolt on a better notebook, re-explain the world every session. I've shipped that pattern, and it buys real continuity. But it is continuity rented by the token. Every session, the same conventions and history get re-injected and re-charged, and none of it ever becomes the model's own.
Antumbra's bet is to push the harness down into the
substrate. Continuity becomes trained, routed
competence plus a private memory the agent boots from,
so a session starts already knowing "this repo uses
deno" without anyone paying to re-say it. There is
no cold start, and the part that used to be a recurring
bill starts becoming an owned asset.
A population of small experts, not one big brain
The base is one frozen, code-capable model (Qwen2.5-Coder-1.5B-Instruct) on a single 24 GB consumer GPU, a 3090 Ti in my case. That constraint is deliberate. I wanted to start as low-level as the hardware allows and prove the idea small, rather than assume a cluster I don't have.
Over that base sits a growing library of specialists, each one a small LoRA adapter that masters a single narrow, recurring job. A specialist isn't a new model; it's a thin overlay on the shared base. They're cheap to train and cheap to keep, so you can have hundreds. And when one graduates, it is frozen: immutable for good. Immutability is the only hard guarantee I know of that a learned skill is never silently smudged when the system trains the next one. The usual way a single fine-tuned model "learns" task B is by quietly getting worse at task A. Freezing is how I refuse that trade.
The cost story falls straight out of this. You stop
renting a flagship for routine work, because routine
work is exactly what graduates into a local specialist
first. The specialists are tiny, and 4-bit QLoRA keeps
them to about a quarter of the memory, which is what
lets the whole thing fit on one card. The trainer and
server are written in Rust on candle, single
process, no Python plane to babysit.
The part I'm proudest of architecturally: the portable asset is the memory, not the adapters. A specialist is a LoRA overlay on one specific base, but the verified history it grew from is yours independent of any model. When a stronger small model lands, and in 2026 they keep landing, I re-derive my specialists from the memory I already banked instead of starting over.
Knowing when it doesn't know
The most valuable thing an AI system can do is know where its own competence ends.
Most systems only ever accumulate what works. The
neglected and, I'd argue, more valuable half is the
boundary of a rule: the fact that a behavior is right
here and wrong one step over, and which feature
decides. "Use deno install, not npm install" is
true in this repo and false in the next one. A good
apprentice doesn't memorize "always do X"; they learn
"do X when we're in this situation."
That sense of scope is what powers the gate. For each incoming task it computes a relative-coverage margin, how strongly the best-matching specialist covers the task versus the runner-up, minus an inhibition term for known failure boundaries. Inside a regime we've proven we cover, it answers locally, for free. Out of distribution, it doesn't bluff; it escalates.
match gate.route(&task)? {
// inside a regime we've proven: answer locally, free
Coverage::Covered(expert) => serve.answer(expert, &task).await?,
// out of distribution: don't guess, escalate deliberately
Coverage::Escalate(reason) => frontier.answer(&task, reason).await?,
}This is where "worth its output" stops being a slogan. You never pay flagship prices for work a 1.5B local expert already nails, and you never pay for confident-but-wrong output, because out-of-scope means abstain and escalate, not invent. Knowing the limit of what you know turns out to be a cost control as much as a quality one. The system even names the governing feature of a boundary itself, by finding the context key whose values among passing runs are disjoint from the failing ones, rather than waiting for me to label the scope by hand.
Reality is the teacher
Antumbra learns from verifiable outcomes in an environment, not by imitating a teacher's text. For coding over your own repos, the reward is the test that passes, the command that runs, the build that goes green, the draft you accept. That's it. The training signal is reward-ranked fine-tuning over the completions that actually verified.
// reality is the teacher: reward only what verified
let reward = match verifier.check(&candidate, &task) {
Outcome::Verified => 1.0, // the test passed; the build is green
Outcome::Failed(_) => 0.0, // a critic may explain why, but it
// cannot turn this into a pass
};A critic exists, but it has a strict job: turn a
failure into a diagnostic ("npm install failed
because this is a Deno project") and name the
governing feature of the boundary. It never rescues a
verified-failed step. Verifiers are ground truth.
Crucially, training happens on the verified outcome,
not on the critic's words, which keeps the learning
grounded in real data and clean of "trained on a
provider's outputs." A frontier model, if it's used at
all, is an optional cold-start accelerator, not the
thing being copied.
Even routing is earned rather than declared. An expert's capability vector is the centroid of the prompts it provably solved in its final training round, so routing is retrieval over demonstrated behavior instead of a hand-written description of what a specialist is supposed to be good at. This is also why coding over your own repositories is the ideal first domain: it's maximally verifiable because you can run it, maximally context-scoped because per-repo conventions are textbook boundaries, and the data is yours.
Your data becomes weights you own
Privacy here isn't a setting you toggle; it's the shape of the system. The data that trains your specialists never leaves your boundary, and even the embedding step that touches your content sits behind a pluggable port so it can stay on your side. You can run the whole engine fully offline, an embedded store and a loopback MCP with a single identity, where nothing ever leaves the machine, or as a private hosted tenant.
The sovereignty argument is the one I actually care about: you end up owning the behavior, the capability, and the data. Your conventions, the things that today live as re-injected system-prompt text, composition over inheritance, builder-only database access, no emojis, become trained, owned boundaries instead of instructions you keep paying to repeat.
When it does run multi-tenant, isolation lives in the
database engine, not in handler code. Record access
binds (tenant, user) to $auth, and every row is
filtered by a permission rule the engine enforces:
-- generated by surql-rs builders, never hand-authored
DEFINE TABLE memory SCHEMALESS PERMISSIONS
FOR select WHERE tenant_id = $auth.tenant
AND (compartment = NONE
OR compartment IN (SELECT VALUE key FROM compartment
WHERE owner = $auth.user)
OR compartment IN (SELECT VALUE compartment FROM grant
WHERE grantee = $auth.user));A forgotten app-side filter cannot leak data, because the app isn't the thing enforcing the boundary. That rule is generated by my own SurrealDB client library's schema builders, never hand-written, so there is no string-built query and no injection surface, the same schema-as-code discipline I wrote about in the surql client article, doing real security work here. Compartments are ownable spaces of memory you can grant and revoke, where a revocation is a tombstone that fails closed immediately and propagates across devices, and a private compartment can consolidate into a private expert that other tenants can't even see, yet routes for you.
Context that compounds instead of repeating
Underneath the experts is a memory store with three networks, world for facts, bank for experiences, opinion for judgments, each with vector recall, reinforcement counts, and provenance. It does two jobs: it boots the agent with relevant context on the way in, and it's the raw material specialists later graduate from.
Consolidation is where context turns into capability. A memory's consolidation score is recurrence × verifiability × stability; the trusted ones graduate into experts with an interleaved replay buffer that resists forgetting, and an expert is retired when a memory it was built on is later contradicted. The memory forgets falsified theses on purpose, instead of letting stale "facts" calcify.
Recall itself is built to be honest about what dense vectors miss. It's hybrid: a sparse BM25 leg beside the 384-d dense vectors, fused and then re-ranked by a cross-encoder, because exact tokens, code identifiers, error codes, tickers, are precisely the things a single embedding silently drops. The embedder stays a generalist deliberately, since this substrate isn't only for code. The net effect is the thing ordinary "AI memory" can't do: the scaffolding shrinks. The more the system metabolizes, the less context you re-stuff, and capability builds up instead of you re-paying the look-it-up-and-re-explain cost on every call.
What it's for
The first and primary use is a coding agent that is actually yours. You bolt Antumbra onto the assistant you already drive; it becomes the brain, your tool stays the hands. A session boots already knowing your repo's conventions and history, answers from the local population when it can, escalates the genuinely new, and on the way out writes back what verified, so the proven, repeating work quietly graduates into a permanent specialist for next time.
Because the substrate is general, it backs more than coding for me. The same memory, gate, and verification discipline drive a market-narrative pipeline and an event-driven trading brain, where "frozen experts plus abstain-when-out-of-distribution" becomes strategy specialists with risk management as architecture: the system simply doesn't trade when it has no edge. In the narrative pipeline, a distilled sentiment specialist replaces a per-item LLM call outright, which is a direct, measurable cost win, one local adapter instead of a remote inference per document.
And it runs two ways on purpose: entirely on your own machine, where nothing leaves, or as a private hosted service others can sign up for, where each tenant's materials sit in separate locked cabinets the engine enforces. For privacy-sensitive work, medical, legal, financial, that property is often the whole point.
What's true today
I hold every milestone to a falsifiable experiment with a kill criterion, so I'll be precise about what has and hasn't held up.
The full loop is real and runs end to end on the GPU: route, serve, verify, capture, consolidate, graduate a frozen expert, refresh the router. Engine-enforced isolation is validated live over the wire. But most of it is proven at under a small verifiable scale, small adapters, small corpora, a 1.5B base. The thing I have not yet demonstrated is large-scale generalization together with a measured no-catastrophic-forgetting result across many experts. That is the next chapter, and it's evidence, not construction. Antumbra is slowly rolling towards this conclusion.
Goals I'm Chasing, Bets I'm Making
The north star is genuinely separate experts wired together by learned cross-attention rather than just routed between, but currently only a linear-blend precursor exists. Automatic boundary discovery works but is still stochastic, and the bottleneck is the hardware, not the algorithm.
This is also somewhat a bet on the future of this industry. The hardware trend is towards more, smaller models, not fewer, bigger ones, and the software trend is towards more efficient training and inference on consumer GPUs, not more expensive frontier prices. If that holds, then my theory of change contends that "small experts on a shared base" is the future-pattern and will increasingly more viable over time. The value of owning your own specialists becomes immeasurable to IP-concerned businesses, and the cost of renting proprietary flagships to owning being able to own that expertise across any model and domain.
Requirements
To run all of Antumbra's current features, including the training the creates the specialists, you need:
- By platform (YMMV):
- For Windows/Linux: a 8GB VRAM GPU (e.g., NVIDIA 30-series) with WSL2 and CUDA toolkit installed.
- For macOS: an Apple Silicon M1 Pro or later with at least 16GB of unified memory.
- At least 16GB of system RAM.
- More of the above is better if you want to train larger specialists or run more of them concurrently.
What if I don't have the hardware?
You can still gain benefit from Antumbra by using the memory layer and the gate to route to a frontier model, which is a default fallback. The system is designed for graceful degradation, so even without a GPU you can use the same interfaces and still get value from the learned routing, memory, and verification layers. You won't get the compounding competence of the specialists without their training, but you can benefit from persistant permanent context that turns your agent from a stateless oracle into a learning apprentice.
Competence vs Time
One caveat that is worth stating plainly: Antumbra is not a magic wand that turns a 1.5B model into a 100B one. It is a system for owning and compounding competence, but that competence has to be earned through verifiable outcomes. So while it wants a capable GPU and a learning-in period, it is not an open-ended oracle, and it only pays off if you do enough repeating, checkable work to be worth instantiating specialists.
Why this is the work I want to be doing
I didn't set out to beat a frontier model on a benchmark. I would lose that competition while being limited to consumer hardware. I set out to own competence and it's boundaries that prove to compound, and to make sure every token the system spends is one worth spending while forcing cost downwards over time, not up.
That's the throughline of everything above. The client library I wrote pushes the substrate lower than harnesses are willing to go. Verification makes the learning provable. The boundary makes the spending transparent, escalate when you must, answer locally when you can, and never bluff in between. An AI system is only worth building if its output is worth more than it cost to produce. Antumbra is my attempt to make that true by design, not by accident.
