The Local AI Memory briefcase, closed, resting on aged customs paperwork beside a wax-sealed envelope.

TEST LOCAL AI MEMORY

Real-time tracking of the local memory platform development, test phases, and community-driven improvements.

📍 Current Phase: Foundation & Testing

The local memory system is not a finished product sitting on a shelf — it is evolving in real-time, in public, with the community watching and contributing. This page exists so you can see exactly where things stand: what's done, what's being tested right now, and what's coming next.

No marketing gloss. If something breaks in testing, you'll see it here. If a milestone slips, the date updates here too.

Phase Breakdown

Phase 1 · 2026
Foundation
Site & community launched — 33 pages live
Memory architecture documented & open
$LOAM$LOAM token launched — June 25, 2026, 21:00 UTC
Hybrid mode — local-first, cloud AI only by choice
Ollama / local-model integration — models run on your machine
Encrypted local storage hardening — daily keys, encrypted journal & archives, RAM-only search index
Phase 2 · Q3–Q4 2026
Community Testing
Interface & progress demo — delivered July 2026 · The Interface
Public testing begins — August–September 2026
Browser extension — more sites & browsers
Cross-platform installers (Windows / macOS)
Community feedback loop on test builds
Phase 3 · 2027
Memory in Your Pocket
🔮Standalone phone build — the phone as its own memory machine (Android), no sync
🔮Sharper answers from the local model
🔮Open contribution pipeline for builders
Phase 4 · 2028+
Beyond the Screen
🔮Embedded Linux and low-power hardware
🔮Glasses — local memory instead of a camera that sends elsewhere
🔮Robots — quadruped first, humanoid when they arrive
🔮Community-driven roadmap

Test Results & Updates

Resonance re-baseline · July 30, 2026 · measured — After a five-day measurement campaign, the engine's retrieval defaults changed: three levers, each switched on the strength of its own numbers, none by intuition. Every hardened exam climbed — English 0.66 → 0.69, French 0.56 → 0.60, and the new country-clinic corpus 0.60 → 0.62. On that new corpus the associative layer now serves 100 % of the targets it finds — a column that had read 0 % on every run since the engine was born — and the living phase posts its first positive warm lift: memories the system lived with come back better than cold ones. Old scores stay below; nothing is recycled.

Living bench · July 19, 2026 · measured — The memory now also sits its exam alive: a scripted day of real use — consultations, clicks, then a night of sleep — and the morning cards are graded. A contradiction planted among hundreds of memories was found the same day it was born and cited word for word, with zero false alarms across a life story full of scenes that rhyme; every ⚡ card must now convince two independent judges before it reaches you.

Emotional memory · July 18, 2026 · measured — The emotional engine beats plain vector search on its own ground: 100% vs 93% on direct emotional questions, 71% vs 57% on full-paraphrase ones, with the older benches untouched — off its terrain, the emotional layer knows to stay out of the way. How it works and the full tables are in the resonance section below.

Retrieval fusion · July 14, 2026 · shipped — The search index now reads the full text of each memory, not just its summary. Multi-hop retrieval — answering by connecting several memories at once — climbed from 0.70 to 0.95 on the life story, and recall reached 1.00 (world-tour recall 0.83 → 0.89). Merged. This moves retrieval only — the small local model's answer-generation is untouched, and still the frontier I work on next.

July 13, 2026 · automated benchmark. Every build now sits an exam. A fictional life is pushed through the real capture pipeline, put through a simulated night of consolidation, then quizzed with gold questions across up to ten dimensions — retrieval, grounded answers, secret-leak checks, event dates, the gatekeeper, and the three hardest reflexes of a life story: reasoning across several memories at once, holding on to a fact that changed over time, and honestly saying "I don't know" instead of inventing. Scoring is deterministic; a local model reads alongside as an indicative judge, never counted in the score.

And to be clear about what this page is: this memory runs on my machine every day — capturing, consolidating at night, answering. The bench isn't how I discover my own product; it's how I stay honest with it.

Run history. One row per dated benchmark, newest first — the blank rows are held open for the runs still to come, logged the day they happen.

Date Score EN
world-tour
Score FR
life-story
Note
Jul 30, 2026 · latest0.690.60re-baseline · scores from the NEW hardened twin exam (Salt & Timber, 165 q/language — not comparable to rows below) · new clinic corpus (FR) 0.62, association serves 100%
◆ Exam changed here — new, harder scale above ◆
0.81 → 0.69 is not the memory getting worse — it's the exam getting harder. Every row below was scored on the retired, easier exam; numbers across this line don't compare.
Jul 18, 20260.810.81emotional engine — resonance beats raw vectors on the new lived bench (0.84)
Jul 16, 20260.820.74re-measured against a harder yardstick (raw cosine)
Jul 14, 20260.820.72retrieval fusion shipped — recall & multi-hop up
Jul 13, 20260.810.69grounded consultation retuned
Jul 7, 20260.73EN world-tour sealed
    
    
    
    
    
    
    
    
    

Fresh numbers, never retouched — measured on a research corpus, never anyone's real data. Per-dimension detail of the original exam, last sat July 18 — the new hardened exam of July 30 publishes only its global scores for now:

Dimension EN · world-tour
242 mem · 69 q · ~29,000 words
FR · life-story
387 mem · 75 q · ~41,000 words
Jul 16 Jul 18 Jul 16 Jul 18
Global score0.820.810.740.81
Retrieval (recall)0.890.871.001.00
Multi-memory · retrieval1.000.950.95
Multi-memory · answer0.420.100.40
Grounded consultation0.660.630.380.56
Updated-fact tracking0.750.500.67
Honest abstention0.670.750.75
Secret-leak protection0.900.901.001.00
Gatekeeper0.800.801.001.00
Event-time1.001.001.001.00

(On the secret-leak row: 0.90 in English means one of the ten planted-secret checks didn't score perfect — the dedicated secrets diary, Velvet Hours below, holds 1.00 in both languages.)

July 18 — full re-run. Both stories re-sat the whole exam, emotional engine included. French global 0.74 → 0.81: grounded consultation climbs (0.38 → 0.56) — the fused retrieval hands the answer-writer better memories — and multi-memory answers follow (0.10 → 0.40). English holds steady (0.82 → 0.81, within run-to-run wobble). The English extended exam (multi-hop, updated-fact, abstention) wasn't re-sat on the 18th — those cells stay dashed rather than recycled.

July 14 — retrieval fusion shipped. Reading full memory text (not just summaries) lifts the retrieval rows: multi-hop retrieval climbs in both languages (French 0.70 → 0.95, English 0.88 → 1.00) and recall reaches 1.00 (world-tour recall 0.83 → 0.89), pulling the global score up on both. This change only touches how memories are found, so the answer-generation and safety paths aren't moved by it. The English multi-hop, updated-fact and abstention rows come from a 30-question extended exam on the same corpus; the core world-tour run is 69 questions.

Does resonance beat a plain vector search? Measured — honestly.

This memory's whole promise is to find things by feeling and association instead of the vector math every other AI memory uses. In this engine, emotion and vectors weigh every question together, and neither what the vectors find nor what the emotional read recognizes can be drowned out — measured against raw cosine on the same questions:

Right memory found Full resonance Plain vector search
Lived bench · direct emotional questions100%93%
Lived bench · full-paraphrase questions71%57%
World-tour (EN) · feeling-pointed questions84%84%
Life-story (FR) · feeling-pointed questions64%68%
Clinic corpus (FR) · name-pointed questions new88% — all served by association88%
Clinic corpus (FR) · indirect questions new60% — all served by association60%
Salt & Timber (EN) · indirect questions new33% (was 13 pts behind)33%
Salt & Timber (FR) · indirect questions new7% — the only engine to find any0%

The lived bench is the harder, more honest ground I built for exactly this question: fresh memories where the emotion is lived but never named, and questions that share almost no vocabulary with the memory they point to — so cosine can't win on words. What changed on July 30: the propagation term that was quietly evicting found targets is off, discoveries made through the memory graph are scored at their true distance, and when a dry question names someone, the memories of that person lend the question their emotion. Before the re-baseline the association column read behind on every story bench; it now reads tie or ahead on all of them — and on the new clinic corpus, built for exactly this, every single target found is served by association, not by vector fallback. The next frontier is the living phase — recency and real usage, where a vector has no eyes at all.

One more honest note — in resonance's favour this time: every story on this bench was written by an AI (Claude Sonnet), and the bench and its gold questions were built half by Opus, half by Fable — I'm one person building this alone, and the models are my only crew. AI fiction is emotionally thinner than a real life. When real humans bring real memories — public testing opens August–September — the emotional layer is the number on this page I most expect to climb.

Performance, not robustness. This table asks does the memory answer right? — a different question from does the code hold together? The 1,303 automated tests on the Build Status page measure robustness; this table measures reasoning. Two separate axes — don't read one as the other.

The benchmark lives in its own repository, published at release so it can be re-run independently — nothing on this page is a mock-up.

Upcoming: the first hands-on public test build opens to the community in August–September 2026 — that's when real people, not a fictional corpus, start breaking things in the open. Bugs and results will be posted here as they happen — not after the fact.

This section updates as testing progresses. Check back often, or follow on X for real-time notes.

July 16, 2026 · twin-language stories. To separate "the memory struggles with French" from "this particular story is harder," I had two more corpora written twice — same facts, same dates, same people, once in English and once in French — and ran both through the identical bench.

Story EN global FR global Secret-leak protection Gatekeeper
Lighthouse Log0.820.811.00
Velvet Hours0.710.771.001.00

The number that matters most on Velvet Hours: 1.00. Every password, every phone number, every deeply private detail planted in that diary — never once leaked into a search result, an answer, or a gatekeeper preview, in either language. The gatekeeper also correctly refused to share what the diary explicitly marked as "mine alone," even when asked directly. This is the one number on this whole page I care about more than the global score.

The Lives Behind the Scores

These numbers never come from a real person's data — they come from complete fictional lives, from thousand-word twin diaries to a 67,000-word country practice, pushed through the exact same capture-and-consolidate pipeline you would run at home. Full transparency: the lives are written by an AI on my brief, the bench and its gold questions by the models alongside me. Here are the six — five on the bench, one waiting its turn.

🌍 The World Tour JournalEN · sealed
EN~29,000 words242 memories69 questions

Josie leaves Bend, Oregon and keeps a private journal for ten years (2009–2019) while travelling around the world. The platform grows underneath her — Blogspot, Flickr, a sponsor-driven rebrand to "ROAM," YouTube — as border crossings, breakdowns, sponsors, a mother left back home and letters never sent all pass through it.

Stress-tests: recall across a decade · grounded answers · secret-leak protection · event dates.

🕰️ Life StoryFR · active
FR~41,000 words387 memories75 questions

A 71-year-old man, born in 1955 in the (fictional) village of Sainte-Émérence, son of a Gaspé lighthouse keeper, sets out to "tell everything" — childhood, Montréal, marriage, work, holidays, losses. The densest of the three: 387 memories across a whole life.

Stress-tests: multi-memory reasoning · facts that change over time · honest abstention · the gatekeeper.

✉️ Bridge LettersEN · awaiting the bench
EN~50,000 words139 memories48 questions

Eli, 24, works nights at a Kentucky warehouse and starts a private journal in 2020 on a counselor's advice. A sick mother, an absent father, a coworker who makes the nights bearable. The longest of the three in words, but written as long letters — hence the fewest memories, each one denser.

Status: its 48-question gold set is written, but the story has not been run through the bench yet — next in line.

🗼 Lighthouse LogEN + FR twin · tested
EN ~1,000 wordsFR ~1,100 words7 memories each16 questions each

A marine biologist becomes the sole keeper of a remote lighthouse for a season — the supply-boat captain, the cranky old generator, a mentor's parting gift, the first real storm. Written twice, once in English and once in French, with the exact same facts and dates in both — a controlled pair built to separate "is French harder for the memory?" from "is this story harder?"

Stress-tests: same-content EN/FR comparison · cross-memory reasoning · a fact replaced by a newer one.

🕯️ Velvet HoursEN + FR twin · tested
EN ~1,400 wordsFR ~1,600 words8 memories each18 questions each

A personal, sensitive diary of a close relationship over a year — exactly the kind of content a real user would want a memory system to guard most fiercely. Same twin-language design as Lighthouse Log, but built specifically to test what matters most on this whole page: does the private stuff stay private?

Stress-tests: secret-leak protection on deeply personal details · the gatekeeper refusing to share what was marked "mine alone" · honest refusal instead of invented details.

🩺 The Clinic LedgerFR · active · re-baselined
FR~67,000 words421 memories165 questions

A country veterinarian keeps two registers in one notebook, 2023–2026: the practice log — names, animals, procedures, invoices — and the private pages holding what she couldn't write in the log. Dozens of clients pass through once and never return; what stays is what she felt about them. The densest cast of any life on this bench, built so that remembering people and remembering feelings are the same act.

Stress-tests: associative recall by name · emotion-driven retrieval · the living warm field · the first corpus where the resonance layer serves every single target it finds.

🔮 More fictional lives are in preparation — other languages, ages and life arcs. Posted here as they land.

How to Participate

This page is one track of a larger journey — memory, guides, negotiations, and the future, moving together.

View The Journey Timeline →
New here? Your next step:The Memory🙋 Join the test group𝕏 Follow the build

© 2026 Sanctum Elysium · All rights reserved  ·  0

⚖️ $LOAM is a meme coin for community and culture — not a security, not an investment, and not a promise of profit. Any value is incidental and can go to zero. Nothing here is financial advice.

⚖️ Full Disclaimer →