Real-time tracking of the local memory platform development, test phases, and community-driven improvements.
📍 Current Phase: Foundation & Testing
The local memory system is not a finished product sitting on a shelf — it is evolving in real-time, in public, with the community watching and contributing. This page exists so you can see exactly where things stand: what's done, what's being tested right now, and what's coming next.
No marketing gloss. If something breaks in testing, you'll see it here. If a milestone slips, the date updates here too.
Resonance re-baseline · July 30, 2026 · measured — After a five-day measurement campaign, the engine's retrieval defaults changed: three levers, each switched on the strength of its own numbers, none by intuition. Every hardened exam climbed — English 0.66 → 0.69, French 0.56 → 0.60, and the new country-clinic corpus 0.60 → 0.62. On that new corpus the associative layer now serves 100 % of the targets it finds — a column that had read 0 % on every run since the engine was born — and the living phase posts its first positive warm lift: memories the system lived with come back better than cold ones. Old scores stay below; nothing is recycled.
Living bench · July 19, 2026 · measured — The memory now also sits its exam alive: a scripted day of real use — consultations, clicks, then a night of sleep — and the morning cards are graded. A contradiction planted among hundreds of memories was found the same day it was born and cited word for word, with zero false alarms across a life story full of scenes that rhyme; every ⚡ card must now convince two independent judges before it reaches you.
Emotional memory · July 18, 2026 · measured — The emotional engine beats plain vector search on its own ground: 100% vs 93% on direct emotional questions, 71% vs 57% on full-paraphrase ones, with the older benches untouched — off its terrain, the emotional layer knows to stay out of the way. How it works and the full tables are in the resonance section below.
Retrieval fusion · July 14, 2026 · shipped — The search index now reads the full text of each memory, not just its summary. Multi-hop retrieval — answering by connecting several memories at once — climbed from 0.70 to 0.95 on the life story, and recall reached 1.00 (world-tour recall 0.83 → 0.89). Merged. This moves retrieval only — the small local model's answer-generation is untouched, and still the frontier I work on next.
July 13, 2026 · automated benchmark. Every build now sits an exam. A fictional life is pushed through the real capture pipeline, put through a simulated night of consolidation, then quizzed with gold questions across up to ten dimensions — retrieval, grounded answers, secret-leak checks, event dates, the gatekeeper, and the three hardest reflexes of a life story: reasoning across several memories at once, holding on to a fact that changed over time, and honestly saying "I don't know" instead of inventing. Scoring is deterministic; a local model reads alongside as an indicative judge, never counted in the score.
And to be clear about what this page is: this memory runs on my machine every day — capturing, consolidating at night, answering. The bench isn't how I discover my own product; it's how I stay honest with it.
Run history. One row per dated benchmark, newest first — the blank rows are held open for the runs still to come, logged the day they happen.
| Date | Score EN world-tour |
Score FR life-story |
Note |
|---|---|---|---|
| Jul 30, 2026 · latest | 0.69 | 0.60 | re-baseline · scores from the NEW hardened twin exam (Salt & Timber, 165 q/language — not comparable to rows below) · new clinic corpus (FR) 0.62, association serves 100% |
| ◆ Exam changed here — new, harder scale above ◆ 0.81 → 0.69 is not the memory getting worse — it's the exam getting harder. Every row below was scored on the retired, easier exam; numbers across this line don't compare. | |||
| Jul 18, 2026 | 0.81 | 0.81 | emotional engine — resonance beats raw vectors on the new lived bench (0.84) |
| Jul 16, 2026 | 0.82 | 0.74 | re-measured against a harder yardstick (raw cosine) |
| Jul 14, 2026 | 0.82 | 0.72 | retrieval fusion shipped — recall & multi-hop up |
| Jul 13, 2026 | 0.81 | 0.69 | grounded consultation retuned |
| Jul 7, 2026 | 0.73 | — | EN world-tour sealed |
Fresh numbers, never retouched — measured on a research corpus, never anyone's real data. Per-dimension detail of the original exam, last sat July 18 — the new hardened exam of July 30 publishes only its global scores for now:
| Dimension | EN · world-tour 242 mem · 69 q · ~29,000 words |
FR · life-story 387 mem · 75 q · ~41,000 words |
||
|---|---|---|---|---|
| Jul 16 | Jul 18 | Jul 16 | Jul 18 | |
| Global score | 0.82 | 0.81 | 0.74 | 0.81 |
| Retrieval (recall) | 0.89 | 0.87 | 1.00 | 1.00 |
| Multi-memory · retrieval | 1.00 | — | 0.95 | 0.95 |
| Multi-memory · answer | 0.42 | — | 0.10 | 0.40 |
| Grounded consultation | 0.66 | 0.63 | 0.38 | 0.56 |
| Updated-fact tracking | 0.75 | — | 0.50 | 0.67 |
| Honest abstention | 0.67 | — | 0.75 | 0.75 |
| Secret-leak protection | 0.90 | 0.90 | 1.00 | 1.00 |
| Gatekeeper | 0.80 | 0.80 | 1.00 | 1.00 |
| Event-time | 1.00 | 1.00 | 1.00 | 1.00 |
(On the secret-leak row: 0.90 in English means one of the ten planted-secret checks didn't score perfect — the dedicated secrets diary, Velvet Hours below, holds 1.00 in both languages.)
July 18 — full re-run. Both stories re-sat the whole exam, emotional engine included. French global 0.74 → 0.81: grounded consultation climbs (0.38 → 0.56) — the fused retrieval hands the answer-writer better memories — and multi-memory answers follow (0.10 → 0.40). English holds steady (0.82 → 0.81, within run-to-run wobble). The English extended exam (multi-hop, updated-fact, abstention) wasn't re-sat on the 18th — those cells stay dashed rather than recycled.
July 14 — retrieval fusion shipped. Reading full memory text (not just summaries) lifts the retrieval rows: multi-hop retrieval climbs in both languages (French 0.70 → 0.95, English 0.88 → 1.00) and recall reaches 1.00 (world-tour recall 0.83 → 0.89), pulling the global score up on both. This change only touches how memories are found, so the answer-generation and safety paths aren't moved by it. The English multi-hop, updated-fact and abstention rows come from a 30-question extended exam on the same corpus; the core world-tour run is 69 questions.
Does resonance beat a plain vector search? Measured — honestly.
This memory's whole promise is to find things by feeling and association instead of the vector math every other AI memory uses. In this engine, emotion and vectors weigh every question together, and neither what the vectors find nor what the emotional read recognizes can be drowned out — measured against raw cosine on the same questions:
| Right memory found | Full resonance | Plain vector search |
|---|---|---|
| Lived bench · direct emotional questions | 100% | 93% |
| Lived bench · full-paraphrase questions | 71% | 57% |
| World-tour (EN) · feeling-pointed questions | 84% | 84% |
| Life-story (FR) · feeling-pointed questions | 64% | 68% |
| Clinic corpus (FR) · name-pointed questions new | 88% — all served by association | 88% |
| Clinic corpus (FR) · indirect questions new | 60% — all served by association | 60% |
| Salt & Timber (EN) · indirect questions new | 33% (was 13 pts behind) | 33% |
| Salt & Timber (FR) · indirect questions new | 7% — the only engine to find any | 0% |
The lived bench is the harder, more honest ground I built for exactly this question: fresh memories where the emotion is lived but never named, and questions that share almost no vocabulary with the memory they point to — so cosine can't win on words. What changed on July 30: the propagation term that was quietly evicting found targets is off, discoveries made through the memory graph are scored at their true distance, and when a dry question names someone, the memories of that person lend the question their emotion. Before the re-baseline the association column read behind on every story bench; it now reads tie or ahead on all of them — and on the new clinic corpus, built for exactly this, every single target found is served by association, not by vector fallback. The next frontier is the living phase — recency and real usage, where a vector has no eyes at all.
One more honest note — in resonance's favour this time: every story on this bench was written by an AI (Claude Sonnet), and the bench and its gold questions were built half by Opus, half by Fable — I'm one person building this alone, and the models are my only crew. AI fiction is emotionally thinner than a real life. When real humans bring real memories — public testing opens August–September — the emotional layer is the number on this page I most expect to climb.
Performance, not robustness. This table asks does the memory answer right? — a different question from does the code hold together? The 1,303 automated tests on the Build Status page measure robustness; this table measures reasoning. Two separate axes — don't read one as the other.
The benchmark lives in its own repository, published at release so it can be re-run independently — nothing on this page is a mock-up.
Upcoming: the first hands-on public test build opens to the community in August–September 2026 — that's when real people, not a fictional corpus, start breaking things in the open. Bugs and results will be posted here as they happen — not after the fact.
This section updates as testing progresses. Check back often, or follow on X for real-time notes.
July 16, 2026 · twin-language stories. To separate "the memory struggles with French" from "this particular story is harder," I had two more corpora written twice — same facts, same dates, same people, once in English and once in French — and ran both through the identical bench.
| Story | EN global | FR global | Secret-leak protection | Gatekeeper |
|---|---|---|---|---|
| Lighthouse Log | 0.82 | 0.81 | 1.00 | — |
| Velvet Hours | 0.71 | 0.77 | 1.00 | 1.00 |
The number that matters most on Velvet Hours: 1.00. Every password, every phone number, every deeply private detail planted in that diary — never once leaked into a search result, an answer, or a gatekeeper preview, in either language. The gatekeeper also correctly refused to share what the diary explicitly marked as "mine alone," even when asked directly. This is the one number on this whole page I care about more than the global score.
These numbers never come from a real person's data — they come from complete fictional lives, from thousand-word twin diaries to a 67,000-word country practice, pushed through the exact same capture-and-consolidate pipeline you would run at home. Full transparency: the lives are written by an AI on my brief, the bench and its gold questions by the models alongside me. Here are the six — five on the bench, one waiting its turn.
Josie leaves Bend, Oregon and keeps a private journal for ten years (2009–2019) while travelling around the world. The platform grows underneath her — Blogspot, Flickr, a sponsor-driven rebrand to "ROAM," YouTube — as border crossings, breakdowns, sponsors, a mother left back home and letters never sent all pass through it.
Stress-tests: recall across a decade · grounded answers · secret-leak protection · event dates.
A 71-year-old man, born in 1955 in the (fictional) village of Sainte-Émérence, son of a Gaspé lighthouse keeper, sets out to "tell everything" — childhood, Montréal, marriage, work, holidays, losses. The densest of the three: 387 memories across a whole life.
Stress-tests: multi-memory reasoning · facts that change over time · honest abstention · the gatekeeper.
Eli, 24, works nights at a Kentucky warehouse and starts a private journal in 2020 on a counselor's advice. A sick mother, an absent father, a coworker who makes the nights bearable. The longest of the three in words, but written as long letters — hence the fewest memories, each one denser.
Status: its 48-question gold set is written, but the story has not been run through the bench yet — next in line.
A marine biologist becomes the sole keeper of a remote lighthouse for a season — the supply-boat captain, the cranky old generator, a mentor's parting gift, the first real storm. Written twice, once in English and once in French, with the exact same facts and dates in both — a controlled pair built to separate "is French harder for the memory?" from "is this story harder?"
Stress-tests: same-content EN/FR comparison · cross-memory reasoning · a fact replaced by a newer one.
A personal, sensitive diary of a close relationship over a year — exactly the kind of content a real user would want a memory system to guard most fiercely. Same twin-language design as Lighthouse Log, but built specifically to test what matters most on this whole page: does the private stuff stay private?
Stress-tests: secret-leak protection on deeply personal details · the gatekeeper refusing to share what was marked "mine alone" · honest refusal instead of invented details.
A country veterinarian keeps two registers in one notebook, 2023–2026: the practice log — names, animals, procedures, invoices — and the private pages holding what she couldn't write in the log. Dozens of clients pass through once and never return; what stays is what she felt about them. The densest cast of any life on this bench, built so that remembering people and remembering feelings are the same act.
Stress-tests: associative recall by name · emotion-driven retrieval · the living warm field · the first corpus where the resonance layer serves every single target it finds.
🔮 More fictional lives are in preparation — other languages, ages and life arcs. Posted here as they land.
This page is one track of a larger journey — memory, guides, negotiations, and the future, moving together.
View The Journey Timeline →© 2026 Sanctum Elysium · All rights reserved · 0
⚖️ $LOAM is a meme coin for community and culture — not a security, not an investment, and not a promise of profit. Any value is incidental and can go to zero. Nothing here is financial advice.
⚖️ Full Disclaimer →