My mom has published 710 pieces in 16 years. Nothing she wrote links to 526 of them.
In May my mom sent me a poster she had made: a graveyard at dusk, six tombstones carrying titles of her own work, and her line across the top, don't treat your content like a graveyard, treat it like a library. She asked for everything she has written in one place, and something to point her back at it when she starts something new. I built her a social media scheduler instead. This weekend I built the thing she actually asked for: one page, 710 pieces, and an open model that runs inside her browser tab so her unpublished drafts never leave her laptop. She graded it without knowing, because the answer key is every link she made by hand between 2011 and 2026.

Remember the app idea I had?
My mom is Mona Andrei. She writes humour. She has kept a blog called Moxie-Dude since 2010, she writes a column for Westmount Magazine, she runs a Substack for single moms, and she wrote a book called Superwoman. In May she sent me a poster she had made: a graveyard at dusk, six tombstones, and on each tombstone the title of something she had written and where it ran. Across the top was her own line. Under it she typed, "Remember the app idea I had?" She wanted two things. Everything she has written in one place, and something that would point her back at what she already wrote when she sits down to write something new. I built her a social media scheduler. It had a landing page and a logo and it ran on sample data. She found a product online that already did that and texted me: "Found it. I will find the hole." This weekend I built what she asked for the first time.
Her graveyard, counted
I pulled everything public: 650 posts from her blog, 37 columns from Westmount Magazine and 23 pieces from her Substack. That is 710 pieces and 353,232 words, cut into 2,809 passages of about 120 words each. Then I counted the links she had made by hand from one of her pieces to another. 526 of the 710 have nothing of hers pointing at them. On her blog she linked back to an earlier post 179 times, and 150 of those point at something from the previous 30 days. In sixteen years she reached back more than a year nine times. Left to memory, a writer links to what she wrote last week. The poster was right, and it was more right than she knew: three of the six pieces on her tombstones are in the library, and all three are in the ground.

The desk
On the left is a sheet of paper. She writes a draft there, or pastes one. When she pauses, the draft is cut into passages the same way her published work was, each one is embedded inside the tab, and every piece she has published is scored by its closest passage. That is all the matching there is. What comes back is five of her own pieces, each with her own paragraph quoted, where it ran, how many years ago, and a button that copies a link she can paste straight into WordPress or Substack. In the recording I typed two sentences about setting off a smoke alarm and ordering pizza. Those are my test sentences, not hers. What comes back is hers: five pieces written between 2013 and 2016, starting with the one about why she should never try to cook red meat again. Nothing on the page is written for her. Every suggestion is a real link and a passage cut out of her own text.

She wrote the answer key without knowing it
This is the part I care most about, because the grader is not me. The answer key is every link she made by hand between 2011 and 2026. The test takes a post where she linked back to an earlier one, removes the link and its text, treats what is left as a draft, asks for every earlier post ranked by closeness, and sees where the post she actually linked to lands. Most of her links are no test of memory at all, because 150 of 179 point at the previous 30 days and showing the newest post first finds them. The 19 that reach back more than 90 days are the ones that matter. Newest-first puts her real target in the top five once. Keyword search gets 7. The model that ships in the page gets 9. Nineteen is a small number and nine against seven is not a difference I would defend, which is exactly what the submission says. I also tried fusing the model with keyword search, which got 10, and did not ship a second ranker for one extra hit. All of these figures come out of the exact vectors the page loads, not from a separate run.

It says when there is nothing
If she writes about something she has never covered, the page says the subject is new and shows nothing, instead of padding the list with the five least-bad matches. Getting that rule right took two attempts and the first one was useless. My first cut-off was a similarity level that nine in ten of her real links cleared, which sounds reasonable and is not, because the nearest unrelated passage usually scores higher than the post she actually linked to, so the page would never have said no to anything. The rule that shipped comes out of her archive instead: for every piece she has published, how close was the closest earlier one? Nineteen in twenty clear a certain level. A draft below that level is further from her past work than almost anything she has ever published, and that is when the page says the subject is new. It is a rule taken from her own writing, not a guarantee.

Her draft never leaves the tab
An unpublished draft is the one thing a writer does not hand to a server she does not control, and here there is no server to hand it to. The model is all-MiniLM-L6-v2, open weights, 23 MB at 8-bit, running through transformers.js on WebAssembly in a web worker inside her browser tab. Its files sit beside the page rather than on a CDN. The whole index of her writing is 2,809 vectors of 384 numbers each, stored as one byte per number, which is 1.08 MB. The page counts its own requests to any other origin and prints the number at the bottom: it is 0. Turn the Wi-Fi off and it still answers, which is not a claim in a README but a check in real Chromium that switches the browser offline and asks again. It also costs nothing to run, which matters more than it sounds, because she is not going to pay a monthly fee to search her own writing and I am not going to pay per request so that she can.

One line per piece, kept only if it is word for word hers
A second open model picks the sentence most worth quoting from each piece. Gemma 3 4B, running locally through Ollama, reads a piece and proposes one line. Then a string comparison decides, and the model does not get a vote: the line is kept only if it can be found word for word in that piece, and what gets stored is the span cut from her text rather than Gemma's retyping of it. It was asked about 694 pieces. 560 lines were kept, 131 of those only after one correction. 134 were thrown out: 23 were not word for word hers, 18 were the wrong length, 91 ran on past the limit, and twice it answered that there was no line, which is an allowed answer. For the other 9 pieces the local server returned an error every time, so they have no line at all, and the page simply shows none. To be clear about the division of labour, because it is easy to blur: Gemma proposes pull lines, and the matching in the page is all-MiniLM.

Checks that fail out loud
Two scripts decide whether any of this is allowed to ship. check_site.mjs opens the page in real Chromium and checks what the page promises: the model loads from local files, a passage of hers finds its own piece first, the quoted passage is word for word from the library, the Treaty of Westphalia gets "Nothing close", the page still answers with the network switched off, any piece can be opened from the keyboard, and not one request leaves the page's own origin. Fourteen checks. check_claims.py reads every number in the README and in the submission post and compares it against the data, and on its first run it caught a stale word count in my own plan. I keep rebuilding this script on every project because a number written on Saturday and a pipeline re-run on Sunday is exactly how a writeup ends up lying about its own data, and I would rather the build tell me than a reader.

What is not claimed
Nine against seven on a sample of nineteen is not a result, and I will not dress it as one. What the test supports is narrow: a 23 MB open model in a browser tab does at least as well as keyword search at the one thing she cannot do from memory, which is reaching back more than three months. It matches subjects, not jokes, and it cannot tell whether a link would be welcome. It reads passages of about 120 words, so a one-line aside can slip past it. It only knows what was public on three outlets when the library was built, and three more outlets named on her own poster are missing because their sites would not let me read them. The cut-off for saying a subject is new is a level taken from her archive, not a guarantee. I did not run a closed embedding model against her links, so I am not going to claim open won on raw accuracy; what open did that closed could not is run inside her tab, offline, with her draft going nowhere. And on whether she likes it: nothing yet. She said yes to being the example, but she has not sat down with it, so there is no quote from her here. I would rather leave that blank than write one for her, which is also the rule the page itself follows.
The Moxie Library: 16 years of my mom's writing in one page, searched by an open model inside her own browser tab
View the project