Serenity: a private assistant that learns
Created
Serenity is my personal assistant agent. She is built on Hermes, an open-source agent framework, and I talk to her in a chat app. This page is about how she is built and how she gets better from use. It is deliberately not about what she does for me day to day.
Why I built her
I wanted three things from an assistant. First, answers that come from my own records, not from what a model happens to recall. Second, that she runs on hardware I own. Third, that she improves from use without me hand-tuning prompts, but never by quietly changing herself in ways I can’t see.
How she’s built
One line from her workspace sums up the design: she reasons, tools act, the record records, memory remembers, Hermes connects.
- Hermes connects. It handles the channels and the scheduling, and holds no business logic.
- She reasons. The model decides what should happen and which tool to call.
- Tools act. Small tools do the work, and they validate every write. The model never writes state directly.
- The record records. An append-only record is the source of truth for anything that is counted or dated.
- Memory remembers. A semantic memory holds patterns and context, but it is never the source of a count. When memory and the record disagree, the record wins.
She runs on a guest in my homelab rather than on my laptop, so the laptop sleeping no longer stops her. Her code ships as versioned releases with a single deploy owner and a rollback path.
Other agents can reach her too. A coding agent asks her one bounded thing at a time, with a lock so only one talks to her at once. When something needs a decision from me, she relays the question to me and passes my answer back, so I never have to sit in an agent’s terminal approving things.
How she learns
The learning loop has a few small parts:
- A private exchange log. A privacy-filtered log of her exchanges is kept on her guest. Some areas are excluded from it or kept text-free.
- Flagging a miss. I can mark a message to record that she got something wrong or couldn’t help.
- A weekly self-review. She reads the week, updates a list of gaps and writes a sanitized summary.
- A pair review. A Claude and a Codex agent then review that summary in panes I can watch and propose ranked picks.
- My pick. I choose. Picks become planned work; the rest go to an idea queue. Declined ideas don’t come back under new names, because the review matches gaps by meaning.
- A check that the loop runs. If a week passes without a pair review, I get told.
Corrections and shortcuts
Two kinds of improvement happen between the weekly reviews, each with a firm limit.
Corrections become lessons. Local rules, not a model, detect a correction. It takes an explicit marker plus a contradiction of a claim that can be checked. Facts and preferences she applies herself. Corrections about how to do something go to the weekly pair instead. A lesson is added only to the conversation turn where it applies; it is never written into her configuration or memory. About a week later a recheck marks it as held, recurred or not enough evidence.
Repeated questions become shortcuts. When I ask a similar question three or more times on at least two days, she can save a small shortcut: a text-only answer recipe, announced once before it is used. A shortcut never adds a tool, a job or a write. Anything bigger, such as an instruction or a new tool, becomes a proposal for the weekly pair and my pick.
The detectors had targets before they shipped: precision 0.90, recall 0.70, and zero reads of conversations that are sealed from the loop. They met them on their test fixtures.
The shape of it
In words, the diagram is:
me ⇄ chat app ⇄ Hermes ⇄ Serenity (the model) → calls → tools → write → the record, with memory beside her; the exchange log → weekly self-review → Claude and Codex pair review → my picks → planned work
What I learned
- Split reasoning from acting. Tools validate; the model doesn’t write state.
- A record beats memory for anything counted.
- Draw the line for self-improvement. Narrow per-turn lessons and announced read-only recipes she can do herself. Any change to code, tools or procedures goes through a visible review and my pick.
- Add lessons per turn; don’t rewrite the prompt. A wrong lesson stays small and easy to remove.
- Set precision targets before shipping a detector. Simple local rules were enough to detect corrections and repeats, and they are easy to check.
- Some conversations are sealed from learning entirely, by design.
- MVP first: get the mechanics working end to end before hardening them.