The original draft speced a hypothetical per-project agent.sh on a per-project cron. Reality turned out different: there's already a production bulk deploy script on the host (fixed project array, docker compose up -d per project, no explicit pull — freshness comes entirely from pull_policy: always) — so this patches that existing loop in place instead of introducing a parallel per-project script that would duplicate its array/failure-tracking logic. Also folds in a real prerequisite gap found while reconciling the proposal with the actual script: pull_policy: always isn't universal across the fleet yet, independent of keep entirely — the bulk script can't redeploy anything missing that line regardless of keep adoption. One instance of this was already fixed this session (a sibling project's docker-compose.yml missed the line when copying another project's compose pattern). Reframes §5's announce/watch benefit accordingly: since the real script is one all-or-nothing sweep, not N independent watchers, the natural fit is waking the whole sweep early (watch-any), not per-project selective redeploy — still deliberately left unbuilt until there's a real signal it's worth it over the already-scheduled sweep. Project names anonymized per policy — generic proj1..proj12 placeholders instead of the real fleet. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
15 KiB
Proposal: announce + a self-deploying agent, with a no-keep fallback
Status: proposed, not built. Self-contained — written so a fresh agent with no prior conversation context can pick this up and implement it without anything explained first.
Context (read this first)
Full background on keep is in README.md and
IMPLEMENTATION.md — read both before touching
code. The relevant existing pieces for this proposal:
-
src/server/db.ts— SQLite (better-sqlite3, WAL mode).vaults,recipients,vault_grants,vault_grants_previous,access_log. -
src/server/routes.ts— recipient-authenticated (signed-request) routes:push,pull,grant,log, plusexistsandrecipients(both used by the CLI to compute who to re-wrap for). -
src/server/auth.ts—requireRecipientAuth(Ed25519 signature over method+path+timestamp+body-hash, verified againstreq.originalUrlwith the query string stripped — see IMPLEMENTATION.md's "What was verified end-to-end" for a bug that already bit this exact mechanism once) andrequireAdminAuth(X-Admin-Passwordheader). -
src/cli/—identity.ts,recipient.ts,vault.ts,signedFetch.ts(note:signedFetch'spathparam must never include a query string — see the comment there for why). -
scripts/deploy.sh(this repo) — pulls<project>/<env>fromkeepinto$DEPLOY_ROOT/<project>/.env, thendocker compose up -dfor one project. Assumes the vault already exists and this machine already has a grant on it. Superseded by the real deploy script below as the thing this proposal actually patches — kept in this repo as a single-project example/fallback, not the production path. -
The real, already-running deploy script — on the host, not in any repo. A fixed-array bulk script, already in production use, looping over roughly a dozen project directories:
ROOT="$(cd "$(dirname "$0")" && pwd)" PROJECTS=(proj1 proj2 proj3 proj4 proj5 proj6 proj7 proj8 proj9 proj10 proj11 proj12) for project in "${PROJECTS[@]}"; do dir="$ROOT/$project" [ -f "$dir/docker-compose.yml" ] || { echo "skip: $project"; continue; } (cd "$dir" && docker compose up -d) || FAILED+=("$project") done docker image prune -fTwo things about it that matter for this proposal: it never runs
docker compose pull— freshness comes entirely from each project's ownpull_policy: alwaysin itsdocker-compose.yml(up -dre-pulls implicitly). And it loops over a fixed, explicit array, not directory auto-discovery — adding a project means editing this array by hand. This is the actual thing to patch, not a new per-project script; see Design §4. -
Prerequisite gap, found while writing this proposal, not yet fixed everywhere:
pull_policy: alwaysisn't universal across the fleet — confirmed present in roughly half the projects checked, confirmed missing in several others (one of which was fixed in this session — a sibling project'sdocker-compose.ymldidn't have it, a regression from copying another project's compose file without carrying this line over), and unchecked for the rest (no local checkout available at proposal time). This is independent ofkeepentirely — the bulk script can't redeploy any project missing this line regardless of whether that project ever adoptskeep, and it should be swept fleet-wide before or alongside this work, not as part of it. -
.gitea/workflows/docker.yml(present inkeepand other sibling projects) — builds and pushes an image to the Gitea container registry on every push tomain. Does not currently deploy anything — that's the manual/scheduled bulk script above today.
Motivation
The CI workflow already produces a fresh image on every push. Deployment
is still a manual "SSH in and run the deploy script" step. The obvious
fix — CI SSHes into the VPS and runs the deploy — was considered and
rejected: it means a standing SSH credential with production-deploy
rights sits in Gitea's CI secrets, reachable by whatever's executing on
the shared runner across every project it builds. A compromised build
step in any one project could then reach every other project's deploy
target. That's a real regression from where keep already left things
(see IMPLEMENTATION.md's resolved "read no longer implies write" — the
whole point was not concentrating reach into one identity).
The fix that doesn't reintroduce that problem: reverse the direction.
CI never reaches into a VPS. Instead, CI tells keep "there's a new
image for this project," and each VPS — which already has to run
something locally to redeploy itself — polls keep for that signal
and acts on it locally, using an identity and grant it already legitimately
holds. keep's server never executes anything remotely; it stays a pure
relay, same posture as everything else it does.
Not every project will have been migrated to keep yet. Bootstrapping
a project (register a recipient, grant it, push the first vault) is a
real, deliberate step per project, and there'll be a period — maybe a
long one — where some projects have it and some don't. The redeploy path
has to work identically well for both, defaulting to exactly what
happens today (a plain docker compose up -d, relying on that project's
own pull_policy: always) for anything not yet bootstrapped, rather than
either blocking deploys entirely or hard-failing on missing vaults.
Design
1. announce — a distinct verb from push/pull
An image tag isn't a secret. It doesn't belong inside the encrypted vault
payload alongside real credentials — that's a category mismatch (deploy
metadata living inside a secrets blob), not just an implementation
convenience either way. announce is its own small mechanism:
CREATE TABLE IF NOT EXISTS vault_announcements (
vault_key TEXT PRIMARY KEY,
tag TEXT NOT NULL,
announced_at INTEGER NOT NULL,
announced_by TEXT NOT NULL -- recipient_id
);
One row per vault — like the vault's own one-previous-version retention,
this is "latest known state," not a history. If a real audit trail of
every announce is wanted later, access_log already records the action;
this table is just the fast-path "what's the current tag" read, kept
separate from the log on purpose (reading current state back out of an
append-only log is the wrong tool for that job — same reasoning
IMPLEMENTATION.md already gives for why vault_grants_previous is a
real table and not a log replay).
2. API
POST /api/vaults/:key/announce -- { tag: string } — requires ANY grant on the vault (read-only is enough; announcing isn't a secrets mutation)
GET /api/vaults/:key/announce -- { tag, announcedAt, announcedBy } | 404 if nothing announced yet — requires a grant, same as pull
Both go through requireRecipientAuth, identical shape to the existing
routes. POST only requires a grant (not write) — reuse hasGrant,
not hasWriteGrant — since this isn't a mutation of the vault's secret
content, just a signal, and CI's identity for a project should be able to
stay read-only even after this feature exists. (This is worth stating
explicitly since a naive implementation might reach for hasWriteGrant
by analogy with push — resist that; it would silently widen CI's
required permissions for no real reason.)
3. CLI
keep announce <vault> --tag <tag> -- POST, requires a grant on <vault>
keep watch <vault> [--interval 60] -- polls GET .../announce every N seconds, prints the tag on change, exits nonzero if the vault doesn't exist / no grant (see fallback below)
keep watch is deliberately dumb — plain polling, not a websocket or
long-poll. Matches the polling pattern already used elsewhere in this
project family, and a 60-second detection lag for a deploy is fine.
Don't build push-based notification for this; it's more machinery for a
problem polling already solves adequately.
4. Patch the existing loop — don't build a parallel script
This is the part the fallback requirement is actually about. The real
deploy script already does the one thing that matters (a loop over a
fixed project list, docker compose up -d per project, relying on
pull_policy: always for freshness) — the change is inserting one
keep-with-fallback step per iteration, in place, not writing a new
per-project agent.sh that duplicates the loop/array/failure-tracking
logic the real script already has:
for project in "${PROJECTS[@]}"; do
dir="$ROOT/$project"
[ -f "$dir/docker-compose.yml" ] || { echo "skip: $project"; continue; }
echo ""
echo "▶ $project"
# Best-effort: does this machine have a working keep grant for
# "$project/production"? If not — not bootstrapped onto keep yet, or
# never will be (a static site with no secrets) — fall through to
# exactly what this script already did before keep existed. Every
# keep call here is allowed to fail; failure means "keep doesn't
# apply to this project," never "stop the deploy."
if keep pull "$project/production" --out "$dir/.env.new" 2>/dev/null; then
mv "$dir/.env.new" "$dir/.env"
else
rm -f "$dir/.env.new"
fi
if (cd "$dir" && docker compose up -d 2>&1); then
echo " ✓ $project up"
else
echo " ✗ $project failed"
FAILED+=("$project")
fi
done
Vault key convention: <project>/production, using the exact directory
name already in PROJECTS — no separate mapping table to keep in sync.
Notice announce/watch aren't actually load-bearing in this version —
whatever already triggers this script (cron, systemd timer, or still
run by hand) keeps triggering it exactly as before, and keep pull
either updates .env or silently doesn't. That's deliberate: the
existing script, patched this way, is already a complete, correct
baseline — keep-bootstrapped projects get fresh secrets on every run,
everything else behaves identically to today. announce sits on top as
a latency optimization for however this script gets triggered, not a
requirement for this patch to be complete — see the next section.
5. Where announce/watch actually earn their keep
Two independent benefits, worth keeping conceptually separate — and the second one looks different than it would for a one-agent-per-project design, because the real script is one bulk all-or-nothing sweep, not N independent watchers:
- Secrets sourcing —
keep pullinstead of a hand-placed.env. Available per-project the moment its vault exists and this machine has a grant on it. This is what §4's patch already delivers, on its own, with zero dependency onannounce. - Waking the sweep early — however the bulk script currently gets
triggered (cron, systemd timer, or still by hand),
announcecould let it wake immediately on a push instead of waiting for the next scheduled run. The natural fit given the script's actual shape:keep watch-any <vault1> <vault2> ...blocking until any watched vault's tag changes, used as the trigger for a full sweep — not a per-project selective redeploy. Selective ("only redeploy the one project that actually changed") would require reshaping the script away from its current fixed-array-sweep model into something that takes a target project as an argument, which is a real behavior change to something already in production use, not a small addition — out of scope here.
Deliberately not building the wake-the-sweep-early mechanism in this
pass. Ship §4's patch first — it already satisfies the fallback
requirement completely and delivers the secrets-sourcing benefit on its
own. Whether waiting up to one scheduled interval for a redeploy is
actually a problem worth solving with announce/watch-any should be
judged after living with the patched bulk script for a while, not
designed in advance of any real signal that the wait matters.
6. CI workflow addition
One step appended to the existing docker.yml, after the existing
build-and-push step, best-effort (must not fail the build if this
project hasn't been bootstrapped onto keep yet):
- name: announce to keep
continue-on-error: true
run: keep announce myapp/production --tag ${{ steps.meta.outputs.short_sha }}
continue-on-error: true is load-bearing, not incidental — a project
without a keep identity registered yet must still build and push its
image successfully. This step failing should be silent/expected for any
not-yet-bootstrapped project, not a red X on the workflow.
Security considerations
announceleaks a little more thanpush/pulldid: the tag itself is now server-visible (previously, if you were being strict, even the fact that a new image existed wasn't somethingkeepneeded to know). This is an intentional, accepted tradeoff — a git short-SHA or version tag isn't sensitive, and the whole feature only works ifkeepcan see it.- Requiring only a read grant (not write) to
announcemeans a compromised read-only CI identity could announce a false tag, pointing the fleet at an attacker-chosen image — worth flagging even thoughkeepnever pulls or runs that image itself, only relays the string. The actualdocker compose pullon each VPS still pulls from the configured registry using the tag it was told, so this is equivalent in severity to that CI identity's build step being compromised in the first place (it could already push a bad image under a legitimate tag) — not a new capability, just worth stating plainly rather than leaving implicit. - The fallback path must never silently downgrade a bootstrapped
project to running with a stale
.envbecause of a transientkeepoutage. As written,keep pullfailing for any reason (network blip,keepserver down, not just "not bootstrapped") falls through to "deploy with whatever.envis already on disk" — for an already-bootstrapped project, that's the previouskeep pull's output, still valid, not a downgrade. This only becomes a real problem if a bootstrapped project's very first deploy attempt coincides withkeepbeing unreachable, in which case there's no.envon disk yet at all anddocker compose upwill fail for the mundane reason that required env vars are missing — a loud, obvious failure, not a silent wrong-secrets one.
Open questions
- Fleet-wide
pull_policy: alwayssweep — a prerequisite this proposal depends on but doesn't itself fix (see Context). Should happen before or alongside this work, as its own small pass, not bundled into this one. - Does
watch/watch-anyever get used, or does the already-scheduled bulk sweep turn out to be good enough forever? Deliberately left unbuilt (see §5) until there's a real signal the wait between scheduled runs is worth solving. - Multi-VPS fan-out for one project — if a project ever runs on more
than one box,
announcealready supports that for free (any number of watchers can poll the same vault's announce endpoint), but this hasn't been exercised and might reveal something not accounted for here.