Add proposal: announce + self-deploying agent, with a no-keep fallback

Design doc for closing the CI->deploy gap without reintroducing the
credential-concentration problem a naive "CI SSHes into the VPS"
approach would bring back. keep stays a pure relay — never executes
anything remotely. CI announces a new tag (a distinct verb from push/
pull, since a tag isn't a secret and doesn't belong in the encrypted
vault payload); each VPS polls and redeploys itself locally, using an
identity and grant it already holds.

Explicitly specs the not-yet-bootstrapped fallback: the agent script's
`keep pull` is allowed to fail for any reason, and on failure just runs
a plain `docker compose pull && up -d` against whatever's already on
disk — the same thing that would have happened before keep existed.
Ships the dumb-periodic-timer version of the agent first; the watch/
announce fast path is speced but deliberately left unbuilt until
there's a real signal that polling latency is worth the complexity.

Not implemented — design stage.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Fredrik Johansson
2026-07-12 20:43:51 +02:00
parent 5f31eb76c7
commit e505fbf9b4

View File

@@ -0,0 +1,264 @@
# Proposal: announce + a self-deploying agent, with a no-`keep` fallback
**Status:** proposed, not built. Self-contained — written so a fresh agent
with no prior conversation context can pick this up and implement it
without anything explained first.
## Context (read this first)
Full background on `keep` is in [README.md](./README.md) and
[IMPLEMENTATION.md](./IMPLEMENTATION.md) — read both before touching
code. The relevant existing pieces for this proposal:
- `src/server/db.ts` — SQLite (better-sqlite3, WAL mode). `vaults`,
`recipients`, `vault_grants`, `vault_grants_previous`, `access_log`.
- `src/server/routes.ts` — recipient-authenticated (signed-request)
routes: `push`, `pull`, `grant`, `log`, plus `exists` and
`recipients` (both used by the CLI to compute who to re-wrap for).
- `src/server/auth.ts``requireRecipientAuth` (Ed25519 signature over
method+path+timestamp+body-hash, verified against `req.originalUrl`
with the query string stripped — see IMPLEMENTATION.md's "What was
verified end-to-end" for a bug that already bit this exact mechanism
once) and `requireAdminAuth` (`X-Admin-Password` header).
- `src/cli/``identity.ts`, `recipient.ts`, `vault.ts`,
`signedFetch.ts` (note: `signedFetch`'s `path` param must never
include a query string — see the comment there for why).
- `scripts/deploy.sh` — pulls `<project>/<env>` from `keep` into
`$DEPLOY_ROOT/<project>/.env`, then `docker compose up -d`. Assumes
the vault already exists and this machine already has a grant on it.
- `.gitea/workflows/docker.yml` (present in `keep`, `wisp`,
`npm-statuspage`) — builds and pushes an image to the Gitea container
registry on every push to `main`. Does **not** currently deploy
anything — that's a manual step today.
## Motivation
The CI workflow already produces a fresh image on every push. Deployment
is still a manual "SSH in and run the deploy script" step. The obvious
fix — CI SSHes into the VPS and runs the deploy — was considered and
rejected: it means a standing SSH credential with production-deploy
rights sits in Gitea's CI secrets, reachable by whatever's executing on
the shared runner across all ~17 projects it builds. A compromised build
step in any one project could then reach every other project's deploy
target. That's a real regression from where `keep` already left things
(see IMPLEMENTATION.md's resolved "read no longer implies write" — the
whole point was *not* concentrating reach into one identity).
The fix that doesn't reintroduce that problem: **reverse the direction.**
CI never reaches into a VPS. Instead, CI tells `keep` "there's a new
image for this project," and each VPS — which already has to run
*something* locally to redeploy itself — polls `keep` for that signal
and acts on it locally, using an identity and grant it already legitimately
holds. `keep`'s server never executes anything remotely; it stays a pure
relay, same posture as everything else it does.
**Not every project will have been migrated to `keep` yet.** Bootstrapping
a project (register a recipient, grant it, push the first vault) is a
real, deliberate step per project, and there'll be a period — maybe a
long one — where some of the ~17 projects have it and some don't. The
self-deploying agent has to work identically well for both, defaulting to
exactly what happens today (a plain `docker compose pull && up -d`
against whatever's already on disk) for anything not yet bootstrapped,
rather than either blocking deploys entirely or hard-failing on missing
vaults.
## Design
### 1. `announce` — a distinct verb from `push`/`pull`
An image tag isn't a secret. It doesn't belong inside the encrypted vault
payload alongside real credentials — that's a category mismatch (deploy
metadata living inside a secrets blob), not just an implementation
convenience either way. `announce` is its own small mechanism:
```sql
CREATE TABLE IF NOT EXISTS vault_announcements (
vault_key TEXT PRIMARY KEY,
tag TEXT NOT NULL,
announced_at INTEGER NOT NULL,
announced_by TEXT NOT NULL -- recipient_id
);
```
One row per vault — like the vault's own one-previous-version retention,
this is "latest known state," not a history. If a real audit trail of
every announce is wanted later, `access_log` already records the action;
this table is just the fast-path "what's the current tag" read, kept
separate from the log on purpose (reading current state back out of an
append-only log is the wrong tool for that job — same reasoning
`IMPLEMENTATION.md` already gives for why `vault_grants_previous` is a
real table and not a log replay).
### 2. API
```
POST /api/vaults/:key/announce -- { tag: string } — requires ANY grant on the vault (read-only is enough; announcing isn't a secrets mutation)
GET /api/vaults/:key/announce -- { tag, announcedAt, announcedBy } | 404 if nothing announced yet — requires a grant, same as pull
```
Both go through `requireRecipientAuth`, identical shape to the existing
routes. `POST` only requires *a* grant (not write) — reuse `hasGrant`,
not `hasWriteGrant` — since this isn't a mutation of the vault's secret
content, just a signal, and CI's identity for a project should be able to
stay read-only even after this feature exists. (This is worth stating
explicitly since a naive implementation might reach for `hasWriteGrant`
by analogy with `push` — resist that; it would silently widen CI's
required permissions for no real reason.)
### 3. CLI
```
keep announce <vault> --tag <tag> -- POST, requires a grant on <vault>
keep watch <vault> [--interval 60] -- polls GET .../announce every N seconds, prints the tag on change, exits nonzero if the vault doesn't exist / no grant (see fallback below)
```
`keep watch` is deliberately dumb — plain polling, not a websocket or
long-poll. Matches `npm-statuspage`'s existing polling pattern, and a
60-second detection lag for a deploy is fine. Don't build push-based
notification for this; it's more machinery for a problem polling already
solves adequately.
### 4. The agent script and its fallback
This is the part the fallback requirement is actually about. One script
per VPS, running on a timer (systemd timer or cron, not a long-lived
daemon — matches this project's bias toward simple, restartable,
stateless-between-runs tooling), one invocation per project it's
responsible for:
```bash
#!/usr/bin/env bash
# agent.sh <project> [env] — self-deploy check for one project. Meant to
# run on a schedule (cron/systemd timer), not as a daemon.
set -euo pipefail
project="${1:?usage: agent.sh <project> [env]}"
env="${2:-production}"
root="${DEPLOY_ROOT:-/opt/docker}"
dir="$root/$project"
vault="$project/$env"
if [ ! -d "$dir" ]; then
echo "error: $dir does not exist" >&2
exit 1
fi
cd "$dir"
# Best-effort: does this machine have a working keep grant on this
# vault at all? If not — project not yet bootstrapped onto keep, or
# this box was never granted access — fall through to the plain path.
# Every keep call below is allowed to fail; failure just means "keep
# isn't available for this project," not "the deploy is broken."
if keep pull "$vault" --out .env.new 2>/dev/null; then
mv .env.new .env
docker compose pull -q
docker compose up -d
else
rm -f .env.new
# No keep vault (or no grant) for this project — deploy exactly as it
# would have worked before keep existed: whatever .env is already
# sitting here (possibly none, if the compose file doesn't need one),
# plain pull + up. This path must never hard-fail just because keep
# doesn't know about this project.
docker compose pull -q
docker compose up -d
fi
```
Notice `announce`/`watch` aren't actually load-bearing in this version —
the agent above just runs on a timer and lets `docker compose pull`
figure out for itself whether there's anything new (a no-op if the
digest hasn't changed, idempotent either way). That's deliberate: **a
dumb periodic `compose pull && up -d` is already a complete, correct
baseline deploy strategy on its own**, keep-bootstrapped or not. `announce`
sits on top as an optimization for bootstrapped projects, not a
requirement for the agent to function — see the next section.
### 5. Where `announce`/`watch` actually earn their keep
For a project that *has* been bootstrapped, there are two independent
benefits, worth keeping conceptually separate:
- **Secrets sourcing** — `keep pull` instead of a hand-placed `.env`.
Available the moment a vault exists and this machine has a grant.
Already true of `deploy.sh` today; unchanged by this proposal.
- **Fast, precise redeploy** — instead of every VPS blindly polling
`docker compose pull` on a fixed timer for every project it hosts
(harmless, but imprecise and up to a full interval late), a
bootstrapped project's agent can `keep watch` and react within
seconds of the CI `announce` step, and only for the project that
actually changed rather than sweeping all of them.
A fancier version of the agent script could branch on `keep watch`
succeeding vs. not to decide between "wait for a watch event" and "just
poll on the dumb timer" — deliberately **not** speccing that split here.
Ship the dumb-timer-plus-opportunistic-`keep-pull` version first (above);
it already satisfies the fallback requirement completely, and whether
the watch-driven fast path is worth the added complexity should be
judged after living with the dumb version for a while, not designed in
advance of any real signal that polling latency is actually a problem.
### 6. CI workflow addition
One step appended to the existing `docker.yml`, after the existing
build-and-push step, best-effort (must not fail the build if this
project hasn't been bootstrapped onto `keep` yet):
```yaml
- name: announce to keep
continue-on-error: true
run: keep announce myapp/production --tag ${{ steps.meta.outputs.short_sha }}
```
`continue-on-error: true` is load-bearing, not incidental — a project
without a `keep` identity registered yet must still build and push its
image successfully. This step failing should be silent/expected for any
not-yet-bootstrapped project, not a red X on the workflow.
## Security considerations
- `announce` leaks a little more than `push`/`pull` did: the *tag*
itself is now server-visible (previously, if you were being strict,
even the fact that a new image existed wasn't something `keep` needed
to know). This is an intentional, accepted tradeoff — a git short-SHA
or version tag isn't sensitive, and the whole feature only works if
`keep` can see it.
- Requiring only a read grant (not write) to `announce` means a
compromised read-only CI identity could announce a *false* tag,
pointing the fleet at an attacker-chosen image — worth flagging even
though `keep` never pulls or runs that image itself, only relays the
string. The actual `docker compose pull` on each VPS still pulls from
the configured registry using the tag it was told, so this is
equivalent in severity to that CI identity's build step being
compromised in the first place (it could already push a bad image
under a legitimate tag) — not a new capability, just worth stating
plainly rather than leaving implicit.
- The fallback path must never silently downgrade a *bootstrapped*
project to running with a stale `.env` because of a transient `keep`
outage. As written, `keep pull` failing for any reason (network
blip, `keep` server down, not just "not bootstrapped") falls through
to "deploy with whatever `.env` is already on disk" — for an
already-bootstrapped project, that's the *previous* `keep pull`'s
output, still valid, not a downgrade. This only becomes a real problem
if a bootstrapped project's very first deploy attempt coincides with
`keep` being unreachable, in which case there's no `.env` on disk yet
at all and `docker compose up` will fail for the mundane reason that
required env vars are missing — a loud, obvious failure, not a silent
wrong-secrets one.
## Open questions
- **Per-project agent scheduling** — one cron line per project, or one
script that loops over every directory under `$DEPLOY_ROOT`? Leaning
toward one-line-per-project (explicit, greppable in crontab, no
surprise when a new project directory silently starts getting picked
up) over auto-discovery.
- **Does `watch` ever get used**, or does the dumb-timer version turn
out to be good enough forever? Deliberately left unbuilt (see "Where
announce/watch actually earn their keep") until there's a real signal
it's worth the added complexity.
- **Multi-VPS fan-out for one project** — if a project ever runs on more
than one box, `announce` already supports that for free (any number of
agents can poll the same vault's announce endpoint), but this hasn't
been exercised and might reveal something not accounted for here.