Files
keep/PROPOSAL-announce-and-agent.md
Fredrik Johansson 58d5d787dd
All checks were successful
Docker / build-and-push (push) Successful in 1m58s
Implement announce/watch: a deploy signal that reverses CI's reach
CI announces a new image tag to keep (any grant is enough, read-only
included); each deploy target polls for it locally instead of CI
holding standing SSH access to production. keep stays a pure relay —
the tag lives outside the encrypted vault payload in its own table,
and keep never executes anything itself.

Adds vault_announcements, POST/GET /vaults/:key/announce, and the
`keep announce`/`keep watch` CLI commands, per
PROPOSAL-announce-and-agent.md (now marked implemented). Verified live:
a read-only identity announcing successfully and being logged, watch
picking up the tag on its first poll, and watch exiting nonzero for an
unknown/ungranted vault.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-12 21:12:55 +02:00

19 KiB
Raw Blame History

Proposal: announce + a self-deploying agent, with a no-keep fallback

Status: implemented (§1§3, §6 shape). announce/watch, the vault_announcements table, and the API routes described below are built — see IMPLEMENTATION.md's "Announce/watch" section for the as-built version of the design rationale, and the CLI surface list there for the exact commands. §4 (patching the real bulk deploy script) and §6 (the CI workflow step) are per-project changes made outside this repo, on the host and in each project's own .gitea/workflows/docker.yml — not something this repo's code can contain. This file is kept as the original design record, superseded by IMPLEMENTATION.md as the source of truth for what's actually built.

Context (read this first)

Full background on keep is in README.md and IMPLEMENTATION.md — read both before touching code. The relevant existing pieces for this proposal:

  • src/server/db.ts — SQLite (better-sqlite3, WAL mode). vaults, recipients, vault_grants, vault_grants_previous, access_log.

  • src/server/routes.ts — recipient-authenticated (signed-request) routes: push, pull, grant, log, plus exists and recipients (both used by the CLI to compute who to re-wrap for).

  • src/server/auth.tsrequireRecipientAuth (Ed25519 signature over method+path+timestamp+body-hash, verified against req.originalUrl with the query string stripped — see IMPLEMENTATION.md's "What was verified end-to-end" for a bug that already bit this exact mechanism once) and requireAdminAuth (X-Admin-Password header).

  • src/cli/identity.ts, recipient.ts, vault.ts, signedFetch.ts (note: signedFetch's path param must never include a query string — see the comment there for why).

  • scripts/deploy.sh (this repo) — pulls <project>/<env> from keep into $DEPLOY_ROOT/<project>/.env, then docker compose up -d for one project. Assumes the vault already exists and this machine already has a grant on it. Superseded by the real deploy script below as the thing this proposal actually patches — kept in this repo as a single-project example/fallback, not the production path.

  • The real, already-running deploy script — on the host, not in any repo. A fixed-array bulk script, already in production use, looping over roughly a dozen project directories:

    ROOT="$(cd "$(dirname "$0")" && pwd)"
    PROJECTS=(proj1 proj2 proj3 proj4 proj5 proj6 proj7 proj8 proj9 proj10 proj11 proj12)
    for project in "${PROJECTS[@]}"; do
      dir="$ROOT/$project"
      [ -f "$dir/docker-compose.yml" ] || { echo "skip: $project"; continue; }
      (cd "$dir" && docker compose up -d) || FAILED+=("$project")
    done
    docker image prune -f
    

    Two things about it that matter for this proposal: it never runs docker compose pull — freshness comes entirely from each project's own pull_policy: always in its docker-compose.yml (up -d re-pulls implicitly). And it loops over a fixed, explicit array, not directory auto-discovery — adding a project means editing this array by hand. This is the actual thing to patch, not a new per-project script; see Design §4.

  • Prerequisite gap, found while writing this proposal, not yet fixed everywhere: pull_policy: always isn't universal across the fleet — confirmed present in roughly half the projects checked, confirmed missing in several others (one of which was fixed in this session — a sibling project's docker-compose.yml didn't have it, a regression from copying another project's compose file without carrying this line over), and unchecked for the rest (no local checkout available at proposal time). This is independent of keep entirely — the bulk script can't redeploy any project missing this line regardless of whether that project ever adopts keep, and it should be swept fleet-wide before or alongside this work, not as part of it.

  • .gitea/workflows/docker.yml (present in keep and other sibling projects) — builds and pushes an image to the Gitea container registry on every push to main. Does not currently deploy anything — that's the manual/scheduled bulk script above today.

Motivation

The CI workflow already produces a fresh image on every push. Deployment is still a manual "SSH in and run the deploy script" step. The obvious fix — CI SSHes into the VPS and runs the deploy — was considered and rejected: it means a standing SSH credential with production-deploy rights sits in Gitea's CI secrets, reachable by whatever's executing on the shared runner across every project it builds. A compromised build step in any one project could then reach every other project's deploy target. That's a real regression from where keep already left things (see IMPLEMENTATION.md's resolved "read no longer implies write" — the whole point was not concentrating reach into one identity).

The fix that doesn't reintroduce that problem: reverse the direction. CI never reaches into a VPS. Instead, CI tells keep "there's a new image for this project," and each VPS — which already has to run something locally to redeploy itself — polls keep for that signal and acts on it locally, using an identity and grant it already legitimately holds. keep's server never executes anything remotely; it stays a pure relay, same posture as everything else it does.

Not every project will have been migrated to keep yet. Bootstrapping a project (register a recipient, grant it, push the first vault) is a real, deliberate step per project, and there'll be a period — maybe a long one — where some projects have it and some don't. The redeploy path has to work identically well for both, defaulting to exactly what happens today (a plain docker compose up -d, relying on that project's own pull_policy: always) for anything not yet bootstrapped, rather than either blocking deploys entirely or hard-failing on missing vaults.

Considered and rejected: keep executes the redeploy itself

Since keep's own server will run on the same VPS as the deploy targets, it's tempting to skip the polling indirection entirely and have keep shell out to the deploy script directly when it receives an announce — no new SSH credential, no new network-reachable pathway, since nothing is reaching into the box that wasn't already running on it.

Rejected anyway, because it trades away a property that's been the throughline of keep's whole design: today, a fully compromised keep server — not just a compromised client identity — can't do anything worse than leak ciphertext and metadata, because the server never holds a decryption key. The moment keep can also execute a command on announce, that stops being true — a bug in the HTTP layer (not even a cryptographic break; an ordinary web app bug, an auth check wrong on some edge case) becomes "attacker triggers arbitrary redeploys," not "attacker sees encrypted blobs they can't read." That's a new class of bad outcome, not a bigger version of an existing one, and it's not worth trading for skipping a polling loop.

There's also a concrete operational cost, not just a principle: actually running docker compose up -d from inside keep's own process would need either the Docker socket mounted into keep's container (/var/run/docker.sock — widely treated as equivalent to handing that container root on the host) or running keep outside a container entirely, throwing away the isolation its own Dockerfile already provides today.

keep stays a pure relay regardless of where it's deployed. Something outside keep's own process boundary does the polling and the executing — see §4.

Design

1. announce — a distinct verb from push/pull

An image tag isn't a secret. It doesn't belong inside the encrypted vault payload alongside real credentials — that's a category mismatch (deploy metadata living inside a secrets blob), not just an implementation convenience either way. announce is its own small mechanism:

CREATE TABLE IF NOT EXISTS vault_announcements (
  vault_key     TEXT PRIMARY KEY,
  tag           TEXT NOT NULL,
  announced_at  INTEGER NOT NULL,
  announced_by  TEXT NOT NULL   -- recipient_id
);

One row per vault — like the vault's own one-previous-version retention, this is "latest known state," not a history. If a real audit trail of every announce is wanted later, access_log already records the action; this table is just the fast-path "what's the current tag" read, kept separate from the log on purpose (reading current state back out of an append-only log is the wrong tool for that job — same reasoning IMPLEMENTATION.md already gives for why vault_grants_previous is a real table and not a log replay).

2. API

POST /api/vaults/:key/announce   -- { tag: string } — requires ANY grant on the vault (read-only is enough; announcing isn't a secrets mutation)
GET  /api/vaults/:key/announce   -- { tag, announcedAt, announcedBy } | 404 if nothing announced yet — requires a grant, same as pull

Both go through requireRecipientAuth, identical shape to the existing routes. POST only requires a grant (not write) — reuse hasGrant, not hasWriteGrant — since this isn't a mutation of the vault's secret content, just a signal, and CI's identity for a project should be able to stay read-only even after this feature exists. (This is worth stating explicitly since a naive implementation might reach for hasWriteGrant by analogy with push — resist that; it would silently widen CI's required permissions for no real reason.)

3. CLI

keep announce <vault> --tag <tag>     -- POST, requires a grant on <vault>
keep watch <vault> [--interval 60]    -- polls GET .../announce every N seconds, prints the tag on change, exits nonzero if the vault doesn't exist / no grant (see fallback below)

keep watch is deliberately dumb — plain polling, not a websocket or long-poll. Matches the polling pattern already used elsewhere in this project family, and a 60-second detection lag for a deploy is fine. Don't build push-based notification for this; it's more machinery for a problem polling already solves adequately.

4. Patch the existing loop — don't build a parallel script

This is the part the fallback requirement is actually about. The real deploy script already does the one thing that matters (a loop over a fixed project list, docker compose up -d per project, relying on pull_policy: always for freshness) — the change is inserting one keep-with-fallback step per iteration, in place, not writing a new per-project agent.sh that duplicates the loop/array/failure-tracking logic the real script already has:

for project in "${PROJECTS[@]}"; do
  dir="$ROOT/$project"
  [ -f "$dir/docker-compose.yml" ] || { echo "skip: $project"; continue; }

  echo ""
  echo "▶ $project"

  # A vault key is NOT assumed to match the deploy-directory name — repo
  # names, deploy directory names, and vault keys are three independent
  # strings in practice (a repo can be named one thing, its deploy
  # directory another, entirely by history/accident). .keep-vault is an
  # explicit, one-time-per-project marker: if it's missing, this project
  # was never bootstrapped onto keep (or never will be — a static site
  # with no secrets), and the loop falls through to exactly what it did
  # before keep existed. Every keep call here is allowed to fail;
  # failure means "keep doesn't apply to this project," never "stop the
  # deploy."
  if [ -f "$dir/.keep-vault" ]; then
    vault="$(cat "$dir/.keep-vault")"
    if keep pull "$vault" --out "$dir/.env.new" 2>/dev/null; then
      mv "$dir/.env.new" "$dir/.env"
    else
      rm -f "$dir/.env.new"
    fi
  fi

  if (cd "$dir" && docker compose up -d 2>&1); then
    echo "  ✓ $project up"
  else
    echo "  ✗ $project failed"
    FAILED+=("$project")
  fi
done

Bootstrapping a project onto keep becomes: choose a vault key (no naming constraint — doesn't need to match the repo or the directory), keep push it, grant whichever recipients need it, and echo "<the vault key>" > "$dir/.keep-vault". One file, human-readable, cat-able to check what's configured, no second array in this script to keep in sync by hand as directories get renamed or reorganized.

Notice announce/watch aren't actually load-bearing in this version — whatever already triggers this script (cron, systemd timer, or still run by hand) keeps triggering it exactly as before, and keep pull either updates .env or silently doesn't. That's deliberate: the existing script, patched this way, is already a complete, correct baseline — keep-bootstrapped projects get fresh secrets on every run, everything else behaves identically to today. announce sits on top as a latency optimization for however this script gets triggered, not a requirement for this patch to be complete — see the next section.

5. Where announce/watch actually earn their keep

Two independent benefits, worth keeping conceptually separate — and the second one looks different than it would for a one-agent-per-project design, because the real script is one bulk all-or-nothing sweep, not N independent watchers:

  • Secrets sourcingkeep pull instead of a hand-placed .env. Available per-project the moment its vault exists and this machine has a grant on it. This is what §4's patch already delivers, on its own, with zero dependency on announce.
  • Waking the sweep early — however the bulk script currently gets triggered (cron, systemd timer, or still by hand), announce could let it wake immediately on a push instead of waiting for the next scheduled run. The natural fit given the script's actual shape: keep watch-any <vault1> <vault2> ... blocking until any watched vault's tag changes, used as the trigger for a full sweep — not a per-project selective redeploy. Selective ("only redeploy the one project that actually changed") would require reshaping the script away from its current fixed-array-sweep model into something that takes a target project as an argument, which is a real behavior change to something already in production use, not a small addition — out of scope here.

Deliberately not building the wake-the-sweep-early mechanism in this pass. Ship §4's patch first — it already satisfies the fallback requirement completely and delivers the secrets-sourcing benefit on its own. Whether waiting up to one scheduled interval for a redeploy is actually a problem worth solving with announce/watch-any should be judged after living with the patched bulk script for a while, not designed in advance of any real signal that the wait matters.

6. CI workflow addition

One step appended to the existing docker.yml, after the existing build-and-push step, best-effort (must not fail the build if this project hasn't been bootstrapped onto keep yet):

- name: announce to keep
  continue-on-error: true
  run: keep announce myapp/production --tag ${{ steps.meta.outputs.short_sha }}

The vault key here (myapp/production) is a literal, hardcoded string in this specific repo's own workflow file — never derived from ${{ github.repository }} or any other automatic source. Same reasoning as §4's .keep-vault file: the repo's own name, its deploy directory's name, and its vault key are three independent strings that happen to often look similar and are not guaranteed to match (see Context — a real example already exists in this fleet). Whoever bootstraps a project onto keep should paste the same deliberately-chosen vault key into both this workflow step and that project's .keep-vault file, not rely on either side inferring it from something else.

continue-on-error: true is load-bearing, not incidental — a project without a keep identity registered yet must still build and push its image successfully. This step failing should be silent/expected for any not-yet-bootstrapped project, not a red X on the workflow.

Security considerations

  • announce leaks a little more than push/pull did: the tag itself is now server-visible (previously, if you were being strict, even the fact that a new image existed wasn't something keep needed to know). This is an intentional, accepted tradeoff — a git short-SHA or version tag isn't sensitive, and the whole feature only works if keep can see it.
  • Requiring only a read grant (not write) to announce means a compromised read-only CI identity could announce a false tag, pointing the fleet at an attacker-chosen image — worth flagging even though keep never pulls or runs that image itself, only relays the string. The actual docker compose pull on each VPS still pulls from the configured registry using the tag it was told, so this is equivalent in severity to that CI identity's build step being compromised in the first place (it could already push a bad image under a legitimate tag) — not a new capability, just worth stating plainly rather than leaving implicit.
  • The fallback path must never silently downgrade a bootstrapped project to running with a stale .env because of a transient keep outage. As written, keep pull failing for any reason (network blip, keep server down, not just "not bootstrapped") falls through to "deploy with whatever .env is already on disk" — for an already-bootstrapped project, that's the previous keep pull's output, still valid, not a downgrade. This only becomes a real problem if a bootstrapped project's very first deploy attempt coincides with keep being unreachable, in which case there's no .env on disk yet at all and docker compose up will fail for the mundane reason that required env vars are missing — a loud, obvious failure, not a silent wrong-secrets one.

Open questions

  • .keep-vault format — a bare vault-key string is enough for today's one-environment-per-directory reality. If a directory ever needs to deploy more than one environment from the same docker-compose.yml (unlikely given the current fleet shape, but not impossible), a bare string stops being enough and it'd need to become a small key=value file or similar. Not designing that ahead of an actual need — ship the bare-string version.
  • Fleet-wide pull_policy: always sweep — a prerequisite this proposal depends on but doesn't itself fix (see Context). Should happen before or alongside this work, as its own small pass, not bundled into this one.
  • Does watch/watch-any ever get used, or does the already-scheduled bulk sweep turn out to be good enough forever? Deliberately left unbuilt (see §5) until there's a real signal the wait between scheduled runs is worth solving.
  • Multi-VPS fan-out for one project — if a project ever runs on more than one box, announce already supports that for free (any number of watchers can poll the same vault's announce endpoint), but this hasn't been exercised and might reveal something not accounted for here.