All checks were successful
Docker / build-and-push (push) Successful in 1m58s
CI announces a new image tag to keep (any grant is enough, read-only included); each deploy target polls for it locally instead of CI holding standing SSH access to production. keep stays a pure relay — the tag lives outside the encrypted vault payload in its own table, and keep never executes anything itself. Adds vault_announcements, POST/GET /vaults/:key/announce, and the `keep announce`/`keep watch` CLI commands, per PROPOSAL-announce-and-agent.md (now marked implemented). Verified live: a read-only identity announcing successfully and being logged, watch picking up the tag on its first poll, and watch exiting nonzero for an unknown/ungranted vault. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
370 lines
19 KiB
Markdown
370 lines
19 KiB
Markdown
# Proposal: announce + a self-deploying agent, with a no-`keep` fallback
|
||
|
||
**Status:** implemented (§1–§3, §6 shape). `announce`/`watch`, the
|
||
`vault_announcements` table, and the API routes described below are
|
||
built — see [IMPLEMENTATION.md](./IMPLEMENTATION.md)'s "Announce/watch"
|
||
section for the as-built version of the design rationale, and the CLI
|
||
surface list there for the exact commands. §4 (patching the real bulk
|
||
deploy script) and §6 (the CI workflow step) are per-project changes
|
||
made outside this repo, on the host and in each project's own
|
||
`.gitea/workflows/docker.yml` — not something this repo's code can
|
||
contain. This file is kept as the original design record, superseded by
|
||
IMPLEMENTATION.md as the source of truth for what's actually built.
|
||
|
||
## Context (read this first)
|
||
|
||
Full background on `keep` is in [README.md](./README.md) and
|
||
[IMPLEMENTATION.md](./IMPLEMENTATION.md) — read both before touching
|
||
code. The relevant existing pieces for this proposal:
|
||
|
||
- `src/server/db.ts` — SQLite (better-sqlite3, WAL mode). `vaults`,
|
||
`recipients`, `vault_grants`, `vault_grants_previous`, `access_log`.
|
||
- `src/server/routes.ts` — recipient-authenticated (signed-request)
|
||
routes: `push`, `pull`, `grant`, `log`, plus `exists` and
|
||
`recipients` (both used by the CLI to compute who to re-wrap for).
|
||
- `src/server/auth.ts` — `requireRecipientAuth` (Ed25519 signature over
|
||
method+path+timestamp+body-hash, verified against `req.originalUrl`
|
||
with the query string stripped — see IMPLEMENTATION.md's "What was
|
||
verified end-to-end" for a bug that already bit this exact mechanism
|
||
once) and `requireAdminAuth` (`X-Admin-Password` header).
|
||
- `src/cli/` — `identity.ts`, `recipient.ts`, `vault.ts`,
|
||
`signedFetch.ts` (note: `signedFetch`'s `path` param must never
|
||
include a query string — see the comment there for why).
|
||
- `scripts/deploy.sh` (this repo) — pulls `<project>/<env>` from `keep`
|
||
into `$DEPLOY_ROOT/<project>/.env`, then `docker compose up -d` for
|
||
one project. Assumes the vault already exists and this machine
|
||
already has a grant on it. **Superseded by the real deploy script
|
||
below as the thing this proposal actually patches** — kept in this
|
||
repo as a single-project example/fallback, not the production path.
|
||
- **The real, already-running deploy script — on the host, not in any
|
||
repo.** A fixed-array bulk script, already in production use, looping
|
||
over roughly a dozen project directories:
|
||
|
||
```bash
|
||
ROOT="$(cd "$(dirname "$0")" && pwd)"
|
||
PROJECTS=(proj1 proj2 proj3 proj4 proj5 proj6 proj7 proj8 proj9 proj10 proj11 proj12)
|
||
for project in "${PROJECTS[@]}"; do
|
||
dir="$ROOT/$project"
|
||
[ -f "$dir/docker-compose.yml" ] || { echo "skip: $project"; continue; }
|
||
(cd "$dir" && docker compose up -d) || FAILED+=("$project")
|
||
done
|
||
docker image prune -f
|
||
```
|
||
|
||
Two things about it that matter for this proposal: **it never runs
|
||
`docker compose pull`** — freshness comes entirely from each
|
||
project's own `pull_policy: always` in its `docker-compose.yml`
|
||
(`up -d` re-pulls implicitly). And **it loops over a fixed, explicit
|
||
array**, not directory auto-discovery — adding a project means
|
||
editing this array by hand. This is the actual thing to patch, not a
|
||
new per-project script; see Design §4.
|
||
- **Prerequisite gap, found while writing this proposal, not yet fixed
|
||
everywhere**: `pull_policy: always` isn't universal across the
|
||
fleet — confirmed present in roughly half the projects checked,
|
||
confirmed missing in several others (one of which was fixed in this
|
||
session — a sibling project's `docker-compose.yml` didn't have it, a
|
||
regression from copying another project's compose file without
|
||
carrying this line over), and unchecked for the rest (no local
|
||
checkout available at proposal time). **This is independent of
|
||
`keep` entirely** — the bulk script can't redeploy any project
|
||
missing this line regardless of whether that project ever adopts
|
||
`keep`, and it should be swept fleet-wide before or alongside this
|
||
work, not as part of it.
|
||
- `.gitea/workflows/docker.yml` (present in `keep` and other sibling
|
||
projects) — builds and pushes an image to the Gitea container
|
||
registry on every push to `main`. Does **not** currently deploy
|
||
anything — that's the manual/scheduled bulk script above today.
|
||
|
||
## Motivation
|
||
|
||
The CI workflow already produces a fresh image on every push. Deployment
|
||
is still a manual "SSH in and run the deploy script" step. The obvious
|
||
fix — CI SSHes into the VPS and runs the deploy — was considered and
|
||
rejected: it means a standing SSH credential with production-deploy
|
||
rights sits in Gitea's CI secrets, reachable by whatever's executing on
|
||
the shared runner across every project it builds. A compromised build
|
||
step in any one project could then reach every other project's deploy
|
||
target. That's a real regression from where `keep` already left things
|
||
(see IMPLEMENTATION.md's resolved "read no longer implies write" — the
|
||
whole point was *not* concentrating reach into one identity).
|
||
|
||
The fix that doesn't reintroduce that problem: **reverse the direction.**
|
||
CI never reaches into a VPS. Instead, CI tells `keep` "there's a new
|
||
image for this project," and each VPS — which already has to run
|
||
*something* locally to redeploy itself — polls `keep` for that signal
|
||
and acts on it locally, using an identity and grant it already legitimately
|
||
holds. `keep`'s server never executes anything remotely; it stays a pure
|
||
relay, same posture as everything else it does.
|
||
|
||
**Not every project will have been migrated to `keep` yet.** Bootstrapping
|
||
a project (register a recipient, grant it, push the first vault) is a
|
||
real, deliberate step per project, and there'll be a period — maybe a
|
||
long one — where some projects have it and some don't. The redeploy path
|
||
has to work identically well for both, defaulting to exactly what
|
||
happens today (a plain `docker compose up -d`, relying on that project's
|
||
own `pull_policy: always`) for anything not yet bootstrapped, rather than
|
||
either blocking deploys entirely or hard-failing on missing vaults.
|
||
|
||
### Considered and rejected: `keep` executes the redeploy itself
|
||
|
||
Since `keep`'s own server will run on the same VPS as the deploy targets,
|
||
it's tempting to skip the polling indirection entirely and have `keep`
|
||
shell out to the deploy script directly when it receives an `announce` —
|
||
no new SSH credential, no new network-reachable pathway, since nothing
|
||
is reaching *into* the box that wasn't already running on it.
|
||
|
||
Rejected anyway, because it trades away a property that's been the
|
||
throughline of `keep`'s whole design: **today, a fully compromised
|
||
`keep` server — not just a compromised client identity — can't do
|
||
anything worse than leak ciphertext and metadata**, because the server
|
||
never holds a decryption key. The moment `keep` can also execute a
|
||
command on `announce`, that stops being true — a bug in the HTTP layer
|
||
(not even a cryptographic break; an ordinary web app bug, an auth check
|
||
wrong on some edge case) becomes "attacker triggers arbitrary
|
||
redeploys," not "attacker sees encrypted blobs they can't read." That's
|
||
a new class of bad outcome, not a bigger version of an existing one, and
|
||
it's not worth trading for skipping a polling loop.
|
||
|
||
There's also a concrete operational cost, not just a principle: actually
|
||
running `docker compose up -d` from inside `keep`'s own process would
|
||
need either the Docker socket mounted into `keep`'s container
|
||
(`/var/run/docker.sock` — widely treated as equivalent to handing that
|
||
container root on the host) or running `keep` outside a container
|
||
entirely, throwing away the isolation its own `Dockerfile` already
|
||
provides today.
|
||
|
||
`keep` stays a pure relay regardless of where it's deployed. Something
|
||
*outside* `keep`'s own process boundary does the polling and the
|
||
executing — see §4.
|
||
|
||
## Design
|
||
|
||
### 1. `announce` — a distinct verb from `push`/`pull`
|
||
|
||
An image tag isn't a secret. It doesn't belong inside the encrypted vault
|
||
payload alongside real credentials — that's a category mismatch (deploy
|
||
metadata living inside a secrets blob), not just an implementation
|
||
convenience either way. `announce` is its own small mechanism:
|
||
|
||
```sql
|
||
CREATE TABLE IF NOT EXISTS vault_announcements (
|
||
vault_key TEXT PRIMARY KEY,
|
||
tag TEXT NOT NULL,
|
||
announced_at INTEGER NOT NULL,
|
||
announced_by TEXT NOT NULL -- recipient_id
|
||
);
|
||
```
|
||
|
||
One row per vault — like the vault's own one-previous-version retention,
|
||
this is "latest known state," not a history. If a real audit trail of
|
||
every announce is wanted later, `access_log` already records the action;
|
||
this table is just the fast-path "what's the current tag" read, kept
|
||
separate from the log on purpose (reading current state back out of an
|
||
append-only log is the wrong tool for that job — same reasoning
|
||
`IMPLEMENTATION.md` already gives for why `vault_grants_previous` is a
|
||
real table and not a log replay).
|
||
|
||
### 2. API
|
||
|
||
```
|
||
POST /api/vaults/:key/announce -- { tag: string } — requires ANY grant on the vault (read-only is enough; announcing isn't a secrets mutation)
|
||
GET /api/vaults/:key/announce -- { tag, announcedAt, announcedBy } | 404 if nothing announced yet — requires a grant, same as pull
|
||
```
|
||
|
||
Both go through `requireRecipientAuth`, identical shape to the existing
|
||
routes. `POST` only requires *a* grant (not write) — reuse `hasGrant`,
|
||
not `hasWriteGrant` — since this isn't a mutation of the vault's secret
|
||
content, just a signal, and CI's identity for a project should be able to
|
||
stay read-only even after this feature exists. (This is worth stating
|
||
explicitly since a naive implementation might reach for `hasWriteGrant`
|
||
by analogy with `push` — resist that; it would silently widen CI's
|
||
required permissions for no real reason.)
|
||
|
||
### 3. CLI
|
||
|
||
```
|
||
keep announce <vault> --tag <tag> -- POST, requires a grant on <vault>
|
||
keep watch <vault> [--interval 60] -- polls GET .../announce every N seconds, prints the tag on change, exits nonzero if the vault doesn't exist / no grant (see fallback below)
|
||
```
|
||
|
||
`keep watch` is deliberately dumb — plain polling, not a websocket or
|
||
long-poll. Matches the polling pattern already used elsewhere in this
|
||
project family, and a 60-second detection lag for a deploy is fine.
|
||
Don't build push-based notification for this; it's more machinery for a
|
||
problem polling already solves adequately.
|
||
|
||
### 4. Patch the existing loop — don't build a parallel script
|
||
|
||
This is the part the fallback requirement is actually about. The real
|
||
deploy script already does the one thing that matters (a loop over a
|
||
fixed project list, `docker compose up -d` per project, relying on
|
||
`pull_policy: always` for freshness) — the change is inserting one
|
||
`keep`-with-fallback step per iteration, in place, not writing a new
|
||
per-project `agent.sh` that duplicates the loop/array/failure-tracking
|
||
logic the real script already has:
|
||
|
||
```bash
|
||
for project in "${PROJECTS[@]}"; do
|
||
dir="$ROOT/$project"
|
||
[ -f "$dir/docker-compose.yml" ] || { echo "skip: $project"; continue; }
|
||
|
||
echo ""
|
||
echo "▶ $project"
|
||
|
||
# A vault key is NOT assumed to match the deploy-directory name — repo
|
||
# names, deploy directory names, and vault keys are three independent
|
||
# strings in practice (a repo can be named one thing, its deploy
|
||
# directory another, entirely by history/accident). .keep-vault is an
|
||
# explicit, one-time-per-project marker: if it's missing, this project
|
||
# was never bootstrapped onto keep (or never will be — a static site
|
||
# with no secrets), and the loop falls through to exactly what it did
|
||
# before keep existed. Every keep call here is allowed to fail;
|
||
# failure means "keep doesn't apply to this project," never "stop the
|
||
# deploy."
|
||
if [ -f "$dir/.keep-vault" ]; then
|
||
vault="$(cat "$dir/.keep-vault")"
|
||
if keep pull "$vault" --out "$dir/.env.new" 2>/dev/null; then
|
||
mv "$dir/.env.new" "$dir/.env"
|
||
else
|
||
rm -f "$dir/.env.new"
|
||
fi
|
||
fi
|
||
|
||
if (cd "$dir" && docker compose up -d 2>&1); then
|
||
echo " ✓ $project up"
|
||
else
|
||
echo " ✗ $project failed"
|
||
FAILED+=("$project")
|
||
fi
|
||
done
|
||
```
|
||
|
||
Bootstrapping a project onto `keep` becomes: choose a vault key (no
|
||
naming constraint — doesn't need to match the repo or the directory),
|
||
`keep push` it, grant whichever recipients need it, and
|
||
`echo "<the vault key>" > "$dir/.keep-vault"`. One file, human-readable,
|
||
`cat`-able to check what's configured, no second array in this script to
|
||
keep in sync by hand as directories get renamed or reorganized.
|
||
|
||
Notice `announce`/`watch` aren't actually load-bearing in this version —
|
||
whatever already triggers this script (cron, systemd timer, or still
|
||
run by hand) keeps triggering it exactly as before, and `keep pull`
|
||
either updates `.env` or silently doesn't. That's deliberate: **the
|
||
existing script, patched this way, is already a complete, correct
|
||
baseline** — keep-bootstrapped projects get fresh secrets on every run,
|
||
everything else behaves identically to today. `announce` sits on top as
|
||
a latency optimization *for however this script gets triggered*, not a
|
||
requirement for this patch to be complete — see the next section.
|
||
|
||
### 5. Where `announce`/`watch` actually earn their keep
|
||
|
||
Two independent benefits, worth keeping conceptually separate — and the
|
||
second one looks different than it would for a one-agent-per-project
|
||
design, because the real script is one bulk all-or-nothing sweep, not N
|
||
independent watchers:
|
||
|
||
- **Secrets sourcing** — `keep pull` instead of a hand-placed `.env`.
|
||
Available per-project the moment its vault exists and this machine
|
||
has a grant on it. This is what §4's patch already delivers, on its
|
||
own, with zero dependency on `announce`.
|
||
- **Waking the sweep early** — however the bulk script currently gets
|
||
triggered (cron, systemd timer, or still by hand), `announce` could
|
||
let it wake immediately on a push instead of waiting for the next
|
||
scheduled run. The natural fit given the script's actual shape: `keep
|
||
watch-any <vault1> <vault2> ...` blocking until *any* watched vault's
|
||
tag changes, used as the trigger for a full sweep — not a per-project
|
||
selective redeploy. Selective ("only redeploy the one project that
|
||
actually changed") would require reshaping the script away from its
|
||
current fixed-array-sweep model into something that takes a target
|
||
project as an argument, which is a real behavior change to something
|
||
already in production use, not a small addition — out of scope here.
|
||
|
||
Deliberately **not** building the wake-the-sweep-early mechanism in this
|
||
pass. Ship §4's patch first — it already satisfies the fallback
|
||
requirement completely and delivers the secrets-sourcing benefit on its
|
||
own. Whether waiting up to one scheduled interval for a redeploy is
|
||
actually a problem worth solving with `announce`/`watch-any` should be
|
||
judged after living with the patched bulk script for a while, not
|
||
designed in advance of any real signal that the wait matters.
|
||
|
||
### 6. CI workflow addition
|
||
|
||
One step appended to the existing `docker.yml`, after the existing
|
||
build-and-push step, best-effort (must not fail the build if this
|
||
project hasn't been bootstrapped onto `keep` yet):
|
||
|
||
```yaml
|
||
- name: announce to keep
|
||
continue-on-error: true
|
||
run: keep announce myapp/production --tag ${{ steps.meta.outputs.short_sha }}
|
||
```
|
||
|
||
The vault key here (`myapp/production`) is a **literal, hardcoded string
|
||
in this specific repo's own workflow file** — never derived from
|
||
`${{ github.repository }}` or any other automatic source. Same reasoning
|
||
as §4's `.keep-vault` file: the repo's own name, its deploy directory's
|
||
name, and its vault key are three independent strings that happen to
|
||
often look similar and are not guaranteed to match (see Context — a real
|
||
example already exists in this fleet). Whoever bootstraps a project onto
|
||
`keep` should paste the same deliberately-chosen vault key into both this
|
||
workflow step and that project's `.keep-vault` file, not rely on either
|
||
side inferring it from something else.
|
||
|
||
`continue-on-error: true` is load-bearing, not incidental — a project
|
||
without a `keep` identity registered yet must still build and push its
|
||
image successfully. This step failing should be silent/expected for any
|
||
not-yet-bootstrapped project, not a red X on the workflow.
|
||
|
||
## Security considerations
|
||
|
||
- `announce` leaks a little more than `push`/`pull` did: the *tag*
|
||
itself is now server-visible (previously, if you were being strict,
|
||
even the fact that a new image existed wasn't something `keep` needed
|
||
to know). This is an intentional, accepted tradeoff — a git short-SHA
|
||
or version tag isn't sensitive, and the whole feature only works if
|
||
`keep` can see it.
|
||
- Requiring only a read grant (not write) to `announce` means a
|
||
compromised read-only CI identity could announce a *false* tag,
|
||
pointing the fleet at an attacker-chosen image — worth flagging even
|
||
though `keep` never pulls or runs that image itself, only relays the
|
||
string. The actual `docker compose pull` on each VPS still pulls from
|
||
the configured registry using the tag it was told, so this is
|
||
equivalent in severity to that CI identity's build step being
|
||
compromised in the first place (it could already push a bad image
|
||
under a legitimate tag) — not a new capability, just worth stating
|
||
plainly rather than leaving implicit.
|
||
- The fallback path must never silently downgrade a *bootstrapped*
|
||
project to running with a stale `.env` because of a transient `keep`
|
||
outage. As written, `keep pull` failing for any reason (network
|
||
blip, `keep` server down, not just "not bootstrapped") falls through
|
||
to "deploy with whatever `.env` is already on disk" — for an
|
||
already-bootstrapped project, that's the *previous* `keep pull`'s
|
||
output, still valid, not a downgrade. This only becomes a real problem
|
||
if a bootstrapped project's very first deploy attempt coincides with
|
||
`keep` being unreachable, in which case there's no `.env` on disk yet
|
||
at all and `docker compose up` will fail for the mundane reason that
|
||
required env vars are missing — a loud, obvious failure, not a silent
|
||
wrong-secrets one.
|
||
|
||
## Open questions
|
||
|
||
- **`.keep-vault` format** — a bare vault-key string is enough for
|
||
today's one-environment-per-directory reality. If a directory ever
|
||
needs to deploy more than one environment from the same
|
||
`docker-compose.yml` (unlikely given the current fleet shape, but not
|
||
impossible), a bare string stops being enough and it'd need to become
|
||
a small key=value file or similar. Not designing that ahead of an
|
||
actual need — ship the bare-string version.
|
||
- **Fleet-wide `pull_policy: always` sweep** — a prerequisite this
|
||
proposal depends on but doesn't itself fix (see Context). Should
|
||
happen before or alongside this work, as its own small pass, not
|
||
bundled into this one.
|
||
- **Does `watch`/`watch-any` ever get used**, or does the
|
||
already-scheduled bulk sweep turn out to be good enough forever?
|
||
Deliberately left unbuilt (see §5) until there's a real signal the
|
||
wait between scheduled runs is worth solving.
|
||
- **Multi-VPS fan-out for one project** — if a project ever runs on more
|
||
than one box, `announce` already supports that for free (any number of
|
||
watchers can poll the same vault's announce endpoint), but this hasn't
|
||
been exercised and might reveal something not accounted for here.
|