diff --git a/PROPOSAL-watch.md b/PROPOSAL-watch.md new file mode 100644 index 0000000..1932b86 --- /dev/null +++ b/PROPOSAL-watch.md @@ -0,0 +1,219 @@ +# Proposal: `flit watch` — one-way folder sync over flit + +**Status:** proposed, not built. Self-contained — written so a fresh agent +with no prior conversation context can pick this up and implement it +without anything explained first. + +## Context (read this first) + +Full background on flit is in [README.md](./README.md) and +[PROTOCOL.md](./PROTOCOL.md) — read both before touching code. The +relevant existing pieces for this proposal, all in `cli/`: + +- `cmd/flit/main.go` — `send(ctx, path)`: creates a `transport.Session`, + generates a one-shot room + QR invite, and on `OnConnected` calls + `sess.SendFile(path)`. **`SendFile` is synchronous and blocking** — it + returns only once the transfer is fully complete (or errors). Its + return value alone is a sufficient "did this actually get there" + signal; no separate completion callback exists or is needed. +- `cmd/flit/daemon.go` — the pattern this proposal reuses almost + wholesale, just aimed the other direction: + - `daemonConfig`/`peerConfig` — TOML config (`~/.flit/daemon.toml`), + a list of trusted peers by `id`/`label`. + - `runPeerLoop(ctx, tc, myID, peer, downloadDir)` — maintains a + **persistent connection to one named peer**, reconnecting with + exponential backoff (2s → 30s cap) on disconnect. Sets + `OnFileOffer` to auto-accept anything the peer sends and save it to + `downloadDir`. + - This is receive-only today. **`flit watch` needs the send-side + sibling of this exact loop** — same persistent-connection/backoff + shape, but triggered by local filesystem events instead of incoming + offers. +- `internal/transport/transport.go` — `Session.Join(ctx, roomName, + trustedPeerID)`, `transport.PairRoomName(a, b)` (deterministic + pairwise room from two peer IDs — this is what makes "known device, + no fresh QR" work at all), `Session.SendFile(path)`, + `Session.PeerID()`. + +## Motivation + +At one point the idea on the table was two-way directory sync — watch +a folder, keep it mirrored to a known host, in both directions. That +was deliberately walked back to **one-way sync**: source → destination +only, push only, never reconciling anything back. That sidesteps the +entire class of problem full bidirectional sync brings with it +(conflict resolution, delete propagation, catch-up ordering). The +moment this proposal starts growing bidirectional-anything, that's a +sign it's drifting toward re-implementing a general sync tool — flit's +own case study already draws that line explicitly ("Syncthing syncs +continuously; too heavy for ad hoc one-offs" — flit is deliberately +*not* trying to be that). This proposal should stay on the "one +watched folder pushes new files to one named peer, and that's the +whole feature" side of that line. + +The actual want: drop a file into a watched folder, it gets encrypted +and pushed to a specific already-known device automatically — no +manual `flit send`, no QR, no re-running anything by hand per file. + +## Design + +### 1. The receiving side needs no changes + +Worth stating plainly since it simplifies scope: an already-running +`flit daemon`, configured with the sender as a trusted peer, already +does exactly the right thing on receipt — auto-accept, save to +`download_dir`. This proposal is entirely about the sending side. + +### 2. `flit watch` — a new subcommand, new config file + +`~/.flit/watch.toml`, same shape as `daemon.toml` deliberately, for +familiarity: + +```toml +signal_url = "wss://your-anchor.example.com/ws" +turn_url = "" +turn_secret = "" + +[[watch]] +dir = "/home/user/backups/documents" +peer_id = "" +peer_label = "home-server" +pattern = "*.pdf,*.docx" # optional glob include-list; default "*" if omitted +stable_after = "10s" # optional quiet period before considering a file ready +``` + +Multiple `[[watch]]` blocks are allowed — different folders can go to +different peers, or the same peer, independently. + +### 3. Local state = the send queue, no separate queue structure needed + +One state file per watch entry: +`~/.flit/watch-state/.json` — +`{ "": { "size": N, "mtime": N, "sha256": "...", "sent_at": N|null } }`. + +**The invariant that makes this simple**: `sent_at` is only ever set +*after* `SendFile` returns without error. A file with `sent_at: null`, +or whose current size/mtime/hash no longer matches what's recorded, is +"needs sending" — that's the entire queue. No separate retry logic: +if `SendFile` fails partway (peer disconnects mid-transfer), the state +was never advanced, so the file is exactly as eligible for retry on +the next connection as it was before the attempt. + +### 4. Startup: reconcile state against reality, don't trust only live events + +On start, before subscribing to filesystem events, walk `dir` +(respecting `pattern`) and diff against the state file. This catches +anything that changed while the watcher wasn't running — a live-only +`fsnotify` subscription would silently miss files dropped in during +downtime. + +### 5. The watch-and-send loop — `runPeerLoop`'s sibling + +```go +// runWatchLoop mirrors daemon.go's runPeerLoop: a persistent connection +// to one named peer, reconnecting with the same backoff. Differs only in +// what drives sends — a local file-ready queue instead of OnFileOffer. +func runWatchLoop(ctx context.Context, tc transport.Config, myID string, w watchConfig, state *watchState) { + room := transport.PairRoomName(myID, w.PeerID) + backoff := 2 * time.Second + + for { + if ctx.Err() != nil { + return + } + sess, err := transport.NewSession(dataDir(), tc) + // ... same connect/reconnect shape as runPeerLoop ... + sess.OnConnected = func(verified bool) { + backoff = 2 * time.Second + drainQueue(sess, w, state) // send everything currently queued, in order + } + // fsnotify events (and the startup reconciliation) push into the + // same queue; if already connected when a new file becomes ready, + // send it immediately rather than waiting for the next OnConnected. + ... + } +} +``` + +`drainQueue` sends queued files one at a time via `sess.SendFile(path)` +— sequential, not concurrent, since flit's session model is one peer +connection at a time and concurrent sends would need to be queued +DataChannel-side anyway (see "Multi-file send" in the flit case study +for the existing precedent of handling several files across one +session, just not built for a filesystem-driven trigger). + +### 6. Debounce — don't send a file mid-write + +`stable_after` is a quiet-period debounce: after the *last* fsnotify +write event for a given file, wait that long before considering it +ready. Without this, a large file being actively written would get +offered (and likely fail hash verification on the receiving end, or +worse, silently transfer a truncated version) the moment its first +bytes land. This is a real requirement, not a nice-to-have — flag it +as such if a fresh implementation is tempted to skip it for a first +cut. + +### 7. Never touches deletes + +The watcher only reacts to create/write events. A file deleted from +the source directory is never deleted from the destination, and never +should be — that's the entire point of one-way sync rather than a +mirror. Don't wire up `fsnotify`'s remove/rename events to anything +that deletes on the far end. + +## Security considerations + +- **This sends files with no per-file human confirmation** — a + deliberate, meaningful departure from interactive `flit send` + (always a conscious action) and even from `daemon`'s auto-accept + (still gated behind that peer being explicitly marked trusted with + `autoAccept` in the keyring). The safety knob here is entirely + `watch.toml` itself: which directories are named, and each entry's + `pattern`. Recommend a **`flit watch --dry-run`** mode that walks the + configured directories and prints what *would* be sent, without + connecting to anything — a way to sanity-check a new watch config + before it goes live and silently starts pushing files. +- **`pattern` should default to something narrower than `*` in + practice**, even though `*` is the technical default if omitted — + worth calling out in whatever docs ship with this, since an + unconstrained watch on a broad directory (e.g. a whole home folder) + turns "sync my documents" into "sync everything that ever lands in + this tree," which is a much bigger promise than most `[[watch]]` + entries are meant to make. +- **The destination is exactly as trusted as any other flit peer** — + this doesn't introduce a new trust primitive, it reuses the existing + Ed25519 identity + pairwise-room model as-is. A `peer_id` here is + the same kind of already-known, already-paired device as anything in + the PWA's "Known devices" list or `daemon.toml`. +- **No encryption-at-rest concern beyond what already exists** — files + are E2E encrypted in transit exactly like every other flit transfer; + this proposal doesn't change flit's crypto at all, only adds a new + trigger for when `SendFile` gets called. + +## Open questions + +- **Large initial backlogs** — if `dir` already has thousands of files + on first run, the startup reconciliation queues all of them at once. + Sequential sending (per §5) means this could take a long time before + catching up to "live." Acceptable for a first cut; a + `--max-startup-queue` or similar throttle could be added later if + this turns out to matter in practice — not designing it ahead of an + actual backlog being a problem. +- **Multiple watchers targeting the same peer simultaneously** — each + `[[watch]]` entry in §5's design opens its own `runWatchLoop`, meaning + its own separate connection/room to the same peer if two entries + share a `peer_id`. `transport.PairRoomName` is deterministic per + peer pair, so two simultaneous sessions to the same peer would + collide on the same room name — this needs either a single shared + connection per unique `peer_id` (multiple watch entries feeding one + `runWatchLoop`/queue) or confirmation that flit's transport already + tolerates multiple simultaneous sessions to the same room gracefully. + **Needs resolving before implementation, not deferred** — this is a + correctness question about the connection model, not a nice-to-have. +- **State file corruption/loss** — if `~/.flit/watch-state/*.json` is + lost (disk issue, manual deletion), the next startup reconciliation + will treat every file as unsent and re-push the entire directory. + Harmless for correctness (the destination just re-receives what it + already has) but potentially slow/wasteful for a large directory. + Worth a passing mention in user-facing docs, not worth engineering + around for a first version.