From 157111d0a3d82e930f3ae3a9a9e25229e2ae33d5 Mon Sep 17 00:00:00 2001 From: Tyler <53561637+im-tyler@users.noreply.github.com> Date: Wed, 23 Sep 2026 18:14:24 -0700 Subject: [PATCH 01/11] feat(contracts): server-status-envelope fixtures + schema correction (X02 S2 tail) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Closes the last 'pending S2 tail' artifact row. Fixtures are encoder-derived: generated from the REAL collectServerStatus via a mock SSH executor (contracts_golden_test.go, TEPLOY_UPDATE_CONTRACTS) with synthetic values; the wire shape was verified against a live `teploy server status --json` capture before pinning. - valid/full.json: complete healthy observation (uptime, load, memory, two disks, docker inventory with container+image, caddy routes, no errors) - valid/partial-caddy-unavailable.json: the class a real target without a caddy container produces — caddy.routes error entry, empty inventory - legacy/pre-mi.json: machine_interface absent (42243e2-era shape) Schema defect fixed in the same commit: server-status-envelope had copied the appStatus root since its S2 draft (08cfb1b's own message admits the copy) and never described the actual serverStatusDTO wire format. Rewritten to the real root (server/host/uptime/load/memory/ disks/docker/caddy/observed_at/errors) with required-key coverage matching the DTO's no-omitempty fields; $defs for every nested type. Consumers see this as additive — the old schema matched nothing on the wire. Corpus rev 5 (MANIFEST); no MI bump, no token change. --- contracts/MANIFEST.md | 3 +- .../server-status-envelope/legacy/pre-mi.json | 93 +++++ .../server-status-envelope/valid/full.json | 94 +++++ .../valid/partial-caddy-unavailable.json | 45 +++ .../schema/server-status-envelope.schema.json | 335 ++++++++++++------ internal/cli/contracts_golden_test.go | 67 ++++ 6 files changed, 523 insertions(+), 114 deletions(-) create mode 100644 contracts/fixtures/server-status-envelope/legacy/pre-mi.json create mode 100644 contracts/fixtures/server-status-envelope/valid/full.json create mode 100644 contracts/fixtures/server-status-envelope/valid/partial-caddy-unavailable.json diff --git a/contracts/MANIFEST.md b/contracts/MANIFEST.md index b3375e6..54e8618 100644 --- a/contracts/MANIFEST.md +++ b/contracts/MANIFEST.md @@ -10,6 +10,7 @@ Neutron/Nucleus dependency and a public mirror. | Corpus rev | Emitting CLI | Machine Interface | Notes | |---|---|---|---| +| 5 | main (X02 S2 tail: server-status fixtures + schema correction) | 2 | server-status-envelope fixtures landed (was "pending live capture"): valid x2 (full healthy observation, partial-caddy-unavailable — the class a target without a caddy container produces) + legacy pre-MI (machine_interface absent, the 42243e2-era shape). Encoder-derived: generated from the REAL `collectServerStatus` via a mock SSH executor (`contracts_golden_test.go`, TEPLOY_UPDATE_CONTRACTS) — synthetic values, real encoder and parse stages; the wire shape was verified against a live `server status --json` run before pinning. Defect fixed in the same commit: the schema had copied the appStatus root since its S2 draft (its own defect-fix commit 08cfb1b said so) and never described the actual serverStatusDTO wire format (server/host/uptime/load/memory/disks/docker/caddy) — rewritten to the real root with strict required-key coverage of the DTO's no-omitempty fields. Additive to consumers (a schema that matched nothing before now matches the wire); no MI bump. | | 4 | main (X02 S2 tail: server-list reshape) | 2 | **The MI 2 bump** (D8 non-additive): `server list --json` now emits the envelope `{machine_interface, servers[], observed_at}` carrying the per-server fields unchanged (name + id/host/user/role/tags/vpn_ip); the pre-reshape bare map-of-servers root is GONE on the wire and is pinned as the artifact's legacy class. New artifact server-list-envelope (schema + valid + legacy fixtures); version-handshake schema maximum 1→2 and its valid fixture renamed mi1→mi2 (app-list valid likewise — both envelopes now report MI 2). Capability tokens unchanged. Coordinated consumer: teploy-dash decodes both shapes during the transition (MaxSupportedMachineInterface 2). | | 3 (amended) | main (C05 plan-record corpus + defect fix) | 1 | C05 added the plan-record artifact + plan-apply token (see git history); amendment: server-status schema now carries its own $defs (its $refs never resolved), and app-list fixtures emit [] where the encoder emits [] (null fixtures failed schema + the real dash decode - found by dash's new contracts CI job, fixed here). | | 1 | post-v0.1.37 main (S2 skeleton) | 1 | First goldens: version handshake, app-list envelope (MI + pre-MI legacy), error envelope (config-invalid, internal, invalid-code), release-record, attempt-name grammar, preview-state eras. | @@ -23,7 +24,7 @@ Neutron/Nucleus dependency and a public mirror. | version-handshake | yes | valid (real `writeVersion` encoder) | teploy-cli | | app-list-envelope | yes | valid (real DTO tags) + legacy pre-MI | teploy-cli | | server-list-envelope | yes (MI 2 reshape) | valid (real `writeServerList` encoder) + legacy bare-map | teploy-cli | -| server-status-envelope | yes (appStatus root) | pending S2 tail (live `server status` capture) | teploy-cli | +| server-status-envelope | yes (serverStatusDTO root, corrected rev 5) | valid x2 (full, partial-caddy-unavailable; real `collectServerStatus` encoder over mock executor) + legacy pre-MI | teploy-cli | | error-envelope | yes | valid x2 + invalid code | teploy-cli | | release-record | yes | valid container | teploy-cli | | attempt-name | yes (pattern) | valid + invalid examples | teploy-cli | diff --git a/contracts/fixtures/server-status-envelope/legacy/pre-mi.json b/contracts/fixtures/server-status-envelope/legacy/pre-mi.json new file mode 100644 index 0000000..14fcada --- /dev/null +++ b/contracts/fixtures/server-status-envelope/legacy/pre-mi.json @@ -0,0 +1,93 @@ +{ + "caddy": { + "available": true, + "routes": [ + { + "handlers": [ + "reverse_proxy" + ], + "hosts": [ + "myapp.example.com" + ], + "id": "", + "server": "srv0", + "status_code": "", + "upstreams": [ + "myapp-web-3:3000" + ] + }, + { + "handlers": [ + "subroute" + ], + "hosts": [ + "myapp.example.com" + ], + "id": "myapp", + "server": "srv0", + "status_code": "", + "upstreams": [] + } + ] + }, + "disks": [ + { + "available_bytes": 750, + "filesystem": "/dev/vda1", + "mountpoint": "/", + "total_bytes": 1000, + "used_bytes": 250, + "used_percent": "25%" + }, + { + "available_bytes": 1500, + "filesystem": "/dev/vdb1", + "mountpoint": "/srv", + "total_bytes": 2000, + "used_bytes": 500, + "used_percent": "26%" + } + ], + "docker": { + "containers": [ + { + "created_at": "2026-09-23 11:55:00 +0000 UTC", + "id": "9f31c02", + "image": "example/myapp:3", + "name": "myapp-web-3", + "process": "web", + "state": "running", + "status": "Up 4 minutes", + "version": "3" + } + ], + "images": [ + { + "created_at": "2026-09-23 11:50:00 +0000 UTC", + "id": "sha256:1a2b3c4d5e6f", + "repository": "example/myapp", + "size": "25MB", + "tag": "3" + } + ], + "installed": true, + "version": "29.0.0" + }, + "errors": [], + "host": "192.0.2.10", + "load": { + "fifteen": 0.3, + "five": 0.2, + "one": 0.1 + }, + "memory": { + "available_bytes": 409600, + "total_bytes": 1024000, + "used_bytes": 614400 + }, + "observed_at": "2026-09-23T12:00:00Z", + "server": "prod", + "uptime": { + "seconds": 3600.5 + } +} diff --git a/contracts/fixtures/server-status-envelope/valid/full.json b/contracts/fixtures/server-status-envelope/valid/full.json new file mode 100644 index 0000000..9271c4f --- /dev/null +++ b/contracts/fixtures/server-status-envelope/valid/full.json @@ -0,0 +1,94 @@ +{ + "machine_interface": 2, + "server": "prod", + "host": "192.0.2.10", + "uptime": { + "seconds": 3600.5 + }, + "load": { + "one": 0.1, + "five": 0.2, + "fifteen": 0.3 + }, + "memory": { + "total_bytes": 1024000, + "used_bytes": 614400, + "available_bytes": 409600 + }, + "disks": [ + { + "filesystem": "/dev/vda1", + "mountpoint": "/", + "total_bytes": 1000, + "used_bytes": 250, + "available_bytes": 750, + "used_percent": "25%" + }, + { + "filesystem": "/dev/vdb1", + "mountpoint": "/srv", + "total_bytes": 2000, + "used_bytes": 500, + "available_bytes": 1500, + "used_percent": "26%" + } + ], + "docker": { + "installed": true, + "version": "29.0.0", + "containers": [ + { + "id": "9f31c02", + "name": "myapp-web-3", + "image": "example/myapp:3", + "state": "running", + "status": "Up 4 minutes", + "created_at": "2026-09-23 11:55:00 +0000 UTC", + "process": "web", + "version": "3" + } + ], + "images": [ + { + "id": "sha256:1a2b3c4d5e6f", + "repository": "example/myapp", + "tag": "3", + "size": "25MB", + "created_at": "2026-09-23 11:50:00 +0000 UTC" + } + ] + }, + "caddy": { + "available": true, + "routes": [ + { + "server": "srv0", + "id": "", + "hosts": [ + "myapp.example.com" + ], + "handlers": [ + "reverse_proxy" + ], + "upstreams": [ + "myapp-web-3:3000" + ], + "status_code": "" + }, + { + "server": "srv0", + "id": "myapp", + "hosts": [ + "myapp.example.com" + ], + "handlers": [ + "subroute" + ], + "upstreams": [], + "status_code": "" + } + ] + }, + "observed_at": "2026-09-23T12:00:00Z", + "errors": [] +} diff --git a/contracts/fixtures/server-status-envelope/valid/partial-caddy-unavailable.json b/contracts/fixtures/server-status-envelope/valid/partial-caddy-unavailable.json new file mode 100644 index 0000000..f75bfbe --- /dev/null +++ b/contracts/fixtures/server-status-envelope/valid/partial-caddy-unavailable.json @@ -0,0 +1,45 @@ +{ + "machine_interface": 2, + "server": "staging", + "host": "192.0.2.20", + "uptime": { + "seconds": 86400 + }, + "load": { + "one": 0, + "five": 0.01, + "fifteen": 0.05 + }, + "memory": { + "total_bytes": 512000, + "used_bytes": 256000, + "available_bytes": 256000 + }, + "disks": [ + { + "filesystem": "/dev/vda1", + "mountpoint": "/", + "total_bytes": 500, + "used_bytes": 100, + "available_bytes": 400, + "used_percent": "20%" + } + ], + "docker": { + "installed": true, + "version": "29.0.0", + "containers": [], + "images": [] + }, + "caddy": { + "available": false, + "routes": [] + }, + "observed_at": "2026-09-23T12:00:00Z", + "errors": [ + { + "scope": "caddy.routes", + "message": "Error response from daemon: No such container: caddy" + } + ] +} diff --git a/contracts/schema/server-status-envelope.schema.json b/contracts/schema/server-status-envelope.schema.json index 7e43687..24f33d9 100644 --- a/contracts/schema/server-status-envelope.schema.json +++ b/contracts/schema/server-status-envelope.schema.json @@ -1,91 +1,166 @@ { "$schema": "https://json-schema.org/draft/2020-12/schema", "$id": "https://teploy.github.io/contracts/schema/server-status-envelope.schema.json", - "title": "server status --json envelope (MI 1, appStatusDTO root)", - "allOf": [ - { + "title": "server status --json envelope (serverStatusDTO root)", + "type": "object", + "required": [ + "machine_interface", + "server", + "host", + "uptime", + "load", + "memory", + "disks", + "docker", + "caddy", + "observed_at", + "errors" + ], + "properties": { + "machine_interface": { + "type": "integer", + "minimum": 1 + }, + "server": { + "type": "string" + }, + "host": { + "type": "string" + }, + "uptime": { + "$ref": "#/$defs/uptime" + }, + "load": { + "$ref": "#/$defs/load" + }, + "memory": { + "$ref": "#/$defs/memory" + }, + "disks": { + "type": "array", + "items": { + "$ref": "#/$defs/disk" + } + }, + "docker": { + "$ref": "#/$defs/dockerInventory" + }, + "caddy": { + "$ref": "#/$defs/caddyObservation" + }, + "observed_at": { + "type": "string", + "format": "date-time" + }, + "errors": { + "type": "array", + "items": { + "$ref": "#/$defs/machineError" + } + } + }, + "$defs": { + "machineError": { "type": "object", "required": [ - "app", - "domain", - "type", - "ingress", - "current_release", - "previous_release", - "containers", - "processes", - "lock", - "maintenance", - "observed_at", - "errors", - "machine_interface" + "scope", + "message" ], "properties": { - "app": { - "type": "string" - }, - "domain": { - "type": "string" - }, - "type": { + "scope": { "type": "string" }, - "ingress": { + "message": { "type": "string" + } + } + }, + "uptime": { + "type": "object", + "required": [ + "seconds" + ], + "properties": { + "seconds": { + "type": "number", + "minimum": 0 + } + } + }, + "load": { + "type": "object", + "required": [ + "one", + "five", + "fifteen" + ], + "properties": { + "one": { + "type": "number", + "minimum": 0 }, - "current_release": { - "$ref": "#/$defs/release" - }, - "previous_release": { - "$ref": "#/$defs/release" - }, - "containers": { - "type": "array", - "items": { - "$ref": "#/$defs/container" - } - }, - "processes": { - "type": "array" - }, - "lock": { - "type": [ - "object", - "null" - ] - }, - "maintenance": { - "type": "boolean" + "five": { + "type": "number", + "minimum": 0 }, - "observed_at": { - "type": "string", - "format": "date-time" + "fifteen": { + "type": "number", + "minimum": 0 + } + } + }, + "memory": { + "type": "object", + "required": [ + "total_bytes", + "used_bytes", + "available_bytes" + ], + "properties": { + "total_bytes": { + "type": "integer", + "minimum": 0 }, - "errors": { - "type": "array", - "items": { - "$ref": "#/$defs/machineError" - } + "used_bytes": { + "type": "integer", + "minimum": 0 }, - "machine_interface": { + "available_bytes": { "type": "integer", - "minimum": 1 + "minimum": 0 } } - } - ], - "$defs": { - "machineError": { + }, + "disk": { "type": "object", "required": [ - "scope", - "message" + "filesystem", + "mountpoint", + "total_bytes", + "used_bytes", + "available_bytes", + "used_percent" ], "properties": { - "scope": { + "filesystem": { "type": "string" }, - "message": { + "mountpoint": { + "type": "string" + }, + "total_bytes": { + "type": "integer", + "minimum": 0 + }, + "used_bytes": { + "type": "integer", + "minimum": 0 + }, + "available_bytes": { + "type": "integer", + "minimum": 0 + }, + "used_percent": { "type": "string" } } @@ -97,7 +172,10 @@ "name", "image", "state", - "status" + "status", + "created_at", + "process", + "version" ], "properties": { "id": { @@ -126,85 +204,116 @@ } } }, - "release": { + "image": { "type": "object", "required": [ + "id", + "repository", + "tag", + "size", + "created_at" + ], + "properties": { + "id": { + "type": "string" + }, + "repository": { + "type": "string" + }, + "tag": { + "type": "string" + }, + "size": { + "type": "string" + }, + "created_at": { + "type": "string" + } + } + }, + "dockerInventory": { + "type": "object", + "required": [ + "installed", "version", - "ports" + "containers", + "images" ], "properties": { + "installed": { + "type": "boolean" + }, "version": { "type": "string" }, - "ports": { + "containers": { "type": "array", "items": { - "type": "integer" + "$ref": "#/$defs/container" + } + }, + "images": { + "type": "array", + "items": { + "$ref": "#/$defs/image" } } } }, - "appStatus": { + "caddyRoute": { "type": "object", "required": [ - "app", - "domain", - "type", - "ingress", - "current_release", - "previous_release", - "containers", - "processes", - "lock", - "maintenance", - "observed_at", - "errors" + "server", + "id", + "hosts", + "handlers", + "upstreams", + "status_code" ], "properties": { - "app": { + "server": { "type": "string" }, - "domain": { - "type": "string" - }, - "type": { - "type": "string" - }, - "ingress": { + "id": { "type": "string" }, - "current_release": { - "$ref": "#/$defs/release" - }, - "previous_release": { - "$ref": "#/$defs/release" - }, - "containers": { + "hosts": { "type": "array", "items": { - "$ref": "#/$defs/container" + "type": "string" } }, - "processes": { - "type": "array" + "handlers": { + "type": "array", + "items": { + "type": "string" + } }, - "lock": { - "type": [ - "object", - "null" - ] + "upstreams": { + "type": "array", + "items": { + "type": "string" + } }, - "maintenance": { + "status_code": { + "type": "string" + } + } + }, + "caddyObservation": { + "type": "object", + "required": [ + "available", + "routes" + ], + "properties": { + "available": { "type": "boolean" }, - "observed_at": { - "type": "string", - "format": "date-time" - }, - "errors": { + "routes": { "type": "array", "items": { - "$ref": "#/$defs/machineError" + "$ref": "#/$defs/caddyRoute" } } } diff --git a/internal/cli/contracts_golden_test.go b/internal/cli/contracts_golden_test.go index 1a2d21a..4ac7f9c 100644 --- a/internal/cli/contracts_golden_test.go +++ b/internal/cli/contracts_golden_test.go @@ -9,7 +9,9 @@ package cli import ( "bytes" + "context" "encoding/json" + "errors" "os" "path/filepath" "strings" @@ -18,6 +20,7 @@ import ( "github.com/useteploy/teploy/internal/config" "github.com/useteploy/teploy/internal/releasemeta" + "github.com/useteploy/teploy/internal/ssh" ) const contractsDir = "../../contracts" @@ -131,6 +134,70 @@ func TestContractsAppListEnvelopeGolden(t *testing.T) { writeFixture(t, "app-list-envelope/legacy/pre-mi.json", legacy) } +// TestContractsServerStatusEnvelopeGolden drives the REAL +// collectServerStatus encoder (the same collection path `server status +// --json` runs) through a mock SSH executor — the DTO values are +// synthetic, the encoder and every parse stage (memory, disks, docker +// inventory, Caddy routes) are the real ones. The wire shape was +// verified against a live `server status --json` capture before this +// fixture was pinned (see MANIFEST rev 5). Two valid classes: a full +// healthy observation, and a partial one with the Caddy probe failing — +// the class a real deployment without a caddy container produces. The +// legacy fixture is the pre-MI shape (machine_interface absent), which a +// 42243e2-era CLI emitted. +func TestContractsServerStatusEnvelopeGolden(t *testing.T) { + observedAt := time.Date(2026, 9, 23, 12, 0, 0, 0, time.UTC) + container := `{"ID":"9f31c02","Names":"myapp-web-3","Image":"example/myapp:3","State":"running","Status":"Up 4 minutes","CreatedAt":"2026-09-23 11:55:00 +0000 UTC","Labels":"teploy.app=myapp,teploy.process=web,teploy.version=3"}` + image := `{"ID":"sha256:1a2b3c4d5e6f","Repository":"example/myapp","Tag":"3","Size":"25MB","CreatedAt":"2026-09-23 11:50:00 +0000 UTC"}` + caddy := `{"servers":{"srv0":{"routes":[{"@id":"myapp","match":[{"host":["myapp.example.com"]}],"handle":[{"handler":"subroute","routes":[{"handle":[{"handler":"reverse_proxy","upstreams":[{"dial":"myapp-web-3:3000"}]}]}]}]}]}}}` + full := ssh.NewMockExecutor("192.0.2.10", + ssh.MockCommand{Match: "cat /proc/uptime", Output: "3600.50 1200.00"}, + ssh.MockCommand{Match: "cat /proc/loadavg", Output: "0.10 0.20 0.30 1/100 1"}, + ssh.MockCommand{Match: "cat /proc/meminfo", Output: "MemTotal: 1000 kB\nMemAvailable: 400 kB\n"}, + ssh.MockCommand{Match: "df -B1 -P", Output: "Filesystem 1-blocks Used Available Capacity Mounted on\n/dev/vda1 1000 250 750 25% /\n/dev/vdb1 2000 500 1500 26% /srv\n"}, + ssh.MockCommand{Match: "docker version", Output: "29.0.0"}, + ssh.MockCommand{Match: "docker ps --all", Output: container}, + ssh.MockCommand{Match: "docker image ls", Output: image}, + ssh.MockCommand{Match: "docker exec caddy", Output: caddy}, + ) + got := collectServerStatus(context.Background(), full, "prod", observedAt) + if len(got.Errors) != 0 { + t.Fatalf("full observation reported errors: %#v", got.Errors) + } + fullStatus := got + writeFixture(t, "server-status-envelope/valid/full.json", got) + + partial := ssh.NewMockExecutor("192.0.2.20", + ssh.MockCommand{Match: "cat /proc/uptime", Output: "86400.00 86400.00"}, + ssh.MockCommand{Match: "cat /proc/loadavg", Output: "0.00 0.01 0.05 1/100 1"}, + ssh.MockCommand{Match: "cat /proc/meminfo", Output: "MemTotal: 500 kB\nMemAvailable: 250 kB\n"}, + ssh.MockCommand{Match: "df -B1 -P", Output: "Filesystem 1-blocks Used Available Capacity Mounted on\n/dev/vda1 500 100 400 20% /\n"}, + ssh.MockCommand{Match: "docker version", Output: "29.0.0"}, + ssh.MockCommand{Match: "docker ps --all", Output: ""}, + ssh.MockCommand{Match: "docker image ls", Output: ""}, + ssh.MockCommand{Match: "docker exec caddy", Err: errors.New("Error response from daemon: No such container: caddy")}, + ) + got = collectServerStatus(context.Background(), partial, "staging", observedAt) + if len(got.Errors) != 1 || got.Errors[0].Scope != "caddy.routes" { + t.Fatalf("partial observation missing its caddy error: %#v", got.Errors) + } + writeFixture(t, "server-status-envelope/valid/partial-caddy-unavailable.json", got) + + // Legacy: the pre-MI envelope (no machine_interface field) a CLI + // between 42243e2 and dda4911 emitted. Same delete-from-map approach + // as the app-list legacy class. + raw, err := json.Marshal(fullStatus) + if err != nil { + t.Fatalf("marshal full status: %v", err) + } + var legacy map[string]any + if err := json.Unmarshal(raw, &legacy); err != nil { + t.Fatalf("unmarshal legacy: %v", err) + } + delete(legacy, "machine_interface") + writeFixture(t, "server-status-envelope/legacy/pre-mi.json", legacy) +} + // TestContractsErrorEnvelopeGolden pins the two wired error classes. func TestContractsErrorEnvelopeGolden(t *testing.T) { writeFixture(t, "error-envelope/valid/config-invalid.json", machineErrorEnvelope{ From 8a20cac3e20ef5af714d0d2d4566e902e3f85d85 Mon Sep 17 00:00:00 2001 From: Tyler <53561637+im-tyler@users.noreply.github.com> Date: Wed, 23 Sep 2026 18:17:10 -0700 Subject: [PATCH 02/11] feat(cli): surface interrupted deploys as uncertain-outcome (C09 tail) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The C09 acceptance wants scripted automation to distinguish rejected, failed, uncertain, canceled and successful outcomes. The machinery produced all of these but the taxonomy classified an interrupted deploy (SIGINT, automation timeout) as plain internal — a canceled deploy and a failed one were indistinguishable on the wire. - classifyMachineError: context.Canceled/DeadlineExceeded (however deeply wrapped) -> uncertain-outcome, message names reconciliation. The C01 crash-window contract: an interrupted window is INSPECT, never a rendered failure. Exit codes unchanged (0/1/2, D10). - wrapDeployOutcomeError: the human path for a failed deploy names the recovery action (re-run is safe; inspect with teploy status) instead of a bare 'context canceled'. - help consistency: deploy --help now states outcome/exit-code semantics and the --json error-envelope codes; version --help documents the --json machine handshake (MI number + capability registry). No deploy/state internals touched; classification stays CLI-side. --- internal/cli/deploy.go | 49 +++++++++++++++++++-------- internal/cli/errevelope.go | 16 +++++++-- internal/cli/errevelope_test.go | 59 +++++++++++++++++++++++++++++++++ internal/cli/version.go | 6 ++++ 4 files changed, 114 insertions(+), 16 deletions(-) diff --git a/internal/cli/deploy.go b/internal/cli/deploy.go index ea765ae..f4cb0ef 100644 --- a/internal/cli/deploy.go +++ b/internal/cli/deploy.go @@ -61,7 +61,15 @@ If a previously-deployed container mounted volumes from a different host path (common when migrating from Dokploy or hand-rolled docker run setups), the deploy aborts safely rather than orphaning data. Pass --migrate-volumes to copy data from the existing source into the teploy-expected path before -swapping traffic.`, +swapping traffic. + +Outcomes and exit codes: 0 = deployed (warnings possible — read the output); +1 = refused before any effect (config/admission problems), failed, or +INTERRUPTED — an interrupted deploy (Ctrl-C, timeout) has an unknown outcome +until reconciled; re-running is safe (the next attempt reconciles partial +state), or inspect with teploy status. Under --json a failure is reported as +a machine error envelope on stderr whose code distinguishes the classes +(config-invalid, conflict, uncertain-outcome, internal).`, Args: cobra.MaximumNArgs(1), RunE: func(cmd *cobra.Command, args []string) error { var serverName string @@ -713,7 +721,7 @@ func deployBuiltImageFenced(ctx context.Context, executor ssh.Executor, appCfg * emitDeployAudit(ctx, appCfg, "deploy.run", version, serverDisplay, deployErr) if deployErr != nil { - return deployErr + return wrapDeployOutcomeError(deployErr) } // 13. Prune old build images (best-effort). @@ -725,6 +733,19 @@ func deployBuiltImageFenced(ctx context.Context, executor ssh.Executor, appCfg * return nil } +// wrapDeployOutcomeError renders a failed deploy's error for the human +// path. An interrupted deploy (Ctrl-C, automation timeout) leaves the +// outcome UNKNOWN — the C01 crash-window contract — so it names the +// recovery action instead of letting a bare "context canceled" read as a +// clean failure. Under --json the machine envelope classifies the same +// error as uncertain-outcome (errevelope.go). +func wrapDeployOutcomeError(err error) error { + if errors.Is(err, context.Canceled) || errors.Is(err, context.DeadlineExceeded) { + return fmt.Errorf("deploy interrupted: %w\nThe outcome is unknown until reconciled. Re-running this deploy is safe (the next attempt reconciles any partial state); inspect first with: teploy status", err) + } + return err +} + // plannedVolumeMounts resolves config volume declarations to the docker // mount map a deploy uses: a host bind (name starting with "/") mounts a // directory the operator owns, exactly as given — teploy never creates or @@ -969,18 +990,18 @@ func runMultiDeploy(flags *Flags, appCfg *config.AppConfig, image, version strin fmt.Printf("\n%d of %d servers failed — rolling back the %d server(s) that succeeded...\n", failCount, len(targets), len(successTargets)) - // Best-effort: attempt to roll back EVERY succeeded server even if one - // rollback fails — otherwise a single rollback failure would fail-fast - // and strand the remaining servers on the new version (M1). The wave - // runs on a bounded DETACHED recovery context (audit T58): the deploy - // context is signal-cancelled exactly when the operator interrupts, - // and recovery work that skips itself because the cancelled context - // disappeared is how a Ctrl-C strands half a fleet on the new version. - rollbackCtx, rollbackCancel := context.WithTimeout(context.WithoutCancel(ctx), 10*time.Minute) - defer rollbackCancel() - rollbackResults := multideploy.ParallelDeployAll(rollbackCtx, successTargets, parallel, func(ctx context.Context, target multideploy.ServerTarget, out io.Writer) error { - return rollbackSingleServer(ctx, appCfg, target, out) - }, os.Stdout) + // Best-effort: attempt to roll back EVERY succeeded server even if one + // rollback fails — otherwise a single rollback failure would fail-fast + // and strand the remaining servers on the new version (M1). The wave + // runs on a bounded DETACHED recovery context (audit T58): the deploy + // context is signal-cancelled exactly when the operator interrupts, + // and recovery work that skips itself because the cancelled context + // disappeared is how a Ctrl-C strands half a fleet on the new version. + rollbackCtx, rollbackCancel := context.WithTimeout(context.WithoutCancel(ctx), 10*time.Minute) + defer rollbackCancel() + rollbackResults := multideploy.ParallelDeployAll(rollbackCtx, successTargets, parallel, func(ctx context.Context, target multideploy.ServerTarget, out io.Writer) error { + return rollbackSingleServer(ctx, appCfg, target, out) + }, os.Stdout) var rolledBack, firstDeploys, rollbackFailed []string for _, r := range rollbackResults { diff --git a/internal/cli/errevelope.go b/internal/cli/errevelope.go index 04505c7..b313f6f 100644 --- a/internal/cli/errevelope.go +++ b/internal/cli/errevelope.go @@ -1,6 +1,7 @@ package cli import ( + "context" "encoding/json" "errors" "fmt" @@ -39,7 +40,9 @@ const ( // (UNMIGRATED: currently internal). codeConflict = "conflict" // The effect's fate is unknown pending reconciliation; never - // rendered or recorded as failure (UNMIGRATED: currently internal). + // rendered or recorded as failure. Wired for interrupted commands + // (context canceled / deadline exceeded — SIGINT, automation + // timeouts); deeper per-phase migration stays on the S2 list. codeUncertainOutcome = "uncertain-outcome" // Success with a flag — traffic switched but the outcome is not // clean (UNMIGRATED: currently internal). @@ -87,12 +90,19 @@ func refuseAdmission(err error) error { // refusals → config-invalid; an absent config is a config failure for a // machine caller the same way a malformed one is. A plan/apply drift // refusal (C05) → conflict — the request is coherent, the world moved. -// Everything else is internal until its site is migrated (S2 +// An interrupted command (context canceled or deadline exceeded — the +// SIGINT path, an automation timeout) → uncertain-outcome: whatever +// effect was in flight has an unknown fate until reconciled (the C01 +// crash-window machinery treats every interrupted window as INSPECT), +// which is precisely the class automation must not render as a plain +// failure. Everything else is internal until its site is migrated (S2 // generalization). func classifyMachineError(err error) string { switch { case errors.Is(err, errPlanDrift): return codeConflict + case errors.Is(err, context.Canceled), errors.Is(err, context.DeadlineExceeded): + return codeUncertainOutcome case errors.Is(err, config.ErrInvalidConfig), errors.Is(err, config.ErrNoConfig), errors.Is(err, errDeployAdmission): @@ -114,6 +124,8 @@ func writeMachineErrorEnvelope(out io.Writer, err error) error { } case codeConflict: message = "plan no longer valid" + case codeUncertainOutcome: + message = "interrupted — outcome unknown until reconciled" } return json.NewEncoder(out).Encode(machineErrorEnvelope{ MachineInterface: MachineInterface, diff --git a/internal/cli/errevelope_test.go b/internal/cli/errevelope_test.go index 079fde1..8381caf 100644 --- a/internal/cli/errevelope_test.go +++ b/internal/cli/errevelope_test.go @@ -2,8 +2,10 @@ package cli import ( "bytes" + "context" "encoding/json" "errors" + "fmt" "strings" "testing" @@ -150,6 +152,63 @@ func TestMachineErrorEnvelopeUnclassifiedIsInternal(t *testing.T) { } } +// TestMachineErrorEnvelopeInterruptedIsUncertain: an interrupted command +// (context canceled — the SIGINT path — or a deadline exceeded) is NOT a +// plain failure: its outcome is unknown until reconciled (C09's +// uncertain/canceled distinction for automation). Wrapped and chained +// errors must classify the same way. +func TestMachineErrorEnvelopeInterruptedIsUncertain(t *testing.T) { + for name, err := range map[string]error{ + "canceled": context.Canceled, + "deadline": context.DeadlineExceeded, + "wrapped": fmt.Errorf("deploying myapp: %w", context.Canceled), + "double-wrapped": fmt.Errorf("running docker run: %w", fmt.Errorf("build: %w", context.DeadlineExceeded)), + } { + var out bytes.Buffer + if writeErr := writeMachineErrorEnvelope(&out, err); writeErr != nil { + t.Fatal(writeErr) + } + var decoded map[string]any + if jsonErr := json.Unmarshal(out.Bytes(), &decoded); jsonErr != nil { + t.Fatalf("%s: envelope not JSON: %q", name, out.String()) + } + if decoded["code"] != "uncertain-outcome" { + t.Fatalf("%s: code = %v, want uncertain-outcome", name, decoded["code"]) + } + if decoded["message"] != "interrupted — outcome unknown until reconciled" { + t.Fatalf("%s: message = %v", name, decoded["message"]) + } + } +} + +// TestDeployInterruptedErrorNamesRecovery: the human-path error for an +// interrupted deploy names the recovery action instead of a bare +// "context canceled" (which reads as a clean failure), while a plain +// deploy failure passes through unchanged. +func TestDeployInterruptedErrorNamesRecovery(t *testing.T) { + interrupted := wrapDeployOutcomeError(fmt.Errorf("deploying myapp: %w", context.Canceled)) + msg := interrupted.Error() + if !strings.Contains(msg, "deploy interrupted") || !strings.Contains(msg, "reconcil") || !strings.Contains(msg, "teploy status") { + t.Fatalf("interrupted deploy error must name the recovery action: %q", msg) + } + var out bytes.Buffer + if err := writeMachineErrorEnvelope(&out, interrupted); err != nil { + t.Fatal(err) + } + var decoded map[string]any + if err := json.Unmarshal(out.Bytes(), &decoded); err != nil { + t.Fatalf("envelope not JSON: %q", out.String()) + } + if decoded["code"] != "uncertain-outcome" { + t.Fatalf("interrupted deploy code = %v, want uncertain-outcome", decoded["code"]) + } + + plain := errors.New("image pull failed") + if got := wrapDeployOutcomeError(plain); got != plain { + t.Fatalf("plain failure must pass through unchanged: %v", got) + } +} + // TestExitCodesPinned: 0/1/2 semantics are unchanged by the envelope — // drift's 2 remains a signal (not a failure) gated on --exit-code, and // every other failure exits 1 after the error is reported. diff --git a/internal/cli/version.go b/internal/cli/version.go index c97b39b..9bf60e5 100644 --- a/internal/cli/version.go +++ b/internal/cli/version.go @@ -21,6 +21,12 @@ func newVersionCmd(flags *Flags, version string) *cobra.Command { return &cobra.Command{ Use: "version", Short: "Show teploy version", + Long: `Show teploy version. + +With --json, emits the machine-interface handshake instead: the version, +the machine_interface number (fail closed if a consumer supports less), +and the capability tokens this build advertises. One call replaces +help-text scraping as the compatibility check.`, RunE: func(cmd *cobra.Command, args []string) error { return writeVersion(cmd.OutOrStdout(), version, flags.JSON) }, From d27ea1ba1253b093aa93acc864ac2b0dbf84c072 Mon Sep 17 00:00:00 2001 From: R02 lane Date: Wed, 23 Sep 2026 18:20:10 -0700 Subject: [PATCH 03/11] docs(r02): supported-workload matrix, first-success tutorial, failure/recovery guide, migration recipes; README superlative sweep - docs/supported-workloads.md: what deploys (single-image apps, multi-process, static, Compose subset, templates, previews, accessories) vs what is refused (multi-image stacks, ambiguous web candidates, per-field Compose refusals) with the refusal behavior named; ingress guarantees per mode; declared operational limits. - docs/first-success.md: install -> server add -> deploy -> verify -> rollback, executed against a scratch SSH+Docker host (colima fixture); each step's failure mode and remedy (admission errors, rsync, health gate, doctor). - docs/failure-and-recovery.md: exit codes + --json error envelope taxonomy (wired vs reserved codes marked), doctor check table, crash-recovery dispositions (RETRY/INSPECT/COMPENSATE/MANUAL) and repair debt, self-heal, DR bundle family (dr create/list/show/restore verified on the fixture; cutover documented). - docs/migration.md: Dokploy/Coolify -> teploy via the C05 Compose import (converts/refuses tables, concept mapping, reversible adoption). - README: replace unsupported superlatives with scoped verified statements (zero-downtime scoped to Caddy blue/green; rollback no longer 'instantly'; drop 'No dependencies'); add Docs index; honest Requirements. --- README.md | 34 ++++++-- docs/failure-and-recovery.md | 160 +++++++++++++++++++++++++++++++++++ docs/first-success.md | 146 ++++++++++++++++++++++++++++++++ docs/migration.md | 96 +++++++++++++++++++++ docs/supported-workloads.md | 61 +++++++++++++ 5 files changed, 488 insertions(+), 9 deletions(-) create mode 100644 docs/failure-and-recovery.md create mode 100644 docs/first-success.md create mode 100644 docs/migration.md create mode 100644 docs/supported-workloads.md diff --git a/README.md b/README.md index 99e11a9..18d411f 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@

teploy

-

Zero-downtime Docker deploys to any server via SSH.
Single binary. No management server. No dependencies.

+

Docker deploys to any Linux server you can SSH into, with blue/green zero-downtime under Caddy.
Single binary. No management server.

@@ -13,7 +13,7 @@ ## Why teploy? -Most deploy tools require either a management server (Coolify, Dokploy) or complex configuration (Kamal). Teploy is a single binary that deploys Docker containers to any server you can SSH into. Three lines of config, one command to deploy. +Most deploy tools require either a management server (Coolify, Dokploy) or longer configuration (Kamal). Teploy is a single binary that deploys Docker containers to any server you can SSH into. Three lines of config, one command to deploy. ```yaml # teploy.yml @@ -26,7 +26,11 @@ server: 1.2.3.4 teploy deploy ``` -Your app is live with HTTPS, zero-downtime deploys, and automatic rollback on failure. +With the default Caddy ingress your app is live with automatic HTTPS, +blue/green zero-downtime deploys, and rollback to the previous version if +the readiness gate fails. (`ingress: host` publishes a raw port instead: +recreate-style deploys with seconds of downtime — see +[docs/supported-workloads.md](docs/supported-workloads.md).) ## Install @@ -81,16 +85,16 @@ standalone. 2. **Starts** a new container alongside the old one 3. **Health checks** the new container 4. **Routes traffic** via Caddy (automatic HTTPS) -5. **Stops** the old container — zero downtime -6. **Rolls back** automatically if anything fails +5. **Stops** the old container — no downtime during the switch under Caddy +6. **Rolls back** to the previous version if the readiness gate fails ## Features | Feature | Description | |---|---| -| **Zero-downtime deploys** | New container starts and passes health checks before old one stops | +| **Zero-downtime deploys** | New container starts and passes health checks before old one stops (Caddy blue/green; `ingress: host` deploys by recreate — brief downtime, documented) | | **Automatic HTTPS** | Caddy provisions and renews TLS certificates | -| **Rollback** | `teploy rollback` reverts to the previous version instantly | +| **Rollback** | `teploy rollback` reverts to the previous version and health-gates it before answering | | **Multi-process** | Run web, worker, and scheduler from the same image | | **Accessories** | Manage Postgres, Redis, etc. alongside your app | | **Environment variables** | `teploy env set KEY=value` — stored securely on server | @@ -631,8 +635,20 @@ the server), and a five-command human-confirmed rebuild runbook. Read ## Requirements -- A server with SSH access (any Linux VPS — Hetzner, DigitalOcean, Linode, etc.) -- That's it. `teploy setup` handles the rest. +- A Linux server with SSH key access (any VPS — Hetzner, DigitalOcean, Linode, etc.) +- That's it for the standard path. `teploy setup` installs Docker, Caddy, and rsync. Already-provisioned Docker hosts work too (an `ingress: host` deploy needs nothing else). + +One app runs one image; multi-image stacks are refused, not approximated — see the full matrix and declared limits in [docs/supported-workloads.md](docs/supported-workloads.md). + +## Docs + +- [First success](docs/first-success.md) — install to verified deploy and rollback, with failure modes and remedies +- [Supported workloads](docs/supported-workloads.md) — what deploys today, what is refused, ingress guarantees, operational limits +- [Failure and recovery](docs/failure-and-recovery.md) — error envelope and exit codes, `teploy doctor`, interrupted deploys and repair debt, DR bundles (`teploy dr`) +- [Migration](docs/migration.md) — Dokploy/Coolify/Compose import: what converts, what refuses, concept mapping +- [CI/CD](docs/ci-deploy.md) — push-to-deploy with Forgejo/GitHub Actions +- [Secrets scanning](docs/secrets-scanning.md) — Gitleaks recipe +- [Resilience](docs/resilience.md) — surviving server loss (topology + runbook) ## Comparison diff --git a/docs/failure-and-recovery.md b/docs/failure-and-recovery.md new file mode 100644 index 0000000..6f1409d --- /dev/null +++ b/docs/failure-and-recovery.md @@ -0,0 +1,160 @@ +# Failure and recovery + +What failure looks like from the outside (exit codes, the structured error +envelope), what `teploy doctor` diagnoses, what happens when a deploy is +interrupted mid-flight, and the disaster-recovery bundle family. The +behavioral contract behind the crash-recovery machinery is +[C01_RECOVERY_STATE_TABLE.md](C01_RECOVERY_STATE_TABLE.md); this page is +the operator-facing version. + +## Exit codes and the error envelope + +Exit codes are stable and minimal: `0` success, `1` command failure. `2` +is reserved as `teploy drift --exit-code`'s CI signal and is never used +by other commands (`doctor` fails with 1, never 2). + +Under `--json`, a failed command additionally writes one JSON document to +**stderr** (stdout keeps only successful output): + +```json +{"machine_interface":2,"code":"config-invalid","message":"invalid teploy configuration","detail":"invalid teploy.yml: 'domain' is required"} +``` + +The closed code taxonomy: + +| Code | Meaning | Wired today | +|---|---|---| +| `config-invalid` | Config (teploy.yml/TOML/destination/Compose) failed to load or validate, or a deploy request was refused before any effect (admission) | Yes — config errors and deploy admission refusals | +| `conflict` | The request is coherent but the world moved under it (plan/apply drift refusal) | Yes | +| `internal` | Everything else, including not-yet-migrated failure sites | Yes (default) | +| `target-unreachable` | SSH/transport failure | Reserved — currently classifies as `internal` | +| `unsupported` | The verb needs a capability this binary lacks | Reserved | +| `uncertain-outcome` | The effect's fate is unknown pending reconciliation | Reserved | +| `degraded` | Success with a flag — traffic switched but the outcome is not clean | Reserved | + +Reserved means the code exists in the schema and consumers must decode it, +but no error site classifies into it yet; treat an unknown code as +`internal` (forward-safe). Machine consumers: stderr is the error channel, +stdout is data, and the exit code stays the pass/fail signal. + +A note on honest outcomes: a deploy whose traffic switched but whose +predecessor retirement partially failed is recorded in `teploy log` as +`DEGRADED` (with the reason), not as a clean `ok` — filter on status, not +just success. + +## teploy doctor + +Read-only diagnostics; a run that fails checks deploys nothing. Checks: + +| Check | What it verifies | +|---|---| +| `git` | Local git present (warn only — prebuilt-image deploys do not need it) | +| `config` | teploy.yml/TOML or the Compose importer accepts the project (grammar errors arrive verbatim) | +| `ssh` | Key-auth connectivity to the app's server (or `--server`) | +| `docker` | Remote Docker reachable | +| `disk` | Root filesystem headroom (fail < 2 GiB free; warn < 10 GiB or > 85% used — deploys write layers, backups and attempt artifacts to `/`) | +| `registry` | The configured `image:` is reachable; auth failures distinguished from unreachable | +| `caddy` | Caddy admin API (Caddy ingress only; skipped-ok for host/external ingress) | +| `compatibility` | Local teploy version vs the server-side teploy binary when one exists (autodeploy installs one at `/deployments/.bin/teploy`) | +| `repair-debt` | Outstanding release-record repair debt (below) | + +`--json` emits `{machine_interface, checks[{name,result,detail,remediation}], summary}`; +result is `ok|warn|fail` and every check always carries all four keys. +Failing checks each print a `fix:` line in human mode. + +## Interrupted deploys + +A deploy is a sequence of durable steps (attempt artifacts, containers, +readiness receipt, traffic switch, fenced state commit, predecessor +retirement, terminal record). The machinery around an owner that dies +mid-sequence: + +- **Fenced per-app locks.** A new owner that takes over a stale lock does + not assume quiescence: it observes the target's real evidence (running + containers by exact name/label, receipts, authority state) and runs the + recovery decision before its first effect. Safe-to-retry is the only + disposition that proceeds automatically; anything ambiguous refuses + with the observed evidence rather than guessing. +- **Dispositions** (what the decision function can order): RETRY + (idempotent tail — record writes, predecessor retirement), INSPECT + (effects landed without receipts — re-observe, never invent success), + COMPENSATE (uncommitted traffic — undo via the recorded predecessor), + MANUAL (unattributable workloads, unreadable authority, traffic on an + uncommitted generation with the predecessor gone). MANUAL means the + operator reconciles by hand; the refusal message names what was seen. +- **Repair debt.** If the post-commit release-record write fails, a + `/deployments//repair-debt.json` marker is persisted; the next + deploy repairs the record from live containers before its own work, + then clears the marker. `teploy status` and `teploy doctor` report + outstanding debt. +- **Predecessor snapshot persistence.** The predecessor container + identities are journaled with the attempt + (`meta/att/./predecessors.json`) before any new container + starts, so compensation after a crash targets exactly the recorded + containers rather than an inference. + +What this means operationally: after a crashed or canceled deploy, the +command to run first is `teploy doctor`, then `teploy status`. If the +deploy refused with a recovery-evidence error, that is the MANUAL +disposition — read the named evidence, `teploy rollback` if a predecessor +is recorded, and only then redeploy. Never respond to an uncertain deploy +by blind re-deploying; the CLI will refuse where it cannot attribute, and +that refusal is the safety working. + +Fleet note: a multi-server rollout is a sequence of per-host recorded +outcomes, not a global transaction — see the staged-rollout section of the +README for canary and failure-budget behavior. + +## Self-heal (steady state) + +`teploy heal enable` installs a systemd-timer probe that restarts an +unhealthy **web** container in place (bounded attempts/backoff) — for +"container up but failing", not for deploys. `teploy heal status` / +`teploy heal disable` manage it. + +## DR bundles (teploy dr) + +`teploy backup` is the data-only family (volume archives, single +accessories). `teploy dr` bundles the whole application: state, release +records, the applied manifest, secret references (or explicitly opted-in +encrypted material), routing identity, and consistency-labeled data +snapshots. + +```bash +# create (S3 or a plain directory on the server) +teploy dr create --bucket my-bucket # or --dir /srv/dr-bundles +teploy dr create --include-secrets # opt in: age ciphertexts, resolved .env, + # accessory credentials (never default) +teploy dr create --include-age-key # + the key itself, so a fresh host can + # decrypt (requires --include-secrets) +teploy dr create --stop-app # quiesced volume snapshots (app restarted after) + +teploy dr list --dir ... # bundle ids +teploy dr show --dir ... # manifest: snapshots + consistency, + # secrets mode, routing, recovery plan + +teploy dr restore --dir ... # isolated restore into /var/tmp/teploy-dr, + # boots scratch engines + app container, + # validates, writes an RPO/RTO receipt. + # Nothing under /deployments is touched. +teploy dr cutover # the explicit mutation: promote a + # validated staged restore over the live app + # (originals kept for two-phase recovery) +``` + +Verified against a scratch host (this branch): `create --dir`, `list`, +`show`, and `restore` — the restore receipt reported staging path, RPO/RTO, +per-check results (`app ... pass image=... running`), and the exact next +command (`teploy dr cutover `). A restore that fails validation +exits non-zero with "nothing live was touched"; missing secret keys fail +before any mutation. After a cutover, run `teploy deploy` to bring the +app container and routing live from the restored state. + +Snapshot consistency is labeled per snapshot: engine dumps are +engine-consistent; raw volume copies are crash-consistent unless the app +was stopped (`--stop-app`) or you assert a volume quiesced by hand +(`--quiesced-volume NAME`). The labels are recorded in the manifest — +read them before trusting a volume snapshot of a writing database. + +For the topology-level version (N+1, state off-box, the dead-server +runbook), see [resilience.md](resilience.md). diff --git a/docs/first-success.md b/docs/first-success.md new file mode 100644 index 0000000..63b7c4e --- /dev/null +++ b/docs/first-success.md @@ -0,0 +1,146 @@ +# First success: install to rollback + +The shortest honest path from nothing to a deployed app you have verified +and can roll back. Every command below was executed against a scratch +Docker host over SSH (a colima VM); where a step needs an environment this +walk-through cannot assume (a public domain, a cloud VPS), it says so. + +## 0. Install the CLI + +```bash +brew install useteploy/tap/teploy # macOS/Linux +# or download a release binary / Scoop on Windows / go install — see README +teploy version +``` + +Verified here with a worktree build (`teploy version` prints the embedded +version; a source build prints `dev`). + +## 1. Have a server + +Any Linux server you can SSH into with key auth, with Docker installable. +`teploy setup ` provisions it: Docker, Caddy, firewall, and by +default host audit hardening (auditd + sudo session recording; skip with +`--no-harden`). + +```bash +teploy setup 203.0.113.10 +``` + +Gated here: `setup` against a fresh public VPS was not re-run for this +walk-through (no spare public host); the command's behavior is covered by +its own tests, and the fixture below used an already-provisioned Docker +host. If your host already runs Docker, a deploy works without `setup` +for `ingress: host` — Caddy is only required for domain routing. + +Register the server so commands can name it (writes +`~/.teploy/servers.yml`): + +```bash +teploy server add box1 203.0.113.10 --user root +``` + +Non-standard SSH port? Use `host:port` — `teploy server add box1 +203.0.113.10:2222 --user root` and `server: box1` in `teploy.yml`. + +## 2. Create the app + +A project directory with a `Dockerfile` listening on one port: + +```dockerfile +FROM nginx:1.27-alpine +COPY index.html /usr/share/nginx/html/index.html +``` + +Either run `teploy init` (interactive: app name, domain or raw port, +server) or write the three-line config yourself: + +```yaml +# teploy.yml +app: demo +domain: demo.example.com # or: ingress: host + port: 8080 for a raw port +server: box1 +``` + +Check yourself before deploying: + +```bash +teploy validate # config grammar + server reference +teploy doctor # read-only: git, config, SSH, Docker, disk, registry, + # Caddy, version compatibility, repair debt +``` + +`doctor` never mutates server state and exits 1 (never 2) when any check +fails; every failing check prints a `fix:` line. A `warn` (e.g. missing +local git) does not fail the run. + +## 3. Deploy + +```bash +teploy deploy +``` + +What you should see (abridged, from the verified run): + +``` +Built image: demo-build- +Deploying demo (version )... +Publishing on 0.0.0.0:80 (host ingress)... +Starting container demo-web- (port 80)... + Readiness: auto — HTTP then TCP fallback (compat, 30s deadline) + Health check passed +Deployed demo version in 1.237s +Receipt: image sha256:..., revision , config manifest ... +``` + +The deploy output names the target, the version, the readiness mode and +deadline, and ends with a receipt (image digest, revision, config +manifest) — the identity of exactly what landed. + +Failure modes at this step, with the product's remedy: + +| Symptom | Meaning | Remedy | +|---|---|---| +| `could not determine version from git ... (use --version flag)` | Project is not a git repository | `git init && git commit`, or `teploy deploy --version v1` | +| `rsync failed: exit status 127` | Server lacks `rsync` | Install it (`apt-get install rsync`); `teploy setup` includes it | +| `authentication failed for root@...; try --user ...` | SSH user/key mismatch | `--user`, `--key`, or `TEPLOY_USER`/`TEPLOY_SSH_KEY` | +| `'domain' is required` | No domain and no `ingress: host` | Set `domain:`, or `ingress: host` + `port:` | +| Health check never passes | Readiness gate refuses to switch traffic | The deploy fails and the previous version keeps serving. Fix the app's health endpoint or set `health: {mode: tcp}` — see README "Config" | +| Doctor fails on `ssh`/`docker`/`disk` | Target not ready | Follow the per-check `fix:` line; disk fails below 2 GiB free | + +## 4. Verify + +```bash +teploy status # containers + version +teploy health # run the readiness probe on the live app +teploy logs --tail 20 # stream logs (Ctrl-C to exit — it follows) +teploy log # deploy history: deploys, rollbacks, failures +``` + +And from any machine that can reach the app: `curl http(s):///`. + +## 5. Change something, deploy again, roll back + +Commit a change and deploy; `teploy status` now shows current and previous +hashes. To revert: + +```bash +teploy rollback +``` + +On Caddy ingress this switches traffic back to the still-known predecessor +version and health-gates it. On `ingress: host` the prior container was +removed at deploy, so rollback redeploys the previous version and +health-gates it (seconds, not an instant switch) — verified on the +fixture: `Rolled back demo to version in 369ms`. + +## 6. Next steps + +- Secrets and env: `teploy env set`, `teploy secret set` (encrypted at + rest), SOPS/age `env_files:`. +- Backups: `teploy backup create --bucket ...`, verified restores via + `teploy accessory verify-backup`. +- Whole-app disaster recovery bundles: [failure-and-recovery.md](failure-and-recovery.md). +- CI: [ci-deploy.md](ci-deploy.md). +- Fleet and surviving server loss: [resilience.md](resilience.md). diff --git a/docs/migration.md b/docs/migration.md new file mode 100644 index 0000000..5e32b49 --- /dev/null +++ b/docs/migration.md @@ -0,0 +1,96 @@ +# Migrating to teploy (from Dokploy / Coolify / raw Compose) + +Both platforms center on Compose files, and teploy can import that subset +of Compose it can preserve exactly. The contract (programme workstream +C05): **every supplied field is preserved, explicitly translated, or +rejected with a named, actionable error — never silently dropped.** The +classification table below is a summary; the executable authority is the +importer and its conformance tests (`internal/config/compose.go`, +`TestLoadCompose_FieldClassificationInventory` in `compose_test.go`). + +## What a migration looks like + +1. `teploy setup ` on a fresh host (Docker + Caddy + firewall). +2. Put your `docker-compose.yml` in the project directory (or run + `teploy init`, which offers to import it). +3. `teploy validate` — this either imports or refuses, naming every + problem field. +4. Fix refusals (below), set `server:` and `domain:`, deploy. +5. Cut traffic over (DNS or proxy) when the app is verified — both stacks + can run side by side; nothing forces a destructive cutover. + +## What converts + +| Compose | Becomes | +|---|---| +| The single non-accessory service with ports | The app (`web` process); its container port from short-form `"host:container"` (host side deliberately not preserved — teploy allocates host ports and routes via Caddy) | +| `build:` on the web service | Build context for the image | +| `image:` on the web service | `image:` (no build) | +| Same-`build:` second service | A `processes:` worker running the same image with the service's `command:` | +| `postgres/redis/mysql/mariadb/mongo/clickhouse/meilisearch/elasticsearch/memcached/rabbitmq/nats` images | Accessories (known default ports) | +| Any other standalone-image service | An accessory | +| `environment:` (map or list) | `env:` verbatim (deploy-time `${VAR}` expansion applies) | +| `volumes:` on web or services | Named volumes (managed under `/deployments//volumes/`) or host binds (source starting `/`) | +| Web `healthcheck.test` exec-list HTTP probe | `health:` path/interval | +| `healthcheck: {disable: true}` / `test: ["NONE"]` | `healthcheck..disable: true` | +| `restart: always` / `unless-stopped` | Tolerated (teploy's own policies match) | +| `deploy:` with only `replicas: 1` / `mode: replicated` | Tolerated as the no-op default | +| `networks: [default]` (or absent) | Tolerated as the implicit default | +| Services under non-default `profiles:` | Skipped, deliberately — `docker compose up` without `--profile` would not deploy them either | +| `depends_on` | Parsed, not translated: teploy already starts every accessory before any app container. Readiness conditions (`service_healthy`) are **not** waited for | + +## What refuses (named errors, before any effect) + +- A second service with a **different build context**: "unsupported + independent build in compose import: `` (build `""`) while + `"web"` builds from ... — teploy runs one image per app and cannot + preserve a separately built service; use the same build context as the + app, a prebuilt image, or write teploy.yml". +- **Multiple port-publishing non-accessory services**: "ambiguous compose + import: multiple non-accessory services publish ports (...)". +- Per-field refusals, one named error each: non-default `networks:`, + `env_file`, `secrets`, `configs`, `extends`, non-default `deploy:`, + `container_name`, `hostname`, `working_dir`, `entrypoint`, + `privileged: true`, non-empty `cap_add`, other `restart:` policies, + long-form/ranged/multi-port publishes, non-web healthchecks without a + home. Each error names the field, why it cannot be preserved, and the + alternative ("remove it or write teploy.yml"). + +What multi-image stacks should do instead: model each independently built +service as its own teploy app on the same server (they share the teploy +network and can address each other by app name), or prebuild images and +run them as accessories. + +## Concept mapping from Dokploy/Coolify + +| There | Here | +|---|---| +| Project / Application | One directory with `teploy.yml` (or an imported compose file) | +| Environment variables UI | `teploy env set KEY=value` (stored server-side) | +| Secrets | `teploy secret set` (age-encrypted at rest) or SOPS/age `env_files:` | +| Traefik + Let's Encrypt | Caddy, written and reloaded by teploy (ACME default; custom `tls:` supported) | +| Domains / routes | `domain:` per app; preview subdomains via `teploy preview` | +| Databases | `accessories:` (managed containers with `--restart always`) | +| Webhook auto-deploy | `teploy autodeploy setup` (+ `autodeploy.paths:` filters for monorepos) | +| Dashboard | [teploy-dash](https://github.com/useteploy/teploy-dash) (optional, read-state + delegate; the CLI stays the source of truth) | +| Backups | `teploy backup` (data-only) and `teploy dr` (whole-app bundles) | + +## Reversible adoption + +Nothing in a teploy migration touches the origin platform: state lives in +`/deployments//` on the host you point at, containers are plain +Docker containers with `teploy.*` labels, and +`teploy remove` retires an app's containers, proxy route, and deploy +state when you want it gone. Rolling back to Dokploy/Coolify means +pointing DNS at the old deployment — run both in parallel until the new +one has proven itself, then decommission the old. + +## After importing + +- `teploy doctor` — the `config` check re-runs the importer; `disk`, + `docker`, `caddy` validate the target. +- `teploy plan` — read-only preview of the first deploy's container and + routing changes. +- Data: migrate database contents with a dump/restore into the new + accessory (`teploy accessory backup/restore` on the source platform's + volume export), then cut over. diff --git a/docs/supported-workloads.md b/docs/supported-workloads.md new file mode 100644 index 0000000..d3bb1c0 --- /dev/null +++ b/docs/supported-workloads.md @@ -0,0 +1,61 @@ +# Supported workloads + +What `teploy deploy` accepts today, and what it refuses — with the refusal +behavior named. If something you need is in the refused column, the answer +is "not yet", not "silently degraded": every refusal below is a named, +actionable error issued before any container effect unless stated +otherwise. + +## Matrix + +| Workload | Status | Notes | +|---|---|---| +| Single-image container app (Dockerfile) | Supported | Default path. Build on the server (`rsync` context + `docker build`), or `build_local: true`. Nixpacks used when no Dockerfile is present (requires Nixpacks on the server or locally). | +| Single-image app from a registry (`image:`) | Supported | Private registries via `teploy registry login`. | +| Multi-process from one image (`processes:` web/worker/cron) | Supported | All processes run from the same image; `healthcheck:` per-process overrides for inherited probes. | +| Static site (`type: static`) | Supported | rsync to the server, served by the managed Caddy. Requires Caddy ingress; `ingress: host`/`external` and `tls:` are rejected for static. | +| Compose file as an importer (subset) | Supported (subset) | `docker-compose.yml` in the project dir imports when no `teploy.yml` exists. Every supplied field is preserved, translated, or rejected with a named error — see [migration.md](migration.md) for the classification summary. | +| Templates (`teploy template install`) | Supported | One-command deploys of reviewed community apps (Postgres+Adminer, WordPress, Immich, ...). Catalog: `teploy template list`. | +| Preview environments (`teploy preview`) | Supported | Branch slugs on `preview-.` against a pre-built image (`teploy build`). Requires Teploy-managed Caddy. | +| Accessories (Postgres, Redis, MySQL, Mariaadb, Mongo, ClickHouse, Meilisearch, Elasticsearch, Memcached, RabbitMQ, NATS, or any standalone image) | Supported | Managed alongside the app with `--restart always`, volumes, ports, env. | +| Multi-image stacks (several independently built services) | Refused | One image per app is the model. A Compose file whose service builds from a different context than the web service refuses at config load: `unsupported independent build in compose import: ... — teploy runs one image per app and cannot preserve a separately built service`. Model as separate teploy apps, or prebuilt images. | +| Multiple web candidates in one Compose file | Refused | `ambiguous compose import: multiple non-accessory services publish ports (...)`. Remove ports from non-app services or write `teploy.yml`. | +| Host port ranges / long-form Compose ports | Refused | Named error from the port grammar; short-form `"host:container"` and non-TCP publishes are what convert. | +| Compose `networks:` (non-default), `env_file`, `secrets`, `configs`, `extends`, `deploy:` (non-default), `container_name`, `hostname`, `working_dir`, `entrypoint`, `privileged`, `cap_add`, `restart:` (other than always/unless-stopped) | Refused per field | One named, actionable error per supplied field whose meaning would be lost (`unsupported compose fields: ...`). Full table: [migration.md](migration.md). | +| Kubernetes-style scheduling, cross-host replica scheduling | Not supported, by design | No scheduler. Fleet semantics are per-host deploys + Caddy LB. See [resilience.md](resilience.md). | + +## Ingress modes and their guarantees + +| Mode | Deploy strategy | Downtime | Rollback | +|---|---|---|---| +| `ingress: caddy` (default) | Blue/green: new container starts, passes the readiness gate, then traffic switches; predecessor stops | None during the switch (drain via `drain_seconds`) | `teploy rollback` switches back to the predecessor version | +| `ingress: external` | Same container lifecycle; Teploy never touches the proxy (the container joins the teploy network with its app-name alias) | Whatever your proxy's cutover does | `teploy rollback` (container-level) | +| `ingress: host` | Recreate: stop old, start new, on a fixed host port | Seconds per deploy (a fixed port cannot be blue/green) | `teploy rollback` redeploys the previous version and health-gates it (there is no instant container switch — the prior container was removed) | + +Static apps (`type: static`) are Caddy-served and use release directories, +not containers; `keep_releases` prunes them. + +## Operational limits + +Declared, not aspirational: + +- **One app = one image.** Every process runs from the same image; there is + no per-service build. This is the boundary the Compose importer enforces. +- **Single-writer deploys per app per host.** Deploys serialize behind a + fenced per-app lock; a crashed holder's effects are reconciled by the next + owner (see [failure-and-recovery.md](failure-and-recovery.md)). +- **Fixed host ports are single-replica.** `ingress: host` and `publish:` + entries cannot be load-balanced across containers on one host; + replicas require Caddy ingress. +- **Readiness is a gate, not a liveness system.** `health:` defines what + "healthy" means before traffic switches; steady-state restart-in-place is + `teploy heal enable` (bounded, systemd-timer driven). +- **The server needs `rsync` on PATH** for source sync (present on normal + Debian/Ubuntu images; `teploy setup` installs it). Missing `rsync` fails + the sync step with `rsync failed: exit status 127` before anything lands. +- **No scheduler.** Surviving server loss is a topology concern — + [resilience.md](resilience.md) is the supported pattern and runbook. +- **Version identity comes from git** (short hash) or `--version`. A project + that is not a git repository must deploy with `--version` — the failure + message says so: `could not determine version from git ... (use --version + flag)`. From acf475f922073181ece58d3bf7283f657225768e Mon Sep 17 00:00:00 2001 From: Tyler <53561637+im-tyler@users.noreply.github.com> Date: Wed, 23 Sep 2026 18:22:44 -0700 Subject: [PATCH 04/11] feat(examples): executable quickstart against a local docker target (C09) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 'make quickstart' (examples/quickstart/run.sh) closes the C09 acceptance gap that stayed open after doctor/MI/exit-code work: the new-user path was documented prose with nothing executable behind it. The maintained fixture (examples/quickstart/app: busybox httpd + teploy.yml, host ingress :18080, tcp health) deploys to the LOCAL colima VM's own SSH endpoint — no remote server, domain, or DNS needed, so the loop runs on a laptop through the real deploy machinery (config loader, build-on-target, health gate, commit, status, remove). Verified green on this machine (colima/aarch64): deploy qs1 -> curl serves v1 -> redeploy qs2 -> curl serves v2 -> status shows both releases -> full cleanup (containers, images, /deployments/quickstart, the run's own known_hosts lines). Honest gating: missing docker/colima/ ssh/curl or a stopped VM prints SKIP and exits 0. One fixture detail worth recording: the app listens on the PORT env teploy injects (httpd ${PORT:-80}) — a fixed-port listener silently breaks host ingress because the publish and the health probe point at the injected port, while docker-proxy still accepts TCP dials (a tcp health gate can pass against a dead backend). --- Makefile | 21 ++++- examples/quickstart/README.md | 32 ++++++++ examples/quickstart/app/Dockerfile | 5 ++ examples/quickstart/app/index.html | 3 + examples/quickstart/app/teploy.yml | 9 +++ examples/quickstart/run.sh | 125 +++++++++++++++++++++++++++++ 6 files changed, 194 insertions(+), 1 deletion(-) create mode 100644 examples/quickstart/README.md create mode 100644 examples/quickstart/app/Dockerfile create mode 100644 examples/quickstart/app/index.html create mode 100644 examples/quickstart/app/teploy.yml create mode 100755 examples/quickstart/run.sh diff --git a/Makefile b/Makefile index 0fe770a..1bbc269 100644 --- a/Makefile +++ b/Makefile @@ -1,4 +1,4 @@ -.PHONY: build test lint vet clean +.PHONY: build test lint vet clean quickstart release-verify release-record release-smoke build: go build -o teploy ./cmd/teploy @@ -14,3 +14,22 @@ vet: clean: rm -f teploy + +# Executable quickstart (C09): deploy the maintained fixture app to the +# local colima VM and verify it answers. Skips honestly (exit 0) when no +# local docker target exists. +quickstart: + ./examples/quickstart/run.sh + +# R01 release receipts: build the goreleaser matrix locally, checksum, +# and diff against recorded expectations (see release/RELEASE_RECEIPT.md). +release-verify: + ./scripts/release-verify.sh verify + +release-record: + ./scripts/release-verify.sh record + +# R01 built-image smoke: build the container image from the verified +# matrix binary and run version + doctor in it. Skips without docker. +release-smoke: + ./scripts/release-verify.sh smoke diff --git a/examples/quickstart/README.md b/examples/quickstart/README.md new file mode 100644 index 0000000..7a624c0 --- /dev/null +++ b/examples/quickstart/README.md @@ -0,0 +1,32 @@ +# teploy quickstart (executable) + +The C09 acceptance line: *a new user deploys the maintained fixture from +documentation without undocumented repair.* This directory is that +documentation, and it runs. + +``` +make quickstart # from the repo root +``` + +`run.sh` deploys `app/` (the maintained fixture: a busybox httpd app with +a `teploy.yml`) to the **local colima VM's own SSH endpoint**, then: + +1. builds the CLI from this checkout, +2. bootstraps the target once (`/deployments` directory; the only + target-side setup, via the VM's passwordless sudo), +3. deploys version `qs1` — build-on-target, health-gated start, host + ingress on `127.0.0.1:18080`, +4. verifies the app answers with the v1 content, +5. redeploys as `qs2` with changed content and verifies the switch, +6. shows `teploy status`, then removes everything it created (app, + containers, images, its own known_hosts lines). + +Requirements: docker CLI, a **running** colima VM (the script never +starts one — `colima start` yourself), ssh/ssh-keyscan/curl/python3. +When any is missing the script prints `SKIP: ...` and exits 0 — a +skipped quickstart is not a failed one. + +No remote server, no domain, no DNS: host ingress publishes a plain port +on the target, which is exactly what makes the loop runnable on a +laptop. The deploy path exercised is the real one — same config loader, +same build/health/commit machinery as production. diff --git a/examples/quickstart/app/Dockerfile b/examples/quickstart/app/Dockerfile new file mode 100644 index 0000000..53ee879 --- /dev/null +++ b/examples/quickstart/app/Dockerfile @@ -0,0 +1,5 @@ +FROM busybox:1.37 +COPY index.html /www/index.html +# teploy injects PORT (the published port) as an env var; listen there so +# the health gate and the published port see the same listener. +CMD ["sh", "-c", "httpd -f -p ${PORT:-80} -h /www"] diff --git a/examples/quickstart/app/index.html b/examples/quickstart/app/index.html new file mode 100644 index 0000000..ea153cc --- /dev/null +++ b/examples/quickstart/app/index.html @@ -0,0 +1,3 @@ + +teploy quickstart +

teploy quickstart v1

diff --git a/examples/quickstart/app/teploy.yml b/examples/quickstart/app/teploy.yml new file mode 100644 index 0000000..107da89 --- /dev/null +++ b/examples/quickstart/app/teploy.yml @@ -0,0 +1,9 @@ +# The maintained quickstart fixture app (C09). Deployed by +# examples/quickstart/run.sh against a local docker target; everything +# teploy needs travels in this directory — no undocumented repair. +app: quickstart +server: colima-vm +ingress: host +port: 18080 +health: + mode: tcp diff --git a/examples/quickstart/run.sh b/examples/quickstart/run.sh new file mode 100755 index 0000000..3c52cb1 --- /dev/null +++ b/examples/quickstart/run.sh @@ -0,0 +1,125 @@ +#!/usr/bin/env bash +# Executable quickstart (C09): deploy the maintained fixture app from +# examples/quickstart/app against a LOCAL docker target — the colima VM's +# own SSH endpoint — with no undocumented repair, then verify the app +# answers, redeploy a second version, and clean up after itself. +# +# Honest gating: this script needs (a) the docker CLI, (b) a running +# colima VM (it will NOT start one), and (c) ssh/ssh-keyscan/curl on +# PATH. When any is missing it prints SKIP and exits 0 — a skipped +# quickstart must not read as a failed one. +# +# What it proves: a new user path from `git clean` checkout to a +# responding application — config in teploy.yml, build on the target, +# health-gated start, published port, idempotent redeploy, status, +# removal. Exit 0 only if every step held. +set -euo pipefail + +REPO_ROOT="$(cd "$(dirname "$0")/../.." && pwd)" +APP_DIR="$REPO_ROOT/examples/quickstart/app" +WORK_DIR="$(mktemp -d "${TMPDIR:-/tmp}/teploy-quickstart.XXXXXX")" +TEPLOY_BIN="" +SSH_HOST=""; SSH_PORT=""; SSH_USER=""; SSH_KEY="" +CONTAINER_KEY_FILE="" # known_hosts lines added by this run +APP_PORT=18080 + +log() { printf '==> %s\n' "$*"; } +skip() { printf 'SKIP: %s\n' "$*"; exit 0; } + +cleanup() { + local code=$? + set +e + if [ -n "$TEPLOY_BIN" ] && [ -n "$SSH_HOST" ]; then + "$TEPLOY_BIN" remove --purge --yes --app quickstart --host "$SSH_HOST" \ + --user "$SSH_USER" --key "$SSH_KEY" >/dev/null 2>&1 + ssh -i "$SSH_KEY" -p "$SSH_PORT" -o BatchMode=yes "$SSH_USER@$SSH_HOST" \ + 'docker rm -f quickstart-web-qs1 quickstart-web-qs2 >/dev/null 2>&1; docker rmi quickstart-build-qs1 quickstart-build-qs2 >/dev/null 2>&1; sudo rm -rf /deployments/quickstart' 2>/dev/null + fi + if [ -n "$CONTAINER_KEY_FILE" ] && [ -f "$HOME/.ssh/known_hosts" ]; then + # Remove only the lines this run appended. + python3 - "$CONTAINER_KEY_FILE" <<'PY' +import sys +added = set(open(sys.argv[1]).read().splitlines()) +path = __import__("os").path.expanduser("~/.ssh/known_hosts") +lines = open(path).read().splitlines() +kept = [l for l in lines if l not in added] +open(path, "w").write("\n".join(kept) + ("\n" if kept else "")) +PY + fi + rm -rf "$WORK_DIR" + exit $code +} +trap cleanup EXIT + +# --- gates ----------------------------------------------------------------- +for bin in docker ssh ssh-keyscan ssh-keygen curl python3; do + command -v "$bin" >/dev/null 2>&1 || skip "$bin not found on PATH" +done +command -v colima >/dev/null 2>&1 || skip "colima not found (this quickstart targets a local colima VM)" +colima status >/dev/null 2>&1 || skip "colima VM not running (start it with: colima start), then re-run" + +# --- resolve the colima VM's SSH endpoint from colima's own config --------- +COLIMA_HOME_DIR="${COLIMA_HOME:-$HOME/.colima}" +SSH_CFG="" +for candidate in "$COLIMA_HOME_DIR/ssh_config" "$COLIMA_HOME_DIR/default/ssh_config" "$COLIMA_HOME_DIR/_lima/colima/ssh_config"; do + [ -f "$candidate" ] && SSH_CFG="$candidate" && break +done +[ -n "$SSH_CFG" ] || skip "no colima ssh_config under $COLIMA_HOME_DIR" +SSH_HOST="$(awk '$1=="Hostname"{print $2}' "$SSH_CFG")" +SSH_PORT="$(awk '$1=="Port"{print $2}' "$SSH_CFG")" +SSH_USER="$(awk '$1=="User"{print $2}' "$SSH_CFG")" +SSH_KEY="$(awk '$1=="IdentityFile"{print $2}' "$SSH_CFG" | tr -d '"')" +for v in "$SSH_HOST" "$SSH_PORT" "$SSH_USER" "$SSH_KEY"; do + [ -n "$v" ] || skip "could not parse endpoint from $SSH_CFG" +done +[ -f "$SSH_KEY" ] || skip "colima SSH key missing ($SSH_KEY) — restart colima to regenerate" +log "target: $SSH_USER@$SSH_HOST:$SSH_PORT (colima VM)" + +ssh -i "$SSH_KEY" -p "$SSH_PORT" -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null \ + -o BatchMode=yes "$SSH_USER@$SSH_HOST" true 2>/dev/null \ + || skip "cannot SSH to the colima VM ($SSH_USER@$SSH_HOST:$SSH_PORT)" + +# --- host key + /deployments bootstrap (the only target-side setup) -------- +mkdir -p "$HOME/.ssh" +touch "$HOME/.ssh/known_hosts" +CONTAINER_KEY_FILE="$WORK_DIR/added-host-keys" +ssh-keyscan -p "$SSH_PORT" "$SSH_HOST" >"$CONTAINER_KEY_FILE" 2>/dev/null +grep -q . "$CONTAINER_KEY_FILE" || skip "ssh-keyscan produced no keys for $SSH_HOST:$SSH_PORT" +cat "$CONTAINER_KEY_FILE" >>"$HOME/.ssh/known_hosts" + +ssh -i "$SSH_KEY" -p "$SSH_PORT" -o BatchMode=yes "$SSH_USER@$SSH_HOST" \ + 'sudo mkdir -p /deployments && sudo chown "$(id -un)" /deployments' 2>/dev/null \ + || skip "cannot bootstrap /deployments on the VM (needs passwordless sudo)" + +# --- build this checkout's CLI ---------------------------------------------- +log "building teploy from this checkout" +TEPLOY_BIN="$WORK_DIR/teploy" +(cd "$REPO_ROOT" && go build -o "$TEPLOY_BIN" ./cmd/teploy) + +TEPLOY=("$TEPLOY_BIN" --host "$SSH_HOST:$SSH_PORT" --user "$SSH_USER" --key "$SSH_KEY") + +# --- deploy v1 -------------------------------------------------------------- +cp -R "$APP_DIR/." "$WORK_DIR/app/" +log "deploying quickstart v1 (build-on-target, health-gated, host ingress :$APP_PORT)" +(cd "$WORK_DIR/app" && "${TEPLOY[@]}" deploy --version qs1) 2>&1 | sed 's/^/ /' + +BODY="$(curl -fsS -m 10 "http://127.0.0.1:$APP_PORT/")" +grep -q "quickstart v1" <<<"$BODY" || { echo "FAIL: v1 content not served: $BODY" >&2; exit 1; } +log "verified: http://127.0.0.1:$APP_PORT/ serves the v1 fixture" + +# --- redeploy v2 (recreate path) -------------------------------------------- +log "redeploying as qs2 with changed content" +sed 's/quickstart v1/quickstart v2/' "$WORK_DIR/app/index.html" >"$WORK_DIR/app/index.html.tmp" +mv "$WORK_DIR/app/index.html.tmp" "$WORK_DIR/app/index.html" +(cd "$WORK_DIR/app" && "${TEPLOY[@]}" deploy --version qs2) 2>&1 | sed 's/^/ /' + +BODY="$(curl -fsS -m 10 "http://127.0.0.1:$APP_PORT/")" +grep -q "quickstart v2" <<<"$BODY" || { echo "FAIL: v2 content not served after redeploy: $BODY" >&2; exit 1; } +log "verified: redeploy switched the served content to v2" + +# --- status ------------------------------------------------------------------ +log "teploy status sees the deployment" +"${TEPLOY[@]}" status --app quickstart 2>&1 | sed 's/^/ /' + +log "quickstart complete: deployed, verified, redeployed, verified again" +log "cleanup follows (teploy remove --purge, known_hosts lines, temp dir)" From 11d0709fbe19b642739809c614d1ff64a0812e41 Mon Sep 17 00:00:00 2001 From: Tyler <53561637+im-tyler@users.noreply.github.com> Date: Wed, 23 Sep 2026 18:26:46 -0700 Subject: [PATCH 05/11] =?UTF-8?q?feat(release):=20reproducibility=20receip?= =?UTF-8?q?ts=20=E2=80=94=20verify/record/smoke=20(R01=20CLI=20slice)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - scripts/release-verify.sh: builds the exact goreleaser matrix locally (CGO_ENABLED=0, -s -w -X main.version, linux/darwin/windows x amd64/arm64), checksums the raw binaries (the reproducible unit — goreleaser's published checksums.txt covers archives, which embed mtimes), records inputs (go toolchain, commit, describe, ldflags, date/host), and diffs rebuilds against the recorded expectation. Honest refusal over vacuous pass: toolchain skew or changed build inputs (cmd/ internal/ go.mod go.sum) exit 1 naming the difference; a source-identical descendant of the recorded commit compares with a note (the expectation file itself necessarily lands in a later commit). Expectations live in release/expectations/, keyed by (version label, go version), committed and reviewed like contracts. - Dockerfile: scratch smoke vehicle COPY-ing the verified matrix binary — image provenance is always the checksummed build. Explicitly NOT a published artifact (goreleaser ships binaries/archives). - release-smoke: builds linux/ image, runs 'version' (must exit 0) and 'doctor --json' (must emit the 9-check machine envelope; exit 1 = documented failing-diagnosis semantics in a bare container, not a crash). Skips (exit 0) without docker. - release/RELEASE_RECEIPT.md: the pattern (template) + the lane's filled local-verification example, clearly marked NOT A RELEASE. - Make targets: release-verify, release-record, release-smoke. No tag, no publish — release cuts stay owner-controlled. --- .gitignore | 3 + Dockerfile | 13 +++ release/RELEASE_RECEIPT.md | 79 +++++++++++++++ scripts/release-verify.sh | 190 +++++++++++++++++++++++++++++++++++++ 4 files changed, 285 insertions(+) create mode 100644 Dockerfile create mode 100644 release/RELEASE_RECEIPT.md create mode 100755 scripts/release-verify.sh diff --git a/.gitignore b/.gitignore index d9682c7..cca8ff8 100644 --- a/.gitignore +++ b/.gitignore @@ -2,6 +2,9 @@ /teploy /cmd/teploy/teploy +# Release-verify local build matrix (scripts/release-verify.sh) +/dist-verify/ + # Go *.exe *.test diff --git a/Dockerfile b/Dockerfile new file mode 100644 index 0000000..eabb7c3 --- /dev/null +++ b/Dockerfile @@ -0,0 +1,13 @@ +# Smoke-test vehicle for the R01 release-verify built-image check +# (scripts/release-verify.sh smoke). NOT a published release artifact — +# goreleaser ships binaries/archives; this image exists so the exact +# verified binary proves it runs in a minimal (scratch) container. +# The binary is COPY'd from the release-verify dist directory, so the +# image provenance is always the checksummed matrix build. +# +# Build (as the smoke script does): +# docker build --build-arg BINARY=dist-verify/teploy_linux_ -t teploy:release-smoke . +FROM scratch +ARG BINARY=dist-verify/teploy_linux_amd64 +COPY ${BINARY} /teploy +ENTRYPOINT ["/teploy"] diff --git a/release/RELEASE_RECEIPT.md b/release/RELEASE_RECEIPT.md new file mode 100644 index 0000000..a2d47f5 --- /dev/null +++ b/release/RELEASE_RECEIPT.md @@ -0,0 +1,79 @@ +# Release receipts (R01) + +Every release — and every local verification of one — leaves a receipt: +the recorded inputs and the checksums that prove a clean machine can +recreate the build. This file holds the PATTERN (template below) and +receipts from local verification runs. A receipt from an actual release +is added by the owner at tag time; nothing in this directory cuts a +release, tags, or publishes. + +## How to produce a receipt + +``` +make release-record # builds the goreleaser matrix, writes + # release/expectations/checksums--.txt +make release-verify # rebuilds from a clean checkout and diffs against + # the recorded expectation (inputs must match) +make release-smoke # builds the image (Dockerfile, scratch) from the + # verified binary and runs version + doctor in it +``` + +`release-record` prints the receipt block; paste it below under a new +heading. Commit the expectation file with it — the receipt and the +expectation are one artifact split across two files. + +## Template + +```markdown +### — — + +- commit : (git describe: ) +- go toolchain : (go.mod: ) +- build env : CGO_ENABLED=0, GOFLAGS unset +- ldflags : -s -w -X main.version= +- matrix : linux/darwin/windows × amd64/arm64 (goreleaser parity) +- checksums : release/expectations/checksums--.txt +- verify : make release-verify → REPRODUCIBLE (this machine, /) +- image smoke : make release-smoke → PASS (linux/, scratch image; + version exit 0; doctor emits the 9-check machine envelope, + exit 1 = documented failing-diagnosis semantics) +- goreleaser : +``` + +--- + +### 0.0.0-localverify-acf475f — 2026-09-23 — LOCAL VERIFICATION, NOT A RELEASE + +First receipt, from the R01 lane's local run (worktree branch +`c09-x02f-r01-cli`, off `d9b652d`; commit `acf475f` = quickstart commit, +tree carrying uncommitted release/ files at record time — hence the +dirty count below). + +- commit : acf475f922073181ece58d3bf7283f657225768e (git describe: v0.1.37-45-gacf475f) +- go toolchain : go1.26.6 (go.mod: 1.26.0) +- build env : CGO_ENABLED=0, GOFLAGS unset +- ldflags : -s -w -X main.version=0.0.0-localverify-acf475f +- matrix : linux/darwin/windows × amd64/arm64 (goreleaser parity) +- checksums : release/expectations/checksums-0.0.0-localverify-acf475f-go1-26-6.txt + - bf185800ea9b731849cd0a4c61940586bb4fdfc5875636966abe01aad4a34d8e teploy_darwin_amd64 + - 553044b9f5d81fc99fb7ebd2369f0435a41daff3abbfe3f27e0fc192024653f8 teploy_darwin_arm64 + - bacf79c2f18585676e64ebbc06a409304fdb91fb72b9c386ae0b51b4adbcc259 teploy_linux_amd64 + - 33446a81515d1d7c4026b5de04aa56e7216389f6833ac56d13c35a25a41fa208 teploy_linux_arm64 + - 7d0255f55e8f6e74e9c85acff6d57b634b323dd5ea84caeb2ce86da966362fc9 teploy_windows_amd64 + - 13d24237797c3484a40388fa21af4834c1affa3ba16beb2a139960764c23de56 teploy_windows_arm64 +- verify : `make release-verify` → REPRODUCIBLE (all 6 binaries + bit-identical to the recording; Darwin/arm64 host, recorded dirty + files: 3) +- image smoke : `make release-smoke` → PASS — linux/arm64 scratch + image from the verified binary; `teploy version` exit 0; `teploy + doctor --json` exit 1 with the full 9-check machine envelope + (machine_interface 2), which is the documented failing-diagnosis + semantics in a bare container, not a crash +- goreleaser : n/a (no tag, no publish — owner-controlled) + +Scope note: checksums cover the raw matrix binaries, the reproducible +unit; goreleaser's published `checksums.txt` covers archives (which +embed mtimes) and is recorded at release time. Bit-identical +reproduction is toolchain-scoped: the expectation file pins the go +version, and `release-verify` refuses to compare (exit 1, INPUTS +DIFFER) rather than report a meaningless mismatch across toolchains. diff --git a/scripts/release-verify.sh b/scripts/release-verify.sh new file mode 100755 index 0000000..1c7b25a --- /dev/null +++ b/scripts/release-verify.sh @@ -0,0 +1,190 @@ +#!/usr/bin/env bash +# R01 release-reproducibility receipts for teploy-cli. +# +# Subcommands: +# verify build the exact goreleaser matrix locally, checksum every +# binary, and diff against the recorded expectation for this +# (version label, go toolchain). Exit 0 on match; exit 1 on +# checksum mismatch or recorded-input mismatch (a different +# go toolchain cannot reproduce the recorded bits — align or +# re-record); exit 1 with NO RECORDED EXPECTATION when none +# exists yet (run `record`). Honest results, never vacuous. +# record build the same matrix and write the expectation file +# (release/expectations/) plus a receipt block for +# release/RELEASE_RECEIPT.md. Recording is a deliberate act — +# expectations are committed and reviewed like any contract. +# smoke build the linux binary for the local docker architecture, +# build the container image (Dockerfile at repo root), and +# run `version` + `doctor` inside it as the built-image +# smoke. Prints SKIP (exit 0) when docker is unavailable. +# +# The matrix mirrors .goreleaser.yml exactly: CGO_ENABLED=0, +# -s -w -X main.version=