Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,7 +87,7 @@ plugin name. See [Resource plugins](resource-plugins.md).
| `ExecuteWorkload` ⇅ | `workload_id`, `commands`, `envs`, `workdir`, `open_stdin` | Exec inside a workload. When `open_stdin` is set, further client messages carry stdin in `repl_cmd` |
| `RunAndWait` ⇅ | `deploy_options`, `cmd`, `async`, `async_timeout` | Lambda: deploy, attach, wait for exit, then remove. The first messages carry the new workload IDs (`TYPEWORKLOADID`), the last output line is `[exitcode] <n>`. With `async`, core sends the IDs, detaches from the stream, forces `open_stdin` off, and logs the output itself under `async_timeout` seconds (default `global_timeout`) |
| `LogStream` ⇊ | `id`, `tail`, `since`, `until`, `follow` | Engine logs for one workload |
| `RawEngine` | `id`, `op`, `params`, `ignore_lock` | Pass an engine-specific operation through to the node's engine. No real engine implements it; containerd, cocoon and process all return `ErrEngineNotImplemented` (the `mock://` engine answers with a canned result) |
| `RawEngine` | `id`, `op`, `params`, `ignore_lock` | Pass an engine-specific operation through to the node's engine. cocoon serves its snapshot ops (see [engines](engines.md#snapshots)); containerd and process return `ErrEngineNotImplemented` (the `mock://` engine answers with a canned result) |

## Files

Expand Down
43 changes: 39 additions & 4 deletions docs/engines.md
Original file line number Diff line number Diff line change
Expand Up @@ -254,7 +254,7 @@ the cocoon daemon's events on it. The eru name stays in the meta file and in cor

| `engine.API` | cocoon |
| --- | --- |
| `VirtualizationCreate` | `vm create --output json --name <id> [--cpu N] [--memory B] [--storage B] [--data-disk …] [--network <name>] [--windows \| --user U] <image>` — no boot; then the meta record is written. A failure after the create removes the VM again. cocoon has no way to apply the deploy's `env`, `dns`, `extra_hosts` or the entrypoint's `commands` and `dir` to a guest, so any of them set draws a warning; core's own `APP_NAME`/`ERU_*` env does not |
| `VirtualizationCreate` | `vm create --output json --name <id> [--cpu N] [--memory B] [--storage B] [--data-disk …] [--network <name>] [--windows \| --user U] <image>` — no boot; then the meta record is written. A failure after the create removes the VM again. cocoon has no way to apply the deploy's `env`, `dns`, `extra_hosts` or the entrypoint's `commands` and `dir` to a guest, so any of them set draws a warning; core's own `APP_NAME`/`ERU_*` env does not. A `snapshot://<name>` image is a clone instead (see [Snapshots](#snapshots)) |
| `VirtualizationStart` | `vm inspect`, `vm start` and `vm inspect` again in one script; the second inspect reports this boot's `console_path`, and both copies of the meta record are rewritten with it and with the VMM pid. An inspect that reports no console keeps the serial socket path. A Windows guest on its first boot gets its address programmed through `vm exec` in the background, after the start has already returned |
| `VirtualizationStop` | `vm stop`, `--force` for a forced stop, `--timeout` when a grace period is given; a workload with no record on the node is `ErrWorkloadNotExists`, as for the other verbs. cocoon's stop is idempotent, so stopping a created or already-stopped guest succeeds |
| `VirtualizationRemove` | `vm rm [--force]`, then the hibernate snapshot and both copies of the meta record; a running guest is refused unless forced |
Expand All @@ -266,12 +266,13 @@ the cocoon daemon's events on it. The eru name stays in the meta file and in cor
| `Execute` / `ExecExitCode` | `vm exec [-i] [-e K=V …] <id> -- <cmd>` through cocoon-agent in pipe mode, stdio on the SSH session, the exit code the guest command's. `ExecResize` is `ErrEngineNotImplemented` (core#660). A `user` and a `working_dir` are applied inside a Linux guest by wrapping the command — `runuser -u U -- env --chdir=D <cmd>` for a bare user name, `setpriv --reuid=U --regid=G --clear-groups -- env --chdir=D <cmd>` when the id is numeric or a group is named — so the directory is entered as the target user; on a Windows guest both are `ErrEngineNotImplemented` |
| `VirtualizationCopyTo` / `CopyFrom` | a one-entry tar through `vm exec … tar -x -P -f -` / `tar -c -P -f -`: the absolute entry name makes tar create the parents, and `tar.exe` ships with Windows 10+. A copy into a guest that is not running is `ErrInvaildWorkloadOps` — the state is checked first, one round trip per file |
| `VirtualizationUpdateResource` | a remap (the cpumem binding refresh core runs after every deploy) is a no-op without a round trip; a realloc is `ErrEngineNotImplemented`, CPU and memory hot-plug wait on cocoon (core#661) |
| `ImagePull` | `image pull <ref>` for OCI VM images and cloud-image URLs, registry auth left to cocoon's own config; a split-qcow2 artifact (the Windows images) is `oras pull`ed and `image import`ed under the same ref, once |
| `ImagePull` | `image pull <ref>` for OCI VM images and cloud-image URLs, registry auth left to cocoon's own config; a split-qcow2 artifact (the Windows images) is `oras pull`ed and `image import`ed under the same ref, once. A `snapshot://<name>` ref is only looked up with `snapshot inspect`, so a missing snapshot fails the pull |
| `ImageList` / `ImageRemove` | `image list --format json` filtered by name prefix / `image rm`. An empty store answers `No images found.` in prose rather than `[]`, and reads as an empty list, not a failed node |
| `ImageLocalDigests` / `ImageRemoteDigest` | `image inspect` / `oras manifest fetch --descriptor`; a cloud-image URL is its own digest, so it is pulled once. A node without `oras` (probed with `command -v`) reports no remote digest, so every deploy runs `image pull`, which cocoon answers from its cache. Only a node that answered yes is remembered — a probe an ssh failure lost is asked again, instead of pinning the node as oras-less for the engine's life |
| `ImageLocalDigests` / `ImageRemoteDigest` | `image inspect` / `oras manifest fetch --descriptor`; a cloud-image URL is its own digest, so it is pulled once, and so is a `snapshot://<name>` ref while `snapshot inspect` finds the snapshot. A node without `oras` (probed with `command -v`) reports no remote digest, so every deploy runs `image pull`, which cocoon answers from its cache. Only a node that answered yes is remembered — a probe an ssh failure lost is asked again, instead of pinning the node as oras-less for the engine's life |
| `ImageBuildFromExist` | `ErrEngineNotImplemented`, and so is `ImagePush`: cocoon has no registry push, so a build from an existing workload can never finish. Saving a snapshot first only left node state behind — core's build always goes on to push the refs and then runs `ImageRemove` over them — so the engine refuses before anything is written |
| `NetworkList` | the CNI conf dir (`/etc/cni/net.d`); `NetworkConnect` / `Disconnect` are `ErrEngineNotImplemented` |
| `ImageBuild`, `ImagesPrune`, `RawEngine` | `ErrEngineNotImplemented` |
| `RawEngine` | the snapshot ops, see [Snapshots](#snapshots) |
| `ImageBuild`, `ImagesPrune` | `ErrEngineNotImplemented` |

### Resources and networks

Expand Down Expand Up @@ -335,6 +336,40 @@ digest, so the two never match and every deploy runs `ImagePull` again. That is
the import script exits at once when `image inspect <ref>` already answers — but it does mean a
Windows deploy always pays one extra round trip.

### Snapshots

A deploy whose image is `snapshot://<name>` clones the VM from that cocoon snapshot on the node
instead of booting an image:

| `image` | cocoon |
| --- | --- |
| `snapshot://<name>` | `vm clone --output json --name <id> [--network <name>] [--data-disk …] <name>` — the snapshot name or id, checked against cocoon's grammar (`^[a-zA-Z0-9][a-zA-Z0-9._:/-]{0,62}$`) before any round trip |

The clone takes its CPU, memory, storage, guest OS and login from the snapshot, so the deploy's
cpumem and storage quotas are not applied (logged at debug) and `user` reaches only the meta record,
where exec reads it. No network in the deploy keeps the snapshot's conflist. The guest resumes with the source VM's NIC files and hostname, so after the
record the engine rewrites them through `vm exec`: one systemd-networkd file per static NIC, matched
by the clone's new MAC, the workload name as hostname, then `systemctl restart systemd-networkd`,
retried for 30 s until cocoon-agent answers. A clone that will not take its address is removed. The meta record is
written from the clone's JSON exactly as after a create, and a failure after the clone removes the
VM again. The clone is already running when `VirtualizationStart` runs: cocoon answers `vm start`
on a running VM with success, so the start only refreshes the record. The volumes become data disks
hot-added to the running guest; cloud-init has already run, so nothing mounts them, and cocoon
refuses to snapshot or hibernate a VM with a hot-added disk. A Windows deploy (`os: windows`) with a
snapshot image is refused with `ErrInvalidEngineArgs`. The image verbs never reach a registry for
such a ref.

`RawEngine` serves the snapshot ops on the workload's node, `Params` a JSON object, the name checked
as above and refused with `ErrInvalidEngineArgs` before any round trip; any other op is
`ErrEngineNotImplemented`:

| `Op` | `Params` | cocoon | `Data` |
| --- | --- | --- | --- |
| `snapshot.save` | `{"name": N}` | `snapshot save --name N <id>`, then `snapshot inspect N`, in one round trip; the guest keeps running | the snapshot JSON |
| `snapshot.list` | — | `snapshot list --format json`; the prose empty banner reads as `[]` | the JSON list |
| `snapshot.inspect` | `{"name": N}` | `snapshot inspect N` | the snapshot JSON |
| `snapshot.remove` | `{"name": N}` | `snapshot rm N` | empty |

### The meta file

The same record the process engine writes, at `<cocoon.root>/<id>.json` and
Expand Down
4 changes: 0 additions & 4 deletions engine/cocoon/cocoon.go
Original file line number Diff line number Diff line change
Expand Up @@ -92,10 +92,6 @@ func (e *Engine) GetParams() *enginetypes.Params {
return e.ep
}

func (e *Engine) RawEngine(context.Context, *enginetypes.RawEngineOptions) (*enginetypes.RawEngineResult, error) {
return nil, coretypes.ErrEngineNotImplemented
}

func (e *Engine) call(ctx context.Context, argv ...string) (*sshrunner.Result, error) {
return sshrunner.Call(ctx, e.runner, argv...)
}
Expand Down
65 changes: 64 additions & 1 deletion engine/cocoon/create.go
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,28 @@ const (

discardTimeout = 30 * time.Second

// reseedScript rewrites a clone's static NICs by MAC and its hostname; the guest still carries the source's.
reseedScript = `bin=$1; vm=$2; shift 2
guest='rm -f /etc/systemd/network/10-*.network
hostnamectl set-hostname "$1" 2>/dev/null || hostname "$1"
sed -i "/^127\.0\.1\.1[[:space:]]/d" /etc/hosts
printf "127.0.1.1 %s\n" "$(hostname)" >> /etc/hosts
shift
while [ $# -ge 3 ]; do
f="/etc/systemd/network/10-$(printf %s "$1" | tr -d :).network"
printf "[Match]\nMACAddress=%s\n\n[Network]\nAddress=%s\n" "$1" "$2" > "$f"
if [ -n "$3" ]; then printf "Gateway=%s\n" "$3" >> "$f"; fi
shift 3
done
systemctl restart systemd-networkd'
tries=0
until "$bin" vm exec "$vm" -- sh -c "$guest" sh "$@"; do
tries=$((tries+1))
[ "$tries" -lt 30 ] || exit 1
sleep 1
done
`

publishRecord = `mkdir -p "$(dirname "$record")"
cp -f "$durable" "$record.tmp"
mv "$record.tmp" "$record"
Expand Down Expand Up @@ -67,7 +89,13 @@ func (e *Engine) VirtualizationCreate(ctx context.Context, opts *enginetypes.Vir
return nil, err
}
ID := utils.RandomID()
argv, err := createArgv(e.cocoon.Binary, ID, opts, resource, rArgs.OS == osWindows, network)
var argv []string
if snapshot, ok := strings.CutPrefix(opts.Image, snapshotScheme); ok {
logger.Debugf(ctx, "vm %s takes its cpu, memory and storage from snapshot %s", opts.Name, snapshot)
argv, err = cloneArgv(e.cocoon.Binary, ID, snapshot, resource.Volumes, rArgs.OS == osWindows, network)
} else {
argv, err = createArgv(e.cocoon.Binary, ID, opts, resource, rArgs.OS == osWindows, network)
}
if err != nil {
return nil, err
}
Expand All @@ -83,6 +111,9 @@ func (e *Engine) VirtualizationCreate(ctx context.Context, opts *enginetypes.Vir
if err == nil {
err = e.record(ctx, ID, opts, vm)
}
if err == nil && strings.HasPrefix(opts.Image, snapshotScheme) {
_, err = e.run(ctx, reseedArgv(e.cocoon.Binary, ID, opts.Name, vm)...)
}
if err != nil {
e.discard(ctx, ID)
return nil, err
Expand Down Expand Up @@ -139,7 +170,39 @@ func createArgv(binary, ID string, opts *enginetypes.VirtualizationCreateOptions
return append(argv, "--name", ID, opts.Image), nil
}

// cloneArgv boots the vm from a snapshot, which fixes its cpu, memory, storage and guest os.
func cloneArgv(binary, ID, snapshot string, volumes []string, windows bool, network string) ([]string, error) {
if windows {
return nil, errors.Wrap(coretypes.ErrInvalidEngineArgs, "a windows guest cannot be cloned from a snapshot")
}
if err := checkSnapshotName(snapshot); err != nil {
return nil, err
}
argv := []string{binary, "vm", "clone", "--output", formatJSON, "--name", ID}
if network != "" {
argv = append(argv, "--network", network)
}
disks, err := dataDisks(volumes, false)
if err != nil {
return nil, err
}
for _, disk := range disks {
argv = append(argv, "--data-disk", disk)
}
return append(argv, snapshot), nil
}

// dataDisks turns the storage plugin's `src:dst:mode:size` volumes into cocoon data disks.
func reseedArgv(binary, ID, hostname string, vm *vmRecord) []string {
args := []string{binary, ID, hostname}
for _, n := range vm.NICs {
if n.MAC != "" && n.Network != nil && n.Network.IP != "" {
args = append(args, n.MAC, n.Network.IP+"/"+strconv.Itoa(n.Network.Prefix), n.Network.Gateway)
}
}
return sshrunner.Shell(reseedScript, args...)
}

func dataDisks(volumes []string, windows bool) ([]string, error) {
disks := make([]string, 0, len(volumes))
for _, volume := range volumes {
Expand Down
130 changes: 130 additions & 0 deletions engine/cocoon/create_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -121,6 +121,136 @@ func TestVirtualizationCreateKeepsTheConflistOfAnInheritedNetwork(t *testing.T)
}
}

func TestVirtualizationCreateClonesASnapshotImage(t *testing.T) {
runner := &sshrunnertest.Fake{Respond: clonedFrom}
e := testEngine(t, runner)

created, err := e.VirtualizationCreate(t.Context(), &enginetypes.VirtualizationCreateOptions{
Name: "app_web_xyz",
Image: snapshotScheme + testSnap,
User: testUser,
Networks: map[string]string{"eru-cni": ""},
EngineParams: resourcetypes.Resources{"cpumem": {"cpu": 1.5, "memory": 1 << 30}, "storage": {"storage": 20 << 30, "volumes": []string{"/data:/data:rw:1073741824"}}},
})
if err != nil {
t.Fatalf("clone: %v", err)
}
lines := runner.Lines()
if len(lines) != 3 {
t.Fatalf("got %d commands, want the clone, the record and the reseed", len(lines))
}
if reseed := sshrunner.Quote(reseedArgv(testBinary, created.ID, "app_web_xyz", mustParseVM(t, clonedVM))); lines[2] != reseed {
t.Errorf("got %q, want the reseed %q", lines[2], reseed)
}
for _, arg := range []string{"'app_web_xyz'", "'02:00:00:00:00:07'", "'10.22.0.7/16'", "'10.22.0.1'"} {
if !strings.Contains(lines[2], arg) {
t.Errorf("the reseed does not carry %s", arg)
}
}
want := sshrunner.Quote([]string{
testBinary, "vm", "clone", "--output", "json", "--name", created.ID,
"--network", "eru-cni", "--data-disk", "size=1073741824,mount=/data", testSnap,
})
if lines[0] != want {
t.Errorf("got %q, want %q", lines[0], want)
}
for _, field := range []string{
durablePath(testRoot, created.ID),
`"user":"` + testUser + `"`,
`"networks":{"eru-cni":"10.22.0.7"}`,
`"cgroup":"/sys/fs/cgroup/cocoon.slice/vm-` + testVMID + `.scope"`,
`"iface":"tap01ARZ3ND-0"`,
} {
if !strings.Contains(lines[1], field) {
t.Errorf("the record command does not carry %s", field)
}
}
}

func TestVirtualizationCreateDiscardsACloneWhoseRecordFailed(t *testing.T) {
runner := &sshrunnertest.Fake{Respond: func(line string) *sshrunner.Result {
switch {
case strings.Contains(line, "'clone'"):
return &sshrunner.Result{Stdout: clonedVM}
case strings.Contains(line, "'rm'"):
return &sshrunner.Result{}
}
return &sshrunner.Result{Code: 1, Stderr: "read-only file system"}
}}
e := testEngine(t, runner)

if _, err := e.VirtualizationCreate(t.Context(), &enginetypes.VirtualizationCreateOptions{Name: "app_web_xyz", Image: snapshotScheme + testSnap}); err == nil {
t.Fatal("a failed record must fail the clone")
}
lines := runner.Lines()
if len(lines) != 3 || !strings.Contains(lines[2], "'rm' '--force'") {
t.Errorf("got %q, want a forced rm after the failed record", lines)
}
}

func TestVirtualizationCreateDiscardsACloneThatWouldNotReseed(t *testing.T) {
runner := &sshrunnertest.Fake{Respond: func(line string) *sshrunner.Result {
switch {
case strings.Contains(line, "'clone'"):
return &sshrunner.Result{Stdout: clonedVM}
case strings.Contains(line, "systemd-networkd"):
return &sshrunner.Result{Code: 1, Stderr: "agent did not answer"}
}
return &sshrunner.Result{}
}}
e := testEngine(t, runner)

if _, err := e.VirtualizationCreate(t.Context(), &enginetypes.VirtualizationCreateOptions{Name: "app_web_xyz", Image: snapshotScheme + testSnap}); err == nil {
t.Fatal("a clone whose network was not rewritten must fail")
}
lines := runner.Lines()
if len(lines) != 4 || !strings.Contains(lines[3], "'rm' '--force'") {
t.Errorf("got %q, want a forced rm after the failed reseed", lines)
}
}

func TestVirtualizationCreateReportsAMissingSnapshot(t *testing.T) {
runner := &sshrunnertest.Fake{Respond: func(string) *sshrunner.Result {
return &sshrunner.Result{Code: 1, Stderr: "inspect snapshot " + testSnap + ": not found"}
}}
e := testEngine(t, runner)

_, err := e.VirtualizationCreate(t.Context(), &enginetypes.VirtualizationCreateOptions{Name: "app_web_xyz", Image: snapshotScheme + testSnap})
if err == nil || !strings.Contains(err.Error(), "not found") {
t.Fatalf("got %v, want cocoon's lookup failure", err)
}
if lines := runner.Lines(); len(lines) != 1 {
t.Errorf("got %q, want no rm for a vm cocoon never made", lines)
}
}

func TestVirtualizationCreateRefusesACloneBeforeTheNode(t *testing.T) {
tests := []struct {
name string
image string
raw string
}{
{"a windows guest", snapshotScheme + testSnap, `{"os":"windows"}`},
{"an empty snapshot name", snapshotScheme, ""},
{"a name cocoon would refuse", snapshotScheme + "-rf", ""},
{"a name with a space", snapshotScheme + "a b", ""},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
runner := &sshrunnertest.Fake{Respond: clonedFrom}
e := testEngine(t, runner)

_, err := e.VirtualizationCreate(t.Context(), &enginetypes.VirtualizationCreateOptions{Name: "app_web_xyz", Image: tt.image, RawArgs: []byte(tt.raw)})
if !errors.Is(err, coretypes.ErrInvalidEngineArgs) {
t.Errorf("got %v, want ErrInvalidEngineArgs", err)
}
if lines := runner.Lines(); len(lines) != 0 {
t.Errorf("got %q, want no round trip", lines)
}
})
}
}

func TestCreateArgvForAWindowsGuest(t *testing.T) {
opts := &enginetypes.VirtualizationCreateOptions{Image: "win11", User: "eru"}
resource := &engine.VirtualizationResource{Quota: 4, Memory: 4 << 30, Volumes: []string{"/scratch:/scratch:rw:2147483648"}}
Expand Down
Loading
Loading