Skip to content

docs(install): add fleet install guide and Ansible playbook - #373

Merged
mayankpande88 merged 2 commits into
mainfrom
docs/fleet-install-ansible
Oct 8, 2026
Merged

mayankpande88 merged 2 commits into
mainfrom
docs/fleet-install-ansible

Conversation

@PrashantBtkl

@PrashantBtkl PrashantBtkl commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does

Adds a guide and an Ansible playbook for installing the agent on many hosts in one run, instead of running install.sh by hand on each one. Docs and examples only; no agent or installer code changes.

Linked issue / context

No public issue. This is the install step of running the agent on standalone VMs.

How to check it

  1. Read docs/fleet-install.md (linked from the README install section).
  2. With Ansible installed, list a test host in inventory.ini (copied from inventory.example.ini), copy vars.example.yml to vars.yml, and run:
    ansible-playbook -i inventory.ini -e @vars.yml install-node-agent.yml --limit <host>
    Expect the final "Check the agent is up and stayed up" task to pass, and /etc/default/nudgebee-node-agent on the host to be mode 0600 and hold your settings.
  3. Run it again: the agent restarts (the installer always restarts it) and the check passes again.
  4. Run with -e node_agent_state=absent: the service, binary and settings file are removed.

How was this tested?

  • yamllint and ansible-lint (production profile) pass; --syntax-check passes.
  • --check run against localhost: host checks pass, install steps are skipped.
  • The settings task was run for real against a scratch path with values containing \ and "; systemd read every value back unchanged. A value with a single quote is rejected with a clear message.
  • Not run: a real install on a remote host. Step 2 above is the one to verify.

eBPF changes?

No.

Checklist

  • make lint / make test: not applicable (no Go changes)
  • CHANGELOG.md updated under ## [Unreleased]
  • No real credentials or internal hostnames (example hosts use RFC 5737 addresses and example.com)

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a fleet installation guide and an Ansible playbook to support installing, upgrading, or removing the nudgebee node agent across multiple Linux hosts. The review feedback highlights several key improvements: ensuring the playbook explicitly starts the service to handle cases where no configuration changes are detected, replacing a shell loop anti-pattern in the documentation with a more robust while read loop, warning users about GitHub rate-limiting risks when using the latest version on large fleets, and adding a comment to guide developers on overriding the installer URL for local testing.

Comment thread examples/ansible/install-node-agent.yml Outdated
Comment thread docs/fleet-install.md Outdated
Comment thread docs/fleet-install.md Outdated
Comment thread examples/ansible/install-node-agent.yml Outdated
@PrashantBtkl
PrashantBtkl force-pushed the docs/fleet-install-ansible branch from aa4fd2e to 19e71ad Compare October 8, 2026 06:54

@mayankpande88 mayankpande88 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes. Both install paths here download install.sh from the prod branch, and that copy is still the upstream installer, so neither works as written (details inline). A real install on one host would have caught it; the description says that wasn't run. Please do one run against a test host before merge.

Also blocking:

  • API_KEY ends up on the sudo command line on every host (inline on install-node-agent.yml:44).
  • Play-level vars: silently override inventory/group/host vars.
  • The service health check passes for an agent that is crash-looping.

Several statements in docs/fleet-install.md don't match main: TRACES_SAMPLING and other flags are dropped by the installer's whitelist, LISTEN can't change the push-mode listen address, --help doesn't list settings, re-runs always restart, and COLLECTOR_ENDPOINT metrics aren't OTLP.

PR/commit hygiene, since this repo is public:

  • The description links a private-repo issue and an org discussion. Please remove both.
  • Please drop the Co-Authored-By: Claude trailer from the commit and the "Generated with" footer from the description. The trailer needs a commit rewrite, not just a description edit.

Minor, not inline: the installer file is left in /var/tmp when the install fails (a block/always would fix that and remove the five repeated when: lines). A typo in node_agent_state runs nothing and reports success. changed_when depends on the exact wording of an installer log line.

Comment thread examples/ansible/install-node-agent.yml Outdated
Comment thread examples/ansible/install-node-agent.yml Outdated
Comment thread examples/ansible/install-node-agent.yml Outdated
Comment thread examples/ansible/install-node-agent.yml Outdated
Comment thread examples/ansible/install-node-agent.yml Outdated
Comment thread docs/fleet-install.md Outdated
Comment thread docs/fleet-install.md Outdated
Comment thread docs/fleet-install.md Outdated
Comment thread docs/fleet-install.md Outdated
Comment thread examples/ansible/inventory.example.ini Outdated
@PrashantBtkl
PrashantBtkl force-pushed the docs/fleet-install-ansible branch from 19e71ad to 73409fe Compare October 8, 2026 07:21
@PrashantBtkl

Copy link
Copy Markdown
Contributor Author

Thanks, all addressed in the latest push (single commit, trailer removed, description cleaned up).

  • Installer ref: the script is now fetched from the same ref as the binary (main for latest, otherwise the version tag). The SSH example does the same. Verified prod still carries the upstream installer.
  • API_KEY / dropped flags: settings are written to /etc/default/nudgebee-node-agent (0600, no_log) instead of the installer's environment. That keeps the key off the command line and makes every flag in flags.go work, including TRACES_SAMPLING and CA_FILE, with no install.sh change. no_log is gone from the install task. Values are written single-quoted; a value with a single quote or newline is rejected. Checked that systemd reads them back unchanged.
  • Variable precedence: play-level vars: removed; every variable uses | default(...) at the point of use, so inventory, group_vars and host_vars work.
  • Health check: waits 12s (two restart intervals) and requires ActiveState == active and NRestarts == 0. Also starts/enables the unit explicitly, and restarts it on a settings-only change.
  • Cleanup: installer removal is in an always, the install steps are in a block, an unknown node_agent_state fails, and changed_when no longer parses installer output.
  • Docs: corrected the claims about --help, "every setting", re-run behaviour, LISTEN, and COLLECTOR_ENDPOINT; added passwordless sudo, the rate-limit note, the vault flags, and RFC 5737 / example.com examples; the SSH loop is now a while read loop.

Still not done: a real install on a remote host.

@mayankpande88 mayankpande88 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All review points addressed. I ran the playbook for real on a Debian 12 VM (systemd 252, arm64) against v0.1.9:

  • Fresh install: passes. /etc/default/nudgebee-node-agent is 0600 root. An API_KEY containing \, " and $ reaches the agent's environment unchanged, and it doesn't appear in the journal or the sudo log. The installer's own env file is empty, and the agent listens on 127.0.0.1:10300.
  • Re-run, nothing changed: passes. The settings task reports ok.
  • TRACES_SAMPLING: "abc": the health check fails on ActiveState, and the journal shows the ParseFloat exit and the crash-loop.
  • state=absent: the unit, binary and settings file are all removed.
  • SSH loop from the docs: installs; the settings file is 0600, the service is active.

Follow-up, not blocking: NRestarts == 0 works only because the installer always restarts the agent (systemd resets the counter on a manual start). If the installer's "No change detected" path is ever fixed, a healthy host with an earlier auto-restart (e.g. one OOM kill) would fail every re-run. Comparing NRestarts before and after the pause avoids that. Minor wording: latest is resolved through the github.com/.../releases/latest redirect, not the GitHub API.

@mayankpande88
mayankpande88 merged commit 2b16822 into main Oct 8, 2026
7 checks passed
@mayankpande88
mayankpande88 deleted the docs/fleet-install-ansible branch October 8, 2026 08:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants