-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathsystemd.yaml
More file actions
350 lines (346 loc) · 15.5 KB
/
Copy pathsystemd.yaml
File metadata and controls
350 lines (346 loc) · 15.5 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
- id: systemd.failed_to_start_unit
technology: systemd
title: "Failed to start <unit>"
summary: >-
systemd reports a unit failed to start. The unit entered a failed state and
its ExecStart never reached a running condition.
applies_to: [log, command_output, error_string]
match:
any_of:
- "Failed to start "
- "Job for .* failed because"
- "entered failed state"
- "Failed with result"
weight: 0.7
root_causes:
- title: "ExecStart command failed or is misconfigured"
description: >-
The binary path is wrong, an argument is invalid, or the program exits
non-zero immediately, so the unit start job fails.
confidence: 0.6
category: configuration
- title: "Missing dependency, environment, or file"
description: >-
A required mount, network, env file, or working directory is absent, so
ExecStart cannot run successfully.
confidence: 0.45
category: configuration
- title: "Permission or sandboxing prevents start"
description: >-
The unit's User/Group lacks access, or hardening directives (e.g.
ProtectSystem, ReadOnlyPaths) block a needed action.
confidence: 0.4
category: authentication
diagnostic_commands:
- command: "systemctl status <unit> --no-pager -l"
explanation: "Shows the unit's active/failed state, last exit code, and recent log lines."
expected_output: "Active: failed (Result: exit-code) with the failing ExecStart line."
platform: "affected host"
- command: "journalctl -u <unit> -b --no-pager"
explanation: "Prints this boot's full log for the unit, including the underlying error."
expected_output: "The program's stderr / startup error message."
- command: "systemd-analyze verify <unit>"
explanation: "Statically checks the unit file for syntax and directive errors."
expected_output: "No output when valid; otherwise specific unit-file warnings."
suggested_fixes:
- title: "Correct ExecStart and reload"
description: >-
Fix the command path/arguments revealed by the logs, run the binary
manually as the unit's user to confirm, then reload and restart.
snippet: |
systemctl daemon-reload
systemctl restart <unit>
- title: "Provide the missing dependency or relax hardening"
description: "Add the required After=/Requires=, env file, or adjust sandboxing directives the logs implicate."
references:
- title: "Why a systemd service fails to start"
url: "https://devopsaitoolkit.com/blog/systemd-failed-to-start-unit"
source: "devopsaitoolkit"
- title: "systemd.service unit configuration"
url: "https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html"
source: "official docs"
best_practices:
- "Validate unit files with systemd-analyze verify before deploying."
- "Test ExecStart manually as the configured User before enabling."
prevention:
- "Pin absolute paths in ExecStart and declare all dependencies explicitly."
tags: [units, service, startup]
- id: systemd.unit_not_found
technology: systemd
title: "Unit <name> not found"
summary: >-
systemd cannot find the requested unit, so the command or a dependency
reference fails with 'Unit ... not found'.
applies_to: [log, command_output, error_string]
match:
any_of:
- "Unit .* not found"
- "Failed to .*: Unit .* not found"
- "not found\\.$"
- "No such file or directory.*\\.service"
weight: 0.74
root_causes:
- title: "Unit file is missing or misnamed"
description: >-
The .service/.socket file does not exist in any unit search path, or its
name/extension is misspelled in the command or a Requires=/After=.
confidence: 0.6
category: configuration
- title: "Unit added but daemon not reloaded"
description: >-
A new unit file was dropped in but systemctl daemon-reload was not run,
so systemd has not picked it up.
confidence: 0.5
category: configuration
- title: "Unit installed in a path systemd does not scan"
description: >-
The file lives outside /etc/systemd/system, /run/systemd/system, or the
lib unit dirs, so it is invisible.
confidence: 0.35
category: configuration
diagnostic_commands:
- command: "systemctl list-unit-files | grep <name>"
explanation: "Checks whether systemd knows about the unit at all."
expected_output: "A matching unit-file line; no match confirms it is not registered."
platform: "affected host"
- command: "systemctl cat <unit>"
explanation: "Prints the resolved unit file and its drop-ins, or errors if absent."
expected_output: "The full unit contents, or 'No files found' when missing."
- command: "systemd-analyze unit-paths"
explanation: "Lists every directory systemd searches for unit files."
expected_output: "The unit search paths; confirm your file is in one of them."
suggested_fixes:
- title: "Place the unit in a search path and reload"
description: >-
Install the file under /etc/systemd/system with the correct name, then
reload the daemon so systemd registers it.
snippet: |
cp myapp.service /etc/systemd/system/
systemctl daemon-reload
- title: "Correct the referenced unit name"
description: "Fix the typo/extension in the command or in Requires=/After= references."
references:
- title: "Resolving systemd 'Unit not found' errors"
url: "https://devopsaitoolkit.com/blog/systemd-unit-not-found"
source: "devopsaitoolkit"
- title: "systemd unit file load path"
url: "https://www.freedesktop.org/software/systemd/man/latest/systemd.unit.html"
source: "official docs"
best_practices:
- "Always run daemon-reload after adding or editing unit files."
- "Reference units by their exact, fully-qualified name."
prevention:
- "Automate unit install + daemon-reload in your provisioning."
tags: [units, not-found, configuration]
- id: systemd.service_exited_nonzero
technology: systemd
title: "Service exited with non-zero status (code=exited)"
summary: >-
The main process of a service exited with a non-zero code; systemd marks the
unit failed with Result: exit-code.
applies_to: [log, command_output, error_string]
match:
any_of:
- "code=exited, status=[1-9]"
- "Main process exited, code=exited"
- "Failed with result 'exit-code'"
- "status=\\d+/[A-Z]+"
weight: 0.75
root_causes:
- title: "Application returned an error on startup or runtime"
description: >-
The program itself exited non-zero due to a config error, failed
dependency, bad credentials, or an unhandled runtime fault.
confidence: 0.6
category: application
- title: "Wrong Type= for the process model"
description: >-
A forking daemon set as Type=simple (or vice versa) makes systemd
misread the exit, reporting failure.
confidence: 0.4
category: configuration
- title: "Missing environment or permissions"
description: >-
A required env var/secret is absent or the User lacks access, so the
process aborts with a non-zero code.
confidence: 0.4
category: configuration
diagnostic_commands:
- command: "systemctl status <unit> --no-pager -l"
explanation: "Reveals the exit code and status systemd recorded for the main process."
expected_output: "Main process exited, code=exited, status=<n>/<SIGNAL or reason>."
platform: "affected host"
- command: "journalctl -u <unit> -n 100 --no-pager"
explanation: "Shows the last 100 log lines from the unit, including the error before exit."
expected_output: "The application's fatal error or stack trace preceding the exit."
- command: "systemctl show <unit> -p Type -p ExecStart -p ExecMainStatus"
explanation: "Displays the configured Type, command, and last main exit status."
expected_output: "Type, ExecStart, and ExecMainStatus values to cross-check the model."
suggested_fixes:
- title: "Fix the underlying application error"
description: >-
Use the journal to find why the process exited, correct the
config/dependency, and confirm by running the command manually.
- title: "Match Type= to the process behavior"
description: "Set Type=simple/exec/forking/notify to how the binary actually runs."
snippet: |
[Service]
Type=simple
ExecStart=/usr/bin/myapp --foreground
references:
- title: "Diagnosing systemd services that exit non-zero"
url: "https://devopsaitoolkit.com/blog/systemd-service-exited-nonzero"
source: "devopsaitoolkit"
- title: "systemd.service: Type= and exit handling"
url: "https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Type="
source: "official docs"
best_practices:
- "Make services run in the foreground and log to stdout/journal."
- "Choose Type= to match the binary's foreground/forking behavior."
prevention:
- "Validate required env/secrets before the service starts."
tags: [service, exit-code, failure]
- id: systemd.dependency_failed
technology: systemd
title: "Dependency failed for <unit>"
summary: >-
systemd skipped or failed a unit because a unit it depends on failed, so the
job result is 'dependency'.
applies_to: [log, command_output, error_string]
match:
any_of:
- "Dependency failed for "
- "Job .* failed with result 'dependency'"
- "Requested dependency .* failed"
weight: 0.74
root_causes:
- title: "A required upstream unit failed"
description: >-
A unit listed in Requires=/Requisite= failed to start, so systemd
refuses to start the dependent unit.
confidence: 0.65
category: configuration
- title: "Mount, network, or device target not reached"
description: >-
The unit depends on a mount/target (e.g. network-online.target) that
never became active.
confidence: 0.45
category: configuration
- title: "Ordering vs requirement confusion"
description: >-
After= (ordering only) is mistaken for a requirement, or a hard Requires=
is used where Wants= would be more resilient.
confidence: 0.35
category: configuration
diagnostic_commands:
- command: "systemctl list-dependencies <unit>"
explanation: "Shows the dependency tree so you can see which required unit is failing."
expected_output: "A tree where the failed dependency is marked with a red/failed state."
platform: "affected host"
- command: "systemctl --failed --no-pager"
explanation: "Lists all currently failed units, exposing the failing upstream dependency."
expected_output: "The dependency unit listed in the failed table."
- command: "journalctl -u <dependency-unit> -b --no-pager"
explanation: "Shows why the upstream dependency itself failed."
expected_output: "The root error from the dependency unit's start attempt."
suggested_fixes:
- title: "Fix the failing upstream unit"
description: >-
Repair the dependency the tree highlights; once it starts cleanly, the
dependent unit can start. Restart both after the fix.
- title: "Adjust dependency strength/ordering"
description: "Use Wants= for soft dependencies and ensure After= orders against the right target."
snippet: |
[Unit]
Wants=network-online.target
After=network-online.target
references:
- title: "Untangling systemd dependency failures"
url: "https://devopsaitoolkit.com/blog/systemd-dependency-failed"
source: "devopsaitoolkit"
- title: "systemd.unit: requirement and ordering dependencies"
url: "https://www.freedesktop.org/software/systemd/man/latest/systemd.unit.html"
source: "official docs"
best_practices:
- "Prefer Wants= over Requires= unless failure must cascade."
- "Order against network-online.target only when truly needed."
prevention:
- "Test dependency chains by starting the leaf unit on a clean boot."
tags: [dependencies, ordering, targets]
- id: systemd.start_limit_hit
technology: systemd
title: "start-limit-hit: unit not restarting"
summary: >-
systemd stopped restarting a unit because it failed too many times within the
configured interval, entering a failed start-limit state.
applies_to: [log, command_output, error_string]
match:
any_of:
- "start-limit-hit"
- "start request repeated too quickly"
- "Failed with result 'start-limit-hit'"
- "StartLimitIntervalSec"
weight: 0.78
root_causes:
- title: "Underlying failure causes a rapid crash loop"
description: >-
The service keeps exiting quickly, and once it crosses StartLimitBurst
within StartLimitIntervalSec, systemd refuses further restarts.
confidence: 0.65
category: application
- title: "Restart settings too aggressive for the workload"
description: >-
A short RestartSec with a low burst limit trips the rate limiter before
a slow-recovering dependency is ready.
confidence: 0.45
category: configuration
- title: "Persistent config error that never resolves"
description: >-
A bad config or missing resource means every start attempt fails
identically until the limit is reached.
confidence: 0.4
category: configuration
diagnostic_commands:
- command: "systemctl status <unit> --no-pager -l"
explanation: "Confirms the start-limit-hit result and shows the recent failure pattern."
expected_output: "Active: failed (Result: start-limit-hit) with repeated start entries."
platform: "affected host"
- command: "journalctl -u <unit> --since '-15min' --no-pager"
explanation: "Reveals the repeating root error behind the crash loop."
expected_output: "The same fatal error logged on each restart attempt."
- command: "systemctl show <unit> -p StartLimitBurst -p StartLimitIntervalSec -p RestartSec"
explanation: "Displays the rate-limit and restart settings in effect."
expected_output: "The configured burst, interval, and RestartSec values."
suggested_fixes:
- title: "Fix the crash cause, then reset the limit"
description: >-
Resolve the repeating failure from the journal, run systemctl
reset-failed, and start the unit again.
snippet: |
systemctl reset-failed <unit>
systemctl start <unit>
- title: "Tune restart rate limiting"
description: "Adjust StartLimitIntervalSec/StartLimitBurst and RestartSec to allow recovery without masking real failures."
snippet: |
[Unit]
StartLimitIntervalSec=300
StartLimitBurst=5
[Service]
Restart=on-failure
RestartSec=10
references:
- title: "Fixing systemd start-limit-hit crash loops"
url: "https://devopsaitoolkit.com/blog/systemd-start-limit-hit"
source: "devopsaitoolkit"
- title: "systemd.unit: StartLimitIntervalSec / StartLimitBurst"
url: "https://www.freedesktop.org/software/systemd/man/latest/systemd.unit.html#StartLimitIntervalSec="
source: "official docs"
warnings:
- message: "Raising the restart limit without fixing the crash just hides a broken service."
severity: medium
best_practices:
- "Always reset-failed only after fixing the underlying crash."
- "Set RestartSec high enough to avoid tight crash loops."
prevention:
- "Add health checks and validate config before enabling auto-restart."
tags: [restart, rate-limit, crash-loop]