Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions docs/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -589,6 +589,27 @@ a crashed node's descriptor leases block DDL until they expire.
All three must be positive. `scope_expiry_interval_secs` has a floor of `10`.
Below that the sweep costs more than the resolution it buys.

**Startup bounds:**

Boot waits for the metadata raft group to apply its first entry, and for the
locally hosted data raft groups to replay their retained logs. Both waits are
bounded, and both are configurable. The defaults are sized for a cold restart of
a data dir that has been running for weeks: the backlog takes minutes to apply,
not seconds.

| Config field | Default |
| ----------------------------------------------- | -------- |
| `tuning.startup.raft_ready_timeout_ms` | `300000` |
| `tuning.startup.data_group_recovery_timeout_ms` | `600000` |

The metadata bound resets on every applied-index advance, so a large replay
finishes; only a group that applies nothing fails. The data-group recovery bound
is a hard deadline: it does not reset, and a group still recovering when it
expires fails the boot. Both accept `1` to `86400000` ms (one day). A value
outside that range is rejected at load with the key, the value and the accepted
range named; a bound written under `[server]` is rejected with the path that is
read named instead.

**Observability settings:**

| Config field | Environment variable | Default |
Expand Down
15 changes: 12 additions & 3 deletions nodedb-test-support/src/single_node.rs
Original file line number Diff line number Diff line change
Expand Up @@ -117,9 +117,18 @@ pub async fn start(
gateway_enable_gate: sequencer.register_gate(StartupPhase::GatewayEnable, "gateway"),
};
// The harness cores replay their WAL before they report ready, so no
// replay receiver is owed here.
nodedb::bootstrap::cluster_ready::await_cluster_ready(shared, ready_rx, Vec::new(), gates)
.await?;
// replay receiver is owed here. The harness runs without a config file, so
// the bounds are the shipped defaults.
let startup = nodedb_types::config::tuning::StartupTuning::default();
nodedb::bootstrap::cluster_ready::await_cluster_ready(
shared,
ready_rx,
Vec::new(),
gates,
startup.raft_ready_timeout(),
startup.data_group_recovery_timeout(),
)
.await?;
Ok(raft)
}

Expand Down
19 changes: 19 additions & 0 deletions nodedb-types/src/config/tuning/config.rs
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ use super::memory::MemoryTuning;
use super::network::{BridgeTuning, ClusterTransportTuning, NetworkTuning, WalTuning};
use super::scheduler::SchedulerTuning;
use super::shutdown::ShutdownTuning;
use super::startup::StartupTuning;

/// Top-level tuning configuration.
///
Expand Down Expand Up @@ -51,6 +52,8 @@ pub struct TuningConfig {
pub bitemporal: BitemporalTuning,
#[serde(default)]
pub maintenance: MaintenanceTuning,
#[serde(default)]
pub startup: StartupTuning,
}

impl TuningConfig {
Expand Down Expand Up @@ -177,4 +180,20 @@ doc_cache_entries = 8192
assert_eq!(cfg.memory.overflow_max_bytes, 2 * 1024 * 1024 * 1024);
assert_eq!(cfg.memory.doc_cache_entries, 8192);
}

#[test]
fn startup_bounds_are_read_from_the_aggregate() {
let cfg: TuningConfig = toml::from_str("").expect("deserialize");
assert_eq!(cfg.startup.raft_ready_timeout_ms, 300_000);
assert_eq!(cfg.startup.data_group_recovery_timeout_ms, 600_000);

let toml_str = r#"
[startup]
raft_ready_timeout_ms = 1234
data_group_recovery_timeout_ms = 5678
"#;
let cfg: TuningConfig = toml::from_str(toml_str).expect("deserialize");
assert_eq!(cfg.startup.raft_ready_timeout_ms, 1234);
assert_eq!(cfg.startup.data_group_recovery_timeout_ms, 5678);
}
}
5 changes: 5 additions & 0 deletions nodedb-types/src/config/tuning/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@ mod memory;
mod network;
mod scheduler;
mod shutdown;
mod startup;

pub use bitemporal::BitemporalTuning;
pub use config::TuningConfig;
Expand All @@ -23,3 +24,7 @@ pub use memory::MemoryTuning;
pub use network::{BridgeTuning, ClusterTransportTuning, NetworkTuning, WalTuning};
pub use scheduler::SchedulerTuning;
pub use shutdown::ShutdownTuning;
pub use startup::{
DEFAULT_DATA_GROUP_RECOVERY_TIMEOUT_MS, DEFAULT_RAFT_READY_TIMEOUT_MS,
DataGroupRecoveryTimeout, RaftReadyTimeout, StartupTuning,
};
126 changes: 126 additions & 0 deletions nodedb-types/src/config/tuning/startup.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,126 @@
// SPDX-License-Identifier: Apache-2.0

//! Startup tuning: boot-time bounds applied by the readiness gates.

use std::time::Duration;

use serde::{Deserialize, Serialize};

/// Default bound for the metadata-group readiness stall.
///
/// A node with a large metadata apply backlog (tens of thousands of entries
/// from a burst of cross-shard writes) needs minutes of replay before the
/// metadata group applies its first entry. A tighter bound turns a slow boot
/// into a restart loop.
pub const DEFAULT_RAFT_READY_TIMEOUT_MS: u64 = 300_000;

/// Default bound for the local data raft groups' replay.
///
/// Sized for a cold restart on a data dir that has been running for weeks: the
/// backlog of committed entries takes minutes to apply, not seconds.
pub const DEFAULT_DATA_GROUP_RECOVERY_TIMEOUT_MS: u64 = 600_000;

/// Metadata-group readiness stall bound, in the type system.
///
/// Two adjacent `Duration` arguments transpose without a compile error. This
/// newtype pins the bound where it is chosen and handed to a gate; a helper
/// that takes plain `Duration` values converts at its own boundary.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct RaftReadyTimeout(pub Duration);

/// Local data-group replay bound, in the type system. See [`RaftReadyTimeout`].
///
/// A separate type, not a shared one: the two bounds measure different waits
/// and are passed by different call sites, so making them interchangeable buys
/// nothing and costs a class of silent mix-ups.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct DataGroupRecoveryTimeout(pub Duration);

impl From<Duration> for RaftReadyTimeout {
fn from(value: Duration) -> Self {
Self(value)
}
}

impl From<Duration> for DataGroupRecoveryTimeout {
fn from(value: Duration) -> Self {
Self(value)
}
}

fn default_raft_ready_timeout_ms() -> u64 {
DEFAULT_RAFT_READY_TIMEOUT_MS
}

fn default_data_group_recovery_timeout_ms() -> u64 {
DEFAULT_DATA_GROUP_RECOVERY_TIMEOUT_MS
}

/// Boot-time bounds for the readiness gates in `bootstrap::cluster_ready`.
///
/// Both are range-checked where the config is loaded: below 1 ms the wait is
/// over before the group it waits for can answer, and above a day it stops
/// bounding anything while the instant it is added to is still finite.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct StartupTuning {
/// How long the metadata raft group may go without applying an entry
/// before the readiness gate fails startup. Every applied-index advance
/// resets the clock, so a large replay finishes; only a stuck group fails.
/// Default: 300_000 (5 minutes).
#[serde(default = "default_raft_ready_timeout_ms")]
pub raft_ready_timeout_ms: u64,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing checks either bound. 0 and u64::MAX both deserialize, and both break the boot (see the note on data_group_recovery.rs:185). Add a range check at config load that names the key, the value, and the valid range.


/// How long the locally hosted data raft groups have to replay their
/// retained logs before startup fails. Default: 600_000 (10 minutes).
#[serde(default = "default_data_group_recovery_timeout_ms")]
pub data_group_recovery_timeout_ms: u64,
}

impl Default for StartupTuning {
fn default() -> Self {
Self {
raft_ready_timeout_ms: default_raft_ready_timeout_ms(),
data_group_recovery_timeout_ms: default_data_group_recovery_timeout_ms(),
}
}
}

impl StartupTuning {
/// Metadata-group readiness stall bound.
pub fn raft_ready_timeout(&self) -> RaftReadyTimeout {
RaftReadyTimeout(Duration::from_millis(self.raft_ready_timeout_ms))
}

/// Data-group recovery bound.
pub fn data_group_recovery_timeout(&self) -> DataGroupRecoveryTimeout {
DataGroupRecoveryTimeout(Duration::from_millis(self.data_group_recovery_timeout_ms))
}
}

#[cfg(test)]
mod tests {
use super::*;

#[test]
fn defaults_are_the_backlogged_recovery_bounds() {
let cfg = StartupTuning::default();
assert_eq!(cfg.raft_ready_timeout_ms, 300_000);
assert_eq!(cfg.data_group_recovery_timeout_ms, 600_000);
assert_eq!(
cfg.raft_ready_timeout(),
RaftReadyTimeout(Duration::from_secs(300))
);
assert_eq!(
cfg.data_group_recovery_timeout(),
DataGroupRecoveryTimeout(Duration::from_secs(600))
);
}

#[test]
fn one_bound_set_leaves_the_other_at_its_default() {
let cfg: StartupTuning =
toml::from_str("raft_ready_timeout_ms = 1234").expect("deserialize");
assert_eq!(cfg.raft_ready_timeout_ms, 1234);
assert_eq!(cfg.data_group_recovery_timeout_ms, 600_000);
}
}
35 changes: 35 additions & 0 deletions nodedb-wal/src/error.rs
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,22 @@ pub enum WalError {
#[error("unsupported WAL format version {version} (supported: {supported})")]
UnsupportedVersion { version: u16, supported: u16 },

/// A segment holds records in a WAL format version this build cannot read.
///
/// The same finding as [`WalError::UnsupportedVersion`], reported by the
/// callers that know which file carries it. A reader sees one header at a
/// time and holds no path, and a version alone does not tell an operator
/// which segment to act on.
#[error(
"WAL segment '{path}' holds records in WAL format version {version}; this build reads \
version {supported}"
)]
SegmentFormatVersion {
path: String,
version: u16,
supported: u16,
},

/// Unknown required record type encountered during replay.
/// Optional unknown record types are safely skipped.
#[error("unknown required record type {record_type} at LSN {lsn}")]
Expand Down Expand Up @@ -214,4 +230,23 @@ pub enum WalError {
},
}

impl WalError {
/// Name the segment a version gap was found in.
///
/// A reader reports [`WalError::UnsupportedVersion`] from header validation
/// alone and holds no path. The callers that opened the file do, and only
/// they can turn the finding into something an operator can act on. Every
/// other error passes through unchanged.
pub(crate) fn with_segment_path(self, path: &std::path::Path) -> Self {
match self {
Self::UnsupportedVersion { version, supported } => Self::SegmentFormatVersion {
path: path.display().to_string(),
version,
supported,
},
other => other,
}
}
}

pub type Result<T> = std::result::Result<T, WalError>;
Loading
Loading