lib

Core libraries for Radroots
git clone https://radroots.dev/git/lib.git
Log | Files | Refs | README

commit 3225d43c15f5124ce7e990aaeb6817b1713b847b
parent 18620a5d289190e36353e32643587d2636d9c34e
Author: triesap <tyson@radroots.org>
Date:   Tue, 11 Aug 2026 23:38:45 +0000

service-sqlite: recover interrupted restore

- reconcile exact marker topologies under writer authority
- preserve error precedence and valid markers across scratch drift
- keep non-writable opens fail-closed and mutation-free
- verify idempotent recovery across native and cross-target gates

Diffstat:
MAGENTS.md | 27+++++++++++++++++++++++----
Mcrates/service_sqlite/README.md | 36+++++++++++++++++++++++++++++++-----
Mcrates/service_sqlite/src/connection.rs | 6++++++
Mcrates/service_sqlite/src/open.rs | 119+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--
Mcrates/service_sqlite/src/restore/marker.rs | 280++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------
Mcrates/service_sqlite/src/restore/mod.rs | 26++++++--------------------
Acrates/service_sqlite/src/restore/recover.rs | 1052+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcrates/service_sqlite/src/restore/stage.rs | 50++++++++++++++++++++++++++++++++++++--------------
Mcrates/service_sqlite/tests/package_boundary.rs | 83++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------
9 files changed, 1605 insertions(+), 74 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -228,10 +228,29 @@ Before editing code: cleanup must remain disarmed after every later error so the marker never loses a bound artifact. Successful finalization returns no host and leaves the old live database and - `replacement_installed` marker for open-time recovery. Until that recovery - exists, every pool open must reject marker or marker-scratch evidence as - `Recovery`. Finalization must not roll back, delete recovery evidence, reopen - SQLite, or expose paths, descriptors, marker controls, or rename controls. + `replacement_installed` marker for the next writable open to recover. Other + open modes reject that evidence as `Recovery`. Finalization itself must not + roll back, delete recovery evidence, reopen SQLite, or expose paths, + descriptors, marker controls, or rename controls. +- Interrupted restore recovery is private and automatic only for + read-write-existing open under exclusive `WriterAuthority`, before any + SQLite connection or await point. Initialize, initialized-open, and + read-only inspection must reject every fixed stage, backup, marker, or + marker-scratch artifact without mutation. Recovery must bind the marker to + the requested database identity, hash and revalidate exact owner-only + single-link artifacts, reject sidecars, and let exact topology decide the + sole action: roll back `prepared` while old live is still installed, then + roll forward once old live is durably retained. Persist every inferred phase + before the next destructive step; retire exact backup before marker; and + admit marker scratch only as a canonical topology-consistent one-edge + successor whose exact bound inode is removed and durably reproduced through + the governed marker-advance path without overwriting the valid marker. + Repeated recovery may + finish already-absent stage or backup cleanup, but every other missing, + replaced, linked, malformed, mismatched, or ambiguous artifact remains + `Recovery` evidence. Do not expose recovery controls, add a background task + or hidden timeout, repair without writer authority, or fold Step 070 + integrity/status APIs and the later process failpoint harness into recovery. - Runtime-management flows consume a sealed `RuntimeContext` for every service instance. They must not reconstruct service paths from raw identifiers, ambient selectors, or manager-owned roots, and registries must not persist diff --git a/crates/service_sqlite/README.md b/crates/service_sqlite/README.md @@ -167,11 +167,37 @@ authority until it either fails before durability or establishes recovery evidence and continues. Once `prepared` is durable, staged-artifact cleanup is disarmed and the bound stage remains available after every later error. Success returns no database host, retains the old live database and final -marker, and requires a new open. Until open-time recovery is implemented, -every initialized, read-write-existing, or read-only-inspection pool open -refuses a marker or marker-scratch as `Recovery` before opening SQLite. -Finalization does not delete evidence, roll back, reconcile an interruption, -or reopen the database; those remain the next recovery checkpoint. +marker, and requires a new open. Read-write-existing open is the sole recovery +path. Under exclusive writer authority and before opening SQLite, it validates +the marker's exact service, instance, source generation, application ID, schema +ceiling, artifact identities, lengths, digests, restrictive modes, and the +absence of database sidecars. Read-only inspection, initialization, and an +initialized open never recover; they reject any stage, backup, marker, or marker +scratch as `Recovery` without mutation. + +Recovery uses exact topology as the durable authority. `prepared` with the old +live database still installed rolls back by removing only the exact stage and +then the marker. Once the exact old live inode has reached the backup name, +recovery advances and rolls forward. A lagging `live_retained` phase installs or +recognizes the exact replacement, advances to `replacement_installed`, then +removes the exact old backup before retiring the marker. Interrupted rollback +and final cleanup accept only the corresponding already-absent exact artifact, +so repeated recovery is idempotent. Every other topology, sidecar, replacement, +link, mode, owner, length, digest, identity, or directory-authority mismatch +fails closed and preserves the evidence. + +A marker scratch is admitted only when it is the canonical one-edge successor +of the current marker and the artifact topology already proves that successor. +Recovery removes only the exact bound scratch inode, synchronizes that removal, +then reproduces the transition through the governed marker-advance path. This +preserves the valid current marker if the scratch pathname was replaced. +Orphaned, malformed, skipped, same-phase, terminal, mismatched, or +topology-inconsistent scratch is never deleted or reinterpreted. Recovery has no +await point or hidden task: once a writable open is polled, each synchronous +filesystem step and its authority checks complete before the open can be +cancelled. A later cancelled SQLite open is retried by rereading the already +durable, marker-free state. Finalization itself does not reconcile or reopen the +database. The crate owns mechanics only. Service-specific tables, SQL, repositories, backup content policy, identity material, process lifecycle, and readiness diff --git a/crates/service_sqlite/src/connection.rs b/crates/service_sqlite/src/connection.rs @@ -59,6 +59,12 @@ enum ServiceSqliteHostCloseState { impl ServiceSqliteHost { /// Opens existing writable state and finishes every pending governed migration. + /// + /// Before opening SQLite, this path holds exclusive writer authority and + /// synchronously reconciles any exact interrupted-restore topology. The + /// recovery sequence has no await point: cancellation cannot split one + /// filesystem step from its authority check. If the surrounding open is + /// cancelled later, a retry re-reads the already durable filesystem state. #[allow(clippy::too_many_arguments)] pub async fn open_read_write_existing( paths: &ServiceSqlitePaths, diff --git a/crates/service_sqlite/src/open.rs b/crates/service_sqlite/src/open.rs @@ -675,12 +675,15 @@ async fn open_connection_pool( return Err(ServiceSqliteError::new(ServiceSqliteErrorKind::Integrity)); } let recovery_guard = match (mode, authority.as_ref(), inspection_guard.as_ref()) { - (OpenMode::Initialize | OpenMode::ReadWriteExisting, Some(authority), None) => { + (OpenMode::Initialize, Some(authority), None) => { authority.validate_for(paths)?; let result = crate::restore::refuse_unresolved_recovery(authority.directory()); authority.validate_for(paths)?; result } + (OpenMode::ReadWriteExisting, Some(authority), None) => { + crate::restore::recover_for_open(paths, identity, authority) + } (OpenMode::ReadOnlyInspection, None, Some(inspection_guard)) => { inspection_guard.validate_for(paths)?; let result = crate::restore::refuse_unresolved_recovery(&inspection_guard.directory); @@ -1589,7 +1592,7 @@ mod tests { convert::Infallible, fs, num::NonZeroU32, - os::unix::fs::{PermissionsExt, symlink}, + os::unix::fs::{MetadataExt, PermissionsExt, symlink}, sync::{ Arc, atomic::{AtomicUsize, Ordering as AtomicOrdering}, @@ -1603,6 +1606,8 @@ mod tests { }; #[cfg(any(target_os = "linux", target_os = "macos"))] use radroots_storage::event::SourceGeneration; + #[cfg(any(target_os = "linux", target_os = "macos"))] + use sha2::{Digest, Sha256}; #[cfg(any(target_os = "linux", target_os = "macos"))] use crate::{ServiceDatabaseMetadata, ServiceSqliteApplicationId}; @@ -1652,6 +1657,18 @@ mod tests { } #[cfg(any(target_os = "linux", target_os = "macos"))] + fn restore_expectation(path: &Path) -> crate::restore::RestoreArtifactExpectation { + let metadata = fs::metadata(path).expect("restore artifact metadata"); + crate::restore::RestoreArtifactExpectation::new( + metadata.dev(), + metadata.ino(), + metadata.len(), + Sha256::digest(fs::read(path).expect("restore artifact bytes")).into(), + ) + .expect("restore artifact expectation") + } + + #[cfg(any(target_os = "linux", target_os = "macos"))] fn base_catalog() -> MigrationCatalog { MigrationCatalog::new([]).expect("empty v1 catalog") } @@ -2136,6 +2153,104 @@ mod tests { } #[cfg(any(target_os = "linux", target_os = "macos"))] + #[tokio::test(flavor = "current_thread")] + async fn writable_open_recovers_while_read_only_and_initialize_paths_do_not_mutate() { + let directory = tempfile::tempdir().expect("temporary directory"); + let policy = ServiceSqliteConnectionOptions::reviewed(); + let (paths, identity, mut authority) = + initialized_authority(directory.path(), "restore-recovery").await; + authority + .release() + .expect("release initialization authority"); + + let staged = paths + .state_database() + .with_file_name(crate::restore::STAGED_FILE_NAME); + fs::copy(paths.state_database(), &staged).expect("copy exact restore stage"); + fs::set_permissions(&staged, fs::Permissions::from_mode(0o600)) + .expect("restrict staged database"); + let authority = WriterAuthority::acquire(&paths, OpenMode::ReadWriteExisting) + .expect("writer authority") + .expect("writable mode retains authority"); + let marker = crate::restore::RestoreRecoveryMarker::prepared( + &database_metadata(&paths), + crate::BackupManifestSha256::from_bytes([23; 32]), + restore_expectation(paths.state_database()), + restore_expectation(&staged), + ) + .expect("prepared restore marker"); + crate::restore::RestoreMarkerBinding::create(&paths, &authority, &marker) + .expect("persist prepared marker"); + drop(authority); + + let state_directory = paths.state_database().parent().expect("state directory"); + let before = directory_snapshot(state_directory); + let read_only = open_existing_connection_pool( + &paths, + &identity, + &base_catalog(), + &base_schema_catalog(), + OpenMode::ReadOnlyInspection, + policy, + ) + .await; + let Err(error) = read_only else { + panic!("read-only open must not recover"); + }; + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert_eq!(directory_snapshot(state_directory), before); + + let initialize = crate::initialize_database( + &paths, + OpenMode::Initialize, + &database_metadata(&paths), + &base_schema_catalog(), + |_| async { Ok::<_, Infallible>(()) }, + ) + .await + .expect_err("initialize must not recover existing evidence"); + assert_eq!(initialize.kind(), ServiceSqliteErrorKind::Recovery); + assert_eq!(directory_snapshot(state_directory), before); + + let initialized_authority = WriterAuthority::acquire(&paths, OpenMode::ReadWriteExisting) + .expect("writer authority") + .expect("writable mode retains authority"); + let initialized = open_initialized_connection_pool( + &paths, + &identity, + &base_catalog(), + &base_schema_catalog(), + policy, + initialized_authority, + ) + .await; + let Err(error) = initialized else { + panic!("initialized open must not recover"); + }; + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert_eq!(directory_snapshot(state_directory), before); + + let writable = open_existing_connection_pool( + &paths, + &identity, + &base_catalog(), + &base_schema_catalog(), + OpenMode::ReadWriteExisting, + policy, + ) + .await + .expect("writable open recovers before SQLite"); + assert!(!staged.exists()); + assert!( + !paths + .state_database() + .with_file_name(crate::restore::MARKER_FILE_NAME) + .exists() + ); + drop(writable.close().await); + } + + #[cfg(any(target_os = "linux", target_os = "macos"))] #[test] fn connection_failure_precedence_is_exact() { assert_eq!( diff --git a/crates/service_sqlite/src/restore/marker.rs b/crates/service_sqlite/src/restore/marker.rs @@ -14,7 +14,8 @@ use serde::{Deserialize, Serialize}; use sha2::{Digest, Sha256}; use crate::{ - BackupManifestSha256, ServiceDatabaseMetadata, ServiceSqliteApplicationId, ServiceSqlitePaths, + BackupManifestSha256, ServiceDatabaseIdentity, ServiceDatabaseMetadata, + ServiceSqliteApplicationId, ServiceSqlitePaths, }; #[cfg(any(target_os = "linux", target_os = "macos"))] @@ -24,11 +25,11 @@ const RESTORE_MARKER_SCHEMA: &str = "radroots.service-sqlite.restore-marker"; const RESTORE_MARKER_SCHEMA_VERSION: u32 = 1; const RESTORE_MARKER_MAX_BYTES: usize = 2_048; const RESTORE_MARKER_CHECKSUM_DOMAIN: &[u8] = b"radroots.service_sqlite.restore_marker.v1\0"; -pub(super) const LIVE_FILE_NAME: &str = radroots_runtime_paths::SERVICE_STATE_DATABASE_FILE_NAME; -pub(super) const STAGED_FILE_NAME: &str = "state.restore-staged.sqlite"; -pub(super) const BACKUP_FILE_NAME: &str = "state.restore-backup.sqlite"; -pub(super) const MARKER_FILE_NAME: &str = "state.restore-marker.v1"; -pub(super) const MARKER_NEXT_FILE_NAME: &str = "state.restore-marker.v1.next"; +pub(crate) const LIVE_FILE_NAME: &str = radroots_runtime_paths::SERVICE_STATE_DATABASE_FILE_NAME; +pub(crate) const STAGED_FILE_NAME: &str = "state.restore-staged.sqlite"; +pub(crate) const BACKUP_FILE_NAME: &str = "state.restore-backup.sqlite"; +pub(crate) const MARKER_FILE_NAME: &str = "state.restore-marker.v1"; +pub(crate) const MARKER_NEXT_FILE_NAME: &str = "state.restore-marker.v1.next"; /// Exact retained identity and content expected for one restore artifact. #[derive(Clone, Copy, PartialEq, Eq)] @@ -371,6 +372,14 @@ impl RestoreRecoveryMarker { fn matches_paths(&self, paths: &ServiceSqlitePaths) -> bool { self.service == *paths.service() && self.instance == *paths.instance() } + + pub(crate) fn matches_identity(&self, identity: &ServiceDatabaseIdentity) -> bool { + self.service == *identity.service() + && self.instance == *identity.instance() + && self.source_generation == identity.source_generation() + && self.application_id == identity.application_id() + && self.state_schema_version <= identity.supported_state_schema_version() + } } impl fmt::Debug for RestoreRecoveryMarker { @@ -747,13 +756,94 @@ mod store { Ok(Some(binding)) } + pub(crate) fn load_for_recovery( + paths: &ServiceSqlitePaths, + authority: &WriterAuthority, + ) -> Result<Option<Self>, ServiceSqliteError> { + authority.validate_for(paths)?; + let layout = RestoreRecoveryLayout::for_paths(paths).map_err(recovery_contract)?; + let directory = authority_checked(authority, paths, || { + authority + .directory() + .try_clone() + .map_err(|_| StoreFailure::Directory) + })? + .map_err(recovery_store)?; + let directory_identity = + authority_checked(authority, paths, || validate_directory(&directory))? + .map_err(authority_store)?; + let marker_file = match authority_checked(authority, paths, || { + open_marker_file(&directory, MARKER_FILE_NAME) + })? { + Ok(file) => file, + Err(StoreFailure::Missing) => { + authority_checked(authority, paths, || { + require_absent(&directory, MARKER_NEXT_FILE_NAME) + })? + .map_err(recovery_store)?; + return Ok(None); + } + Err(cause) => return Err(recovery_store(cause)), + }; + let marker_identity = + authority_checked(authority, paths, || file_identity(&marker_file))? + .map_err(recovery_store)?; + let marker = authority_checked(authority, paths, || read_marker(&marker_file))? + .map_err(recovery_store)?; + if !marker.matches_paths(paths) { + return Err(recovery_contract( + RestoreMarkerContractError::InvalidIdentity, + )); + } + let binding = Self { + directory, + directory_identity, + marker_file, + marker_identity, + marker, + }; + authority_checked(authority, paths, || { + binding.validate_inner(paths, false, ServiceSqliteErrorKind::Recovery) + })??; + if layout + .marker + .file_name() + .is_none_or(|name| name != MARKER_FILE_NAME) + { + return Err(recovery_contract(RestoreMarkerContractError::InvalidLayout)); + } + authority.validate_for(paths)?; + Ok(Some(binding)) + } + pub(crate) fn advance( self, paths: &ServiceSqlitePaths, authority: &WriterAuthority, next: RestoreRecoveryPhase, ) -> Result<Self, ServiceSqliteError> { - self.advance_with_operations(paths, authority, next, &SystemStoreOperations) + self.advance_with_operations( + paths, + authority, + next, + &SystemStoreOperations, + ServiceSqliteErrorKind::Restore, + ) + } + + pub(crate) fn advance_for_recovery( + self, + paths: &ServiceSqlitePaths, + authority: &WriterAuthority, + next: RestoreRecoveryPhase, + ) -> Result<Self, ServiceSqliteError> { + self.advance_with_operations( + paths, + authority, + next, + &SystemStoreOperations, + ServiceSqliteErrorKind::Recovery, + ) } fn advance_with_operations( @@ -762,30 +852,33 @@ mod store { authority: &WriterAuthority, next: RestoreRecoveryPhase, operations: &dyn StoreOperations, + operation_kind: ServiceSqliteErrorKind, ) -> Result<Self, ServiceSqliteError> { - authority_checked(authority, paths, || self.validate_for_restore(paths))??; + authority_checked(authority, paths, || { + self.validate_inner(paths, true, operation_kind) + })??; let current = authority_checked(authority, paths, || read_marker(&self.marker_file))? - .map_err(restore_store)?; + .map_err(|cause| operation_store(operation_kind, cause))?; if current.canonical_bytes() != self.marker.canonical_bytes() { - return Err(restore_store(StoreFailure::Conflict)); + return Err(operation_store(operation_kind, StoreFailure::Conflict)); } let next_marker = self .marker .transitioned_to(next) - .map_err(restore_contract)?; + .map_err(|cause| operation_contract(operation_kind, cause))?; if next_marker.canonical_bytes() == self.marker.canonical_bytes() { return Ok(self); } authority_checked(authority, paths, || { require_absent(&self.directory, MARKER_NEXT_FILE_NAME) })? - .map_err(restore_store)?; + .map_err(|cause| operation_store(operation_kind, cause))?; let (scratch, scratch_identity) = authority_checked(authority, paths, || { let file = create_marker_file(&self.directory, MARKER_NEXT_FILE_NAME)?; let identity = file_identity(&file)?; Ok::<_, StoreFailure>((file, identity)) })? - .map_err(restore_store)?; + .map_err(|cause| operation_store(operation_kind, cause))?; let scratch_write = authority_checked(authority, paths, || { write_and_sync( &scratch, @@ -802,10 +895,10 @@ mod store { MARKER_NEXT_FILE_NAME, scratch_identity, )?; - return Err(restore_store(cause)); + return Err(operation_store(operation_kind, cause)); } authority_checked(authority, paths, || { - self.validate_inner(paths, false, ServiceSqliteErrorKind::Restore) + self.validate_inner(paths, false, operation_kind) })??; let scratch_matches = authority_checked(authority, paths, || { let current = open_marker_file(&self.directory, MARKER_NEXT_FILE_NAME)?; @@ -814,7 +907,7 @@ mod store { && file_identity(&current)? == scratch_identity, ) })? - .map_err(restore_store)?; + .map_err(|cause| operation_store(operation_kind, cause))?; if !scratch_matches { cleanup_with_authority( authority, @@ -823,7 +916,7 @@ mod store { MARKER_NEXT_FILE_NAME, scratch_identity, )?; - return Err(restore_store(StoreFailure::Conflict)); + return Err(operation_store(operation_kind, StoreFailure::Conflict)); } let replacement = authority_checked(authority, paths, || { operations @@ -838,7 +931,7 @@ mod store { MARKER_NEXT_FILE_NAME, scratch_identity, )?; - return Err(restore_store(StoreFailure::Rename)); + return Err(operation_store(operation_kind, StoreFailure::Rename)); } let parent_sync = authority_checked(authority, paths, || { operations @@ -846,7 +939,7 @@ mod store { .map_err(|_| StoreFailure::Sync) })?; if parent_sync.is_err() { - return Err(restore_store(StoreFailure::Sync)); + return Err(operation_store(operation_kind, StoreFailure::Sync)); } let (marker_file, marker_identity, reread) = authority_checked(authority, paths, || { @@ -855,9 +948,9 @@ mod store { let marker = read_marker(&file)?; Ok::<_, StoreFailure>((file, identity, marker)) })? - .map_err(restore_store)?; + .map_err(|cause| operation_store(operation_kind, cause))?; if reread.canonical_bytes() != next_marker.canonical_bytes() { - return Err(restore_store(StoreFailure::Conflict)); + return Err(operation_store(operation_kind, StoreFailure::Conflict)); } let binding = Self { directory: self.directory, @@ -866,7 +959,9 @@ mod store { marker_identity, marker: next_marker, }; - authority_checked(authority, paths, || binding.validate_for_restore(paths))??; + authority_checked(authority, paths, || { + binding.validate_inner(paths, true, operation_kind) + })??; Ok(binding) } @@ -883,6 +978,24 @@ mod store { authority, next, &FailingStoreOperations { failure }, + ServiceSqliteErrorKind::Restore, + ) + } + + #[cfg(test)] + pub(crate) fn test_advance_for_recovery_with_failure( + self, + paths: &ServiceSqlitePaths, + authority: &WriterAuthority, + next: RestoreRecoveryPhase, + failure: TestStoreFailure, + ) -> Result<Self, ServiceSqliteError> { + self.advance_with_operations( + paths, + authority, + next, + &FailingStoreOperations { failure }, + ServiceSqliteErrorKind::Recovery, ) } @@ -890,6 +1003,113 @@ mod store { &self.marker } + pub(crate) fn interrupted_transition( + &self, + paths: &ServiceSqlitePaths, + authority: &WriterAuthority, + ) -> Result<Option<RestoreRecoveryPhase>, ServiceSqliteError> { + authority_checked(authority, paths, || { + self.validate_inner(paths, false, ServiceSqliteErrorKind::Recovery) + })??; + let scratch = match authority_checked(authority, paths, || { + open_marker_file(&self.directory, MARKER_NEXT_FILE_NAME) + })? { + Ok(file) => file, + Err(StoreFailure::Missing) => return Ok(None), + Err(cause) => return Err(recovery_store(cause)), + }; + let scratch_marker = authority_checked(authority, paths, || read_marker(&scratch))? + .map_err(recovery_store)?; + let next = scratch_marker.phase(); + let expected = self.marker.transitioned_to(next); + if next == self.marker.phase() + || match expected { + Ok(expected) => expected.canonical_bytes() != scratch_marker.canonical_bytes(), + Err(_) => true, + } + { + return Err(recovery_store(StoreFailure::Conflict)); + } + Ok(Some(next)) + } + + pub(crate) fn promote_interrupted_transition( + self, + paths: &ServiceSqlitePaths, + authority: &WriterAuthority, + expected_phase: RestoreRecoveryPhase, + ) -> Result<Self, ServiceSqliteError> { + self.promote_interrupted_transition_with_hook(paths, authority, expected_phase, || {}) + } + + fn promote_interrupted_transition_with_hook( + self, + paths: &ServiceSqlitePaths, + authority: &WriterAuthority, + expected_phase: RestoreRecoveryPhase, + before_exact_removal: impl FnOnce(), + ) -> Result<Self, ServiceSqliteError> { + authority.validate_for(paths)?; + authority_checked(authority, paths, || { + self.validate_inner(paths, false, ServiceSqliteErrorKind::Recovery) + })??; + let scratch = authority_checked(authority, paths, || { + open_marker_file(&self.directory, MARKER_NEXT_FILE_NAME) + })? + .map_err(recovery_store)?; + let scratch_identity = authority_checked(authority, paths, || file_identity(&scratch))? + .map_err(recovery_store)?; + let scratch_marker = authority_checked(authority, paths, || read_marker(&scratch))? + .map_err(recovery_store)?; + let expected = self + .marker + .transitioned_to(expected_phase) + .map_err(recovery_contract)?; + if scratch_marker.canonical_bytes() != expected.canonical_bytes() { + return Err(recovery_store(StoreFailure::Conflict)); + } + before_exact_removal(); + // Preserve the valid current marker even if the scratch pathname + // was replaced after an interrupted advance. Remove only the + // exact validated scratch, then recreate the governed transition. + let removal = authority_checked(authority, paths, || { + remove_exact_and_sync(&self.directory, MARKER_NEXT_FILE_NAME, scratch_identity) + })?; + removal.map_err(recovery_store)?; + self.advance_for_recovery(paths, authority, expected_phase) + } + + #[cfg(test)] + pub(crate) fn test_promote_interrupted_transition_after_hook( + self, + paths: &ServiceSqlitePaths, + authority: &WriterAuthority, + expected_phase: RestoreRecoveryPhase, + before_exact_removal: impl FnOnce(), + ) -> Result<Self, ServiceSqliteError> { + self.promote_interrupted_transition_with_hook( + paths, + authority, + expected_phase, + before_exact_removal, + ) + } + + pub(crate) fn retire( + self, + paths: &ServiceSqlitePaths, + authority: &WriterAuthority, + ) -> Result<(), ServiceSqliteError> { + authority.validate_for(paths)?; + authority_checked(authority, paths, || { + self.validate_inner(paths, true, ServiceSqliteErrorKind::Recovery) + })??; + authority_checked(authority, paths, || { + remove_exact_and_sync(&self.directory, MARKER_FILE_NAME, self.marker_identity) + })? + .map_err(recovery_store) + } + pub(crate) fn validate( &self, paths: &ServiceSqlitePaths, @@ -1183,6 +1403,20 @@ mod store { } } + fn remove_exact_and_sync( + directory: &File, + name: &str, + expected: FileIdentity, + ) -> Result<(), StoreFailure> { + let current = open_marker_file(directory, name)?; + if file_identity(&current)? != expected { + return Err(StoreFailure::Conflict); + } + unlinkat(directory, name, AtFlags::empty()).map_err(|_| StoreFailure::Conflict)?; + directory.sync_all().map_err(|_| StoreFailure::Sync)?; + require_absent(directory, name) + } + #[derive(Clone, Copy, Debug, PartialEq, Eq)] enum StoreFailure { Directory, @@ -1253,6 +1487,8 @@ mod store { #[cfg(any(target_os = "linux", target_os = "macos"))] pub(crate) use store::RestoreMarkerBinding; +#[cfg(all(test, any(target_os = "linux", target_os = "macos")))] +pub(crate) use store::TestStoreFailure; #[cfg(test)] mod tests { diff --git a/crates/service_sqlite/src/restore/mod.rs b/crates/service_sqlite/src/restore/mod.rs @@ -2,37 +2,23 @@ mod finalize; mod marker; +mod recover; mod stage; pub use finalize::finalize_staged_restore; pub use stage::{StagedServiceRestore, stage_verified_restore}; #[cfg(any(target_os = "linux", target_os = "macos"))] -pub(crate) fn refuse_unresolved_recovery( - directory: &impl std::os::fd::AsFd, -) -> Result<(), crate::ServiceSqliteError> { - use rustix::{ - fs::{AtFlags, statat}, - io::Errno, - }; +pub(crate) use recover::refuse_unresolved_recovery; - for name in [marker::MARKER_FILE_NAME, marker::MARKER_NEXT_FILE_NAME] { - match statat(directory, name, AtFlags::SYMLINK_NOFOLLOW) { - Err(Errno::NOENT) => {} - Ok(_) | Err(_) => { - return Err(crate::ServiceSqliteError::new( - crate::ServiceSqliteErrorKind::Recovery, - )); - } - } - } - Ok(()) -} +#[cfg(any(target_os = "linux", target_os = "macos"))] +pub(crate) use recover::recover_for_open; #[allow(unused_imports)] pub(crate) use marker::{ + BACKUP_FILE_NAME, LIVE_FILE_NAME, MARKER_FILE_NAME, MARKER_NEXT_FILE_NAME, RestoreArtifactExpectation, RestoreMarkerContractError, RestoreRecoveryLayout, - RestoreRecoveryMarker, RestoreRecoveryPhase, + RestoreRecoveryMarker, RestoreRecoveryPhase, STAGED_FILE_NAME, }; #[cfg(any(target_os = "linux", target_os = "macos"))] diff --git a/crates/service_sqlite/src/restore/recover.rs b/crates/service_sqlite/src/restore/recover.rs @@ -0,0 +1,1052 @@ +//! Deterministic writer-authorized reconciliation of interrupted restores. + +#[cfg(any(target_os = "linux", target_os = "macos"))] +use core::fmt; + +#[cfg(any(target_os = "linux", target_os = "macos"))] +use { + super::{ + RestoreArtifactExpectation, RestoreMarkerBinding, RestoreRecoveryPhase, + marker::{ + BACKUP_FILE_NAME, LIVE_FILE_NAME, MARKER_FILE_NAME, MARKER_NEXT_FILE_NAME, + STAGED_FILE_NAME, + }, + }, + crate::{ + ServiceDatabaseIdentity, ServiceSqliteError, ServiceSqliteErrorKind, ServiceSqlitePaths, + WriterAuthority, + }, + rustix::{ + fs::{ + AtFlags, FileType, Mode, OFlags, RenameFlags, fstat, openat, renameat_with, statat, + unlinkat, + }, + io::Errno, + process::geteuid, + }, + sha2::{Digest, Sha256}, + std::{error::Error, fs::File, os::unix::fs::FileExt}, +}; + +#[cfg(any(target_os = "linux", target_os = "macos"))] +const HASH_BUFFER_BYTES: usize = 64 * 1_024; + +#[cfg(any(target_os = "linux", target_os = "macos"))] +pub(crate) fn recover_for_open( + paths: &ServiceSqlitePaths, + identity: &ServiceDatabaseIdentity, + authority: &WriterAuthority, +) -> Result<(), ServiceSqliteError> { + authority.validate_for(paths)?; + let Some(mut marker) = RestoreMarkerBinding::load_for_recovery(paths, authority)? else { + authority_checked(authority, paths, || { + refuse_unresolved_recovery(authority.directory()) + })?; + return Ok(()); + }; + if !marker.marker().matches_identity(identity) { + return Err(recovery_error(RecoveryFailureKind::Intent)); + } + + let observed = authority_checked(authority, paths, || { + observe_artifacts(authority.directory(), marker.marker()) + })?; + if let Some(next) = marker.interrupted_transition(paths, authority)? { + let topology_matches = match (marker.marker().phase(), next) { + (RestoreRecoveryPhase::Prepared, RestoreRecoveryPhase::LiveRetained) => { + observed.proves_live_retained() + } + (RestoreRecoveryPhase::LiveRetained, RestoreRecoveryPhase::ReplacementInstalled) => { + observed.proves_replacement_installed() + } + _ => false, + }; + if !topology_matches { + return Err(recovery_error(RecoveryFailureKind::Topology)); + } + marker = marker.promote_interrupted_transition(paths, authority, next)?; + } + + loop { + let observed = authority_checked(authority, paths, || { + observe_artifacts(authority.directory(), marker.marker()) + })?; + match marker.marker().phase() { + RestoreRecoveryPhase::Prepared if observed.can_roll_back_prepared() => { + if let Some(staged) = observed.staged.as_ref() { + authority_checked(authority, paths, || { + remove_exact_artifact( + authority.directory(), + STAGED_FILE_NAME, + staged, + marker.marker().staged(), + ) + })?; + } + marker.retire(paths, authority)?; + authority_checked(authority, paths, || { + refuse_unresolved_recovery(authority.directory()) + })?; + return Ok(()); + } + RestoreRecoveryPhase::Prepared if observed.proves_live_retained() => { + authority_checked(authority, paths, || { + authority.directory().sync_all().map_err(|source| { + recovery_source(RecoveryFailureKind::DirectorySync, source) + }) + })?; + marker = marker.advance_for_recovery( + paths, + authority, + RestoreRecoveryPhase::LiveRetained, + )?; + } + RestoreRecoveryPhase::LiveRetained if observed.needs_replacement_install() => { + let staged = observed + .staged + .as_ref() + .ok_or_else(|| recovery_error(RecoveryFailureKind::Topology))?; + authority_checked(authority, paths, || { + verify_named_artifact( + authority.directory(), + STAGED_FILE_NAME, + staged, + marker.marker().staged(), + ) + })?; + authority_checked(authority, paths, || { + renameat_with( + authority.directory(), + STAGED_FILE_NAME, + authority.directory(), + LIVE_FILE_NAME, + RenameFlags::NOREPLACE, + ) + .map_err(|source| { + recovery_source(RecoveryFailureKind::InstallReplacement, source) + }) + })?; + authority_checked(authority, paths, || { + verify_named_artifact( + authority.directory(), + LIVE_FILE_NAME, + staged, + marker.marker().staged(), + ) + })?; + authority_checked(authority, paths, || { + authority.directory().sync_all().map_err(|source| { + recovery_source(RecoveryFailureKind::DirectorySync, source) + }) + })?; + marker = marker.advance_for_recovery( + paths, + authority, + RestoreRecoveryPhase::ReplacementInstalled, + )?; + } + RestoreRecoveryPhase::LiveRetained if observed.proves_replacement_installed() => { + authority_checked(authority, paths, || { + authority.directory().sync_all().map_err(|source| { + recovery_source(RecoveryFailureKind::DirectorySync, source) + }) + })?; + marker = marker.advance_for_recovery( + paths, + authority, + RestoreRecoveryPhase::ReplacementInstalled, + )?; + } + RestoreRecoveryPhase::ReplacementInstalled + if observed.proves_replacement_installed_or_cleanup() => + { + let live = match observed.live { + LiveArtifact::Replacement(ref file) => file, + _ => return Err(recovery_error(RecoveryFailureKind::Topology)), + }; + authority_checked(authority, paths, || { + verify_named_artifact( + authority.directory(), + LIVE_FILE_NAME, + live, + observed.marker_staged, + ) + })?; + if let Some(backup) = observed.backup.as_ref() { + authority_checked(authority, paths, || { + remove_exact_artifact( + authority.directory(), + BACKUP_FILE_NAME, + backup, + observed.marker_live, + ) + })?; + } + marker.retire(paths, authority)?; + authority_checked(authority, paths, || { + refuse_unresolved_recovery(authority.directory()) + })?; + authority_checked(authority, paths, || { + verify_named_artifact( + authority.directory(), + LIVE_FILE_NAME, + live, + observed.marker_staged, + ) + })?; + return Ok(()); + } + _ => return Err(recovery_error(RecoveryFailureKind::Topology)), + } + } +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +pub(crate) fn refuse_unresolved_recovery( + directory: &impl std::os::fd::AsFd, +) -> Result<(), ServiceSqliteError> { + for name in [ + STAGED_FILE_NAME, + BACKUP_FILE_NAME, + MARKER_FILE_NAME, + MARKER_NEXT_FILE_NAME, + ] { + match statat(directory, name, AtFlags::SYMLINK_NOFOLLOW) { + Err(Errno::NOENT) => {} + Ok(_) | Err(_) => { + return Err(ServiceSqliteError::new(ServiceSqliteErrorKind::Recovery)); + } + } + } + Ok(()) +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +struct ObservedArtifacts { + live: LiveArtifact, + staged: Option<File>, + backup: Option<File>, + marker_live: RestoreArtifactExpectation, + marker_staged: RestoreArtifactExpectation, +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +impl ObservedArtifacts { + fn can_roll_back_prepared(&self) -> bool { + matches!(self.live, LiveArtifact::Original) && self.backup.is_none() + } + + fn proves_live_retained(&self) -> bool { + matches!(self.live, LiveArtifact::Absent) && self.staged.is_some() && self.backup.is_some() + } + + fn needs_replacement_install(&self) -> bool { + self.proves_live_retained() + } + + fn proves_replacement_installed(&self) -> bool { + matches!(self.live, LiveArtifact::Replacement(_)) + && self.staged.is_none() + && self.backup.is_some() + } + + fn proves_replacement_installed_or_cleanup(&self) -> bool { + matches!(self.live, LiveArtifact::Replacement(_)) && self.staged.is_none() + } +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +enum LiveArtifact { + Absent, + Original, + Replacement(File), +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn observe_artifacts( + directory: &File, + marker: &super::RestoreRecoveryMarker, +) -> Result<ObservedArtifacts, ServiceSqliteError> { + for sidecar in [ + "state.sqlite-wal", + "state.sqlite-shm", + "state.sqlite-journal", + ] { + require_absent(directory, sidecar)?; + } + let live = match open_optional(directory, LIVE_FILE_NAME)? { + None => LiveArtifact::Absent, + Some(file) if artifact_has_identity(&file, marker.live())? => { + verify_artifact(&file, marker.live())?; + LiveArtifact::Original + } + Some(file) if artifact_has_identity(&file, marker.staged())? => { + verify_artifact(&file, marker.staged())?; + LiveArtifact::Replacement(file) + } + Some(_) => return Err(recovery_error(RecoveryFailureKind::Artifact)), + }; + let staged = observe_expected(directory, STAGED_FILE_NAME, marker.staged())?; + let backup = observe_expected(directory, BACKUP_FILE_NAME, marker.backup())?; + Ok(ObservedArtifacts { + live, + staged, + backup, + marker_live: marker.live(), + marker_staged: marker.staged(), + }) +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn observe_expected( + directory: &File, + name: &str, + expected: RestoreArtifactExpectation, +) -> Result<Option<File>, ServiceSqliteError> { + let Some(file) = open_optional(directory, name)? else { + return Ok(None); + }; + verify_artifact(&file, expected)?; + Ok(Some(file)) +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn open_optional(directory: &File, name: &str) -> Result<Option<File>, ServiceSqliteError> { + match openat( + directory, + name, + OFlags::RDONLY | OFlags::NOFOLLOW | OFlags::CLOEXEC | OFlags::NONBLOCK, + Mode::empty(), + ) { + Ok(file) => Ok(Some(File::from(file))), + Err(Errno::NOENT) => Ok(None), + Err(source) => Err(recovery_source(RecoveryFailureKind::Artifact, source)), + } +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn artifact_has_identity( + file: &File, + expected: RestoreArtifactExpectation, +) -> Result<bool, ServiceSqliteError> { + let status = + fstat(file).map_err(|source| recovery_source(RecoveryFailureKind::Artifact, source))?; + let device = + u64::try_from(status.st_dev).map_err(|_| recovery_error(RecoveryFailureKind::Artifact))?; + Ok((device, status.st_ino) == (expected.device(), expected.inode())) +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn verify_artifact( + file: &File, + expected: RestoreArtifactExpectation, +) -> Result<(), ServiceSqliteError> { + let status = + fstat(file).map_err(|source| recovery_source(RecoveryFailureKind::Artifact, source))?; + let device = + u64::try_from(status.st_dev).map_err(|_| recovery_error(RecoveryFailureKind::Artifact))?; + let length = + u64::try_from(status.st_size).map_err(|_| recovery_error(RecoveryFailureKind::Artifact))?; + if !FileType::from_raw_mode(status.st_mode).is_file() + || u64::from(status.st_nlink) != 1 + || status.st_uid != geteuid().as_raw() + || u32::from(status.st_mode) & 0o777 != 0o600 + || (device, status.st_ino) != (expected.device(), expected.inode()) + || length != expected.byte_length() + || length == 0 + || length > i64::MAX as u64 + || hash_exact(file, expected.byte_length())? != expected.sha256() + { + return Err(recovery_error(RecoveryFailureKind::Artifact)); + } + Ok(()) +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn verify_named_artifact( + directory: &File, + name: &str, + held: &File, + expected: RestoreArtifactExpectation, +) -> Result<(), ServiceSqliteError> { + verify_artifact(held, expected)?; + let current = open_optional(directory, name)? + .ok_or_else(|| recovery_error(RecoveryFailureKind::Artifact))?; + verify_artifact(&current, expected) +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn remove_exact_artifact( + directory: &File, + name: &str, + held: &File, + expected: RestoreArtifactExpectation, +) -> Result<(), ServiceSqliteError> { + verify_named_artifact(directory, name, held, expected)?; + unlinkat(directory, name, AtFlags::empty()) + .map_err(|source| recovery_source(RecoveryFailureKind::Cleanup, source))?; + directory + .sync_all() + .map_err(|source| recovery_source(RecoveryFailureKind::DirectorySync, source))?; + require_absent(directory, name) +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn require_absent(directory: &File, name: &str) -> Result<(), ServiceSqliteError> { + match statat(directory, name, AtFlags::SYMLINK_NOFOLLOW) { + Err(Errno::NOENT) => Ok(()), + Ok(_) | Err(_) => Err(recovery_error(RecoveryFailureKind::Topology)), + } +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn hash_exact(file: &File, expected_length: u64) -> Result<[u8; 32], ServiceSqliteError> { + let mut hasher = Sha256::new(); + let mut buffer = [0_u8; HASH_BUFFER_BYTES]; + let mut offset = 0_u64; + while offset < expected_length { + let requested = usize::try_from((expected_length - offset).min(HASH_BUFFER_BYTES as u64)) + .map_err(|_| recovery_error(RecoveryFailureKind::Hash))?; + let read = file + .read_at(&mut buffer[..requested], offset) + .map_err(|source| recovery_source(RecoveryFailureKind::Hash, source))?; + if read == 0 { + return Err(recovery_error(RecoveryFailureKind::Hash)); + } + hasher.update(&buffer[..read]); + offset = offset + .checked_add( + u64::try_from(read).map_err(|_| recovery_error(RecoveryFailureKind::Hash))?, + ) + .ok_or_else(|| recovery_error(RecoveryFailureKind::Hash))?; + } + let mut extra = [0_u8; 1]; + if file + .read_at(&mut extra, expected_length) + .map_err(|source| recovery_source(RecoveryFailureKind::Hash, source))? + != 0 + { + return Err(recovery_error(RecoveryFailureKind::Hash)); + } + Ok(hasher.finalize().into()) +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn authority_checked<T>( + authority: &WriterAuthority, + paths: &ServiceSqlitePaths, + operation: impl FnOnce() -> Result<T, ServiceSqliteError>, +) -> Result<T, ServiceSqliteError> { + authority.validate_for(paths)?; + let result = operation(); + authority.validate_for(paths)?; + result +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +enum RecoveryFailureKind { + Intent, + Topology, + Artifact, + Hash, + InstallReplacement, + DirectorySync, + Cleanup, +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +struct RecoveryFailure { + kind: RecoveryFailureKind, + source: Option<Box<dyn Error + Send + Sync + 'static>>, +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +impl fmt::Debug for RecoveryFailure { + fn fmt(&self, formatter: &mut fmt::Formatter<'_>) -> fmt::Result { + formatter + .debug_struct("RecoveryFailure") + .field("kind", &self.kind) + .field("source", &self.source.as_ref().map(|_| "[redacted]")) + .finish() + } +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +impl fmt::Display for RecoveryFailure { + fn fmt(&self, formatter: &mut fmt::Formatter<'_>) -> fmt::Result { + formatter.write_str(match self.kind { + RecoveryFailureKind::Intent => "restore recovery intent does not match", + RecoveryFailureKind::Topology => "restore recovery topology is ambiguous", + RecoveryFailureKind::Artifact => "restore recovery artifact is invalid", + RecoveryFailureKind::Hash => "restore recovery artifact hash failed", + RecoveryFailureKind::InstallReplacement => { + "restore recovery replacement installation failed" + } + RecoveryFailureKind::DirectorySync => "restore recovery directory durability failed", + RecoveryFailureKind::Cleanup => "restore recovery cleanup failed", + }) + } +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +impl Error for RecoveryFailure { + fn source(&self) -> Option<&(dyn Error + 'static)> { + self.source + .as_deref() + .map(|source| source as &(dyn Error + 'static)) + } +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn recovery_error(kind: RecoveryFailureKind) -> ServiceSqliteError { + ServiceSqliteError::with_source( + ServiceSqliteErrorKind::Recovery, + RecoveryFailure { kind, source: None }, + ) +} + +#[cfg(any(target_os = "linux", target_os = "macos"))] +fn recovery_source( + kind: RecoveryFailureKind, + source: impl Error + Send + Sync + 'static, +) -> ServiceSqliteError { + ServiceSqliteError::with_source( + ServiceSqliteErrorKind::Recovery, + RecoveryFailure { + kind, + source: Some(Box::new(source)), + }, + ) +} + +#[cfg(all(test, any(target_os = "linux", target_os = "macos")))] +mod tests { + use std::{ + fs::{self, File, OpenOptions}, + io::Write, + num::NonZeroU32, + os::unix::fs::{FileTypeExt, OpenOptionsExt, PermissionsExt}, + path::{Path, PathBuf}, + process::Command, + }; + + use radroots_runtime_paths::{ + InstanceId, RadrootsHostEnvironment, RadrootsPathProfile, RadrootsPathResolver, + RadrootsPlatform, RuntimeContext, RuntimeContextBootstrap, RuntimeContextSource, ServiceId, + }; + use radroots_storage::event::SourceGeneration; + + use super::*; + use crate::restore::RestoreRecoveryMarker; + use crate::{ + BackupManifestSha256, OpenMode, ServiceDatabaseMetadata, ServiceSqliteApplicationId, + }; + + const OLD_BYTES: &[u8] = b"old-live-state"; + const NEW_BYTES: &[u8] = b"new-restored-state"; + + struct Fixture { + _root: tempfile::TempDir, + paths: ServiceSqlitePaths, + identity: ServiceDatabaseIdentity, + authority: WriterAuthority, + } + + impl Fixture { + fn new() -> Self { + let root = tempfile::tempdir().expect("temporary root"); + let paths = service_paths(root.path()); + let state_directory = paths.state_database().parent().expect("state directory"); + fs::create_dir_all(state_directory).expect("create state directory"); + fs::set_permissions(state_directory, fs::Permissions::from_mode(0o700)) + .expect("restrict state directory"); + write_new(paths.state_database(), OLD_BYTES); + write_new(&artifact_path(&paths, STAGED_FILE_NAME), NEW_BYTES); + let authority = WriterAuthority::acquire(&paths, OpenMode::ReadWriteExisting) + .expect("writer authority") + .expect("writable mode retains authority"); + let metadata = database_metadata(&paths); + let marker = RestoreRecoveryMarker::prepared( + &metadata, + BackupManifestSha256::from_bytes([19; 32]), + expectation(paths.state_database()), + expectation(&artifact_path(&paths, STAGED_FILE_NAME)), + ) + .expect("prepared marker"); + RestoreMarkerBinding::create(&paths, &authority, &marker) + .expect("persist prepared marker"); + Self { + _root: root, + paths, + identity: metadata.identity(), + authority, + } + } + + fn load_marker(&self) -> RestoreMarkerBinding { + RestoreMarkerBinding::load_for_recovery(&self.paths, &self.authority) + .expect("load marker") + .expect("marker exists") + } + + fn retain_live(&self, advance: bool) { + fs::rename( + self.paths.state_database(), + artifact_path(&self.paths, BACKUP_FILE_NAME), + ) + .expect("retain live"); + sync_state_directory(&self.paths); + if advance { + self.load_marker() + .advance( + &self.paths, + &self.authority, + RestoreRecoveryPhase::LiveRetained, + ) + .expect("advance live-retained marker"); + } + } + + fn install_stage(&self, advance: bool) { + fs::rename( + artifact_path(&self.paths, STAGED_FILE_NAME), + self.paths.state_database(), + ) + .expect("install stage"); + sync_state_directory(&self.paths); + if advance { + self.load_marker() + .advance( + &self.paths, + &self.authority, + RestoreRecoveryPhase::ReplacementInstalled, + ) + .expect("advance replacement marker"); + } + } + + fn write_scratch(&self, next: RestoreRecoveryPhase) { + let marker = self.load_marker(); + let next = marker + .marker() + .transitioned_to(next) + .expect("legal next phase"); + write_new( + &artifact_path(&self.paths, MARKER_NEXT_FILE_NAME), + next.canonical_bytes(), + ); + sync_state_directory(&self.paths); + } + + fn recover(&self) -> Result<(), ServiceSqliteError> { + recover_for_open(&self.paths, &self.identity, &self.authority) + } + } + + #[test] + fn prepared_with_old_live_rolls_back_and_retries_idempotently() { + let fixture = Fixture::new(); + fixture.recover().expect("rollback prepared restore"); + assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), OLD_BYTES); + assert_no_recovery_evidence(&fixture.paths); + fixture.recover().expect("repeated recovery is a no-op"); + } + + #[test] + fn interrupted_prepared_rollback_without_stage_finishes_marker_cleanup() { + let fixture = Fixture::new(); + fs::remove_file(artifact_path(&fixture.paths, STAGED_FILE_NAME)).expect("remove stage"); + sync_state_directory(&fixture.paths); + fixture.recover().expect("finish interrupted rollback"); + assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), OLD_BYTES); + assert_no_recovery_evidence(&fixture.paths); + } + + #[test] + fn prepared_with_proven_first_rename_rolls_forward() { + let fixture = Fixture::new(); + fixture.retain_live(false); + fixture.recover().expect("recover after first rename"); + assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), NEW_BYTES); + assert_no_recovery_evidence(&fixture.paths); + } + + #[test] + fn live_retained_with_proven_second_rename_rolls_forward() { + let fixture = Fixture::new(); + fixture.retain_live(true); + fixture.install_stage(false); + fixture.recover().expect("recover after second rename"); + assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), NEW_BYTES); + assert_no_recovery_evidence(&fixture.paths); + } + + #[test] + fn replacement_installed_retires_backup_and_marker() { + let fixture = Fixture::new(); + fixture.retain_live(true); + fixture.install_stage(true); + fixture.recover().expect("retire completed recovery"); + assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), NEW_BYTES); + assert_no_recovery_evidence(&fixture.paths); + } + + #[test] + fn replacement_cleanup_without_backup_finishes_idempotently() { + let fixture = Fixture::new(); + fixture.retain_live(true); + fixture.install_stage(true); + fs::remove_file(artifact_path(&fixture.paths, BACKUP_FILE_NAME)).expect("remove backup"); + sync_state_directory(&fixture.paths); + fixture.recover().expect("finish interrupted cleanup"); + assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), NEW_BYTES); + assert_no_recovery_evidence(&fixture.paths); + fixture.recover().expect("repeated cleanup is a no-op"); + } + + #[test] + fn topology_consistent_marker_scratch_is_promoted_at_both_edges() { + let first = Fixture::new(); + first.retain_live(false); + first.write_scratch(RestoreRecoveryPhase::LiveRetained); + first.recover().expect("promote first scratch"); + assert_eq!(fs::read(first.paths.state_database()).unwrap(), NEW_BYTES); + assert_no_recovery_evidence(&first.paths); + + let second = Fixture::new(); + second.retain_live(true); + second.install_stage(false); + second.write_scratch(RestoreRecoveryPhase::ReplacementInstalled); + second.recover().expect("promote second scratch"); + assert_eq!(fs::read(second.paths.state_database()).unwrap(), NEW_BYTES); + assert_no_recovery_evidence(&second.paths); + } + + #[test] + fn inferred_advance_failures_are_recovery_and_authority_keeps_precedence() { + let recovery = Fixture::new(); + let error = recovery + .load_marker() + .test_advance_for_recovery_with_failure( + &recovery.paths, + &recovery.authority, + RestoreRecoveryPhase::LiveRetained, + crate::restore::marker::TestStoreFailure::ScratchSync, + ) + .expect_err("recovery marker sync failure must be classified"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert!(artifact_path(&recovery.paths, MARKER_FILE_NAME).exists()); + assert!(!artifact_path(&recovery.paths, MARKER_NEXT_FILE_NAME).exists()); + + let authority = Fixture::new(); + let state_directory = authority.paths.state_database().parent().unwrap(); + let error = authority + .load_marker() + .test_advance_for_recovery_with_failure( + &authority.paths, + &authority.authority, + RestoreRecoveryPhase::LiveRetained, + crate::restore::marker::TestStoreFailure::AuthorityDriftAndScratchSync, + ) + .expect_err("authority drift must dominate the marker failure"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Authority); + assert!(artifact_path(&authority.paths, MARKER_FILE_NAME).exists()); + fs::set_permissions(state_directory, fs::Permissions::from_mode(0o700)) + .expect("restore state directory mode"); + } + + #[test] + fn replaced_interrupted_scratch_never_overwrites_the_valid_marker() { + let fixture = Fixture::new(); + fixture.retain_live(false); + fixture.write_scratch(RestoreRecoveryPhase::LiveRetained); + let marker = fixture.load_marker(); + let scratch = artifact_path(&fixture.paths, MARKER_NEXT_FILE_NAME); + let retained = scratch.with_file_name("retained-valid-marker-next"); + let replacement_bytes = b"foreign-marker-next"; + let scratch_for_hook = scratch.clone(); + let error = marker + .test_promote_interrupted_transition_after_hook( + &fixture.paths, + &fixture.authority, + RestoreRecoveryPhase::LiveRetained, + move || { + fs::rename(&scratch_for_hook, &retained).expect("retain exact scratch"); + write_new(&scratch_for_hook, replacement_bytes); + }, + ) + .expect_err("replaced scratch must fail without marker replacement"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert_eq!(fs::read(&scratch).unwrap(), replacement_bytes); + assert!( + scratch + .with_file_name("retained-valid-marker-next") + .exists() + ); + let current = fixture.load_marker(); + assert_eq!(current.marker().phase(), RestoreRecoveryPhase::Prepared); + } + + #[test] + fn topology_inconsistent_scratch_and_tampered_artifacts_fail_without_cleanup() { + let scratch = Fixture::new(); + scratch.write_scratch(RestoreRecoveryPhase::LiveRetained); + let error = scratch + .recover() + .expect_err("scratch before the first rename is ambiguous"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert!(artifact_path(&scratch.paths, MARKER_FILE_NAME).exists()); + assert!(artifact_path(&scratch.paths, MARKER_NEXT_FILE_NAME).exists()); + assert!(artifact_path(&scratch.paths, STAGED_FILE_NAME).exists()); + + let tampered = Fixture::new(); + fs::write( + artifact_path(&tampered.paths, STAGED_FILE_NAME), + b"tampered-restored", + ) + .expect("tamper stage"); + let before = fs::read_dir( + tampered + .paths + .state_database() + .parent() + .expect("state directory"), + ) + .unwrap() + .count(); + let error = tampered + .recover() + .expect_err("tampered stage must fail closed"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert_eq!( + fs::read_dir( + tampered + .paths + .state_database() + .parent() + .expect("state directory") + ) + .unwrap() + .count(), + before + ); + assert!(artifact_path(&tampered.paths, MARKER_FILE_NAME).exists()); + } + + #[test] + fn sidecar_mode_and_identity_mismatch_preserve_recovery_evidence() { + let sidecar = Fixture::new(); + write_new(&artifact_path(&sidecar.paths, "state.sqlite-wal"), b"wal"); + let error = sidecar + .recover() + .expect_err("restore recovery rejects sidecars"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert!(artifact_path(&sidecar.paths, MARKER_FILE_NAME).exists()); + + let mode = Fixture::new(); + fs::set_permissions( + artifact_path(&mode.paths, STAGED_FILE_NAME), + fs::Permissions::from_mode(0o640), + ) + .expect("change stage mode"); + let error = mode + .recover() + .expect_err("insecure artifact mode fails closed"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert!(artifact_path(&mode.paths, MARKER_FILE_NAME).exists()); + + let mismatch = Fixture::new(); + let wrong = ServiceDatabaseIdentity::new( + &mismatch.paths, + SourceGeneration::new([99; 32]).expect("different generation"), + NonZeroU32::new(1).expect("schema version"), + mismatch.identity.application_id(), + ); + let error = recover_for_open(&mismatch.paths, &wrong, &mismatch.authority) + .expect_err("marker intent mismatch fails closed"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert!(artifact_path(&mismatch.paths, MARKER_FILE_NAME).exists()); + } + + #[test] + fn symlink_hardlink_fifo_and_foreign_replacement_fail_without_deletion() { + use std::os::unix::fs::symlink; + + let symlinked = Fixture::new(); + let staged = artifact_path(&symlinked.paths, STAGED_FILE_NAME); + let held = staged.with_file_name("held-stage"); + fs::rename(&staged, &held).expect("retain original stage"); + symlink(symlinked.paths.state_database(), &staged).expect("replace with symlink"); + let error = symlinked + .recover() + .expect_err("symlink replacement fails closed"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert!( + fs::symlink_metadata(&staged) + .unwrap() + .file_type() + .is_symlink() + ); + assert!(held.exists()); + + let hardlinked = Fixture::new(); + let staged = artifact_path(&hardlinked.paths, STAGED_FILE_NAME); + let link = staged.with_file_name("stage-hardlink"); + fs::hard_link(&staged, &link).expect("create hard link"); + let error = hardlinked + .recover() + .expect_err("multiple-link stage fails closed"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert!(staged.exists()); + assert!(link.exists()); + + let fifo = Fixture::new(); + let staged = artifact_path(&fifo.paths, STAGED_FILE_NAME); + fs::remove_file(&staged).expect("remove original stage"); + assert!( + Command::new("mkfifo") + .arg(&staged) + .status() + .expect("run mkfifo") + .success() + ); + fs::set_permissions(&staged, fs::Permissions::from_mode(0o600)).expect("restrict FIFO"); + let error = fifo + .recover() + .expect_err("nonblocking FIFO admission fails closed"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert!(fs::symlink_metadata(&staged).unwrap().file_type().is_fifo()); + + let replaced = Fixture::new(); + let staged = artifact_path(&replaced.paths, STAGED_FILE_NAME); + let original = staged.with_file_name("original-stage"); + fs::rename(&staged, &original).expect("retain original stage"); + write_new(&staged, b"foreign-stage"); + let foreign = fs::read(&staged).unwrap(); + let error = replaced + .recover() + .expect_err("foreign replacement fails closed"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert_eq!(fs::read(&staged).unwrap(), foreign); + assert_eq!(fs::read(&original).unwrap(), NEW_BYTES); + assert!(artifact_path(&replaced.paths, MARKER_FILE_NAME).exists()); + } + + #[test] + fn authority_drift_precedes_recovery_and_keeps_evidence() { + let fixture = Fixture::new(); + let state_directory = fixture.paths.state_database().parent().unwrap(); + fs::set_permissions(state_directory, fs::Permissions::from_mode(0o770)) + .expect("drift state directory mode"); + let error = fixture + .recover() + .expect_err("authority drift must stop recovery"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Authority); + assert!(artifact_path(&fixture.paths, MARKER_FILE_NAME).exists()); + assert!(artifact_path(&fixture.paths, STAGED_FILE_NAME).exists()); + fs::set_permissions(state_directory, fs::Permissions::from_mode(0o700)) + .expect("restore state directory mode"); + } + + #[test] + fn orphan_stage_backup_or_scratch_without_marker_is_never_repaired() { + for name in [STAGED_FILE_NAME, BACKUP_FILE_NAME, MARKER_NEXT_FILE_NAME] { + let fixture = Fixture::new(); + fs::remove_file(artifact_path(&fixture.paths, MARKER_FILE_NAME)) + .expect("remove marker"); + if name != STAGED_FILE_NAME { + fs::remove_file(artifact_path(&fixture.paths, STAGED_FILE_NAME)) + .expect("remove default stage"); + write_new(&artifact_path(&fixture.paths, name), b"orphan"); + } + sync_state_directory(&fixture.paths); + let error = fixture + .recover() + .expect_err("orphan recovery evidence must fail closed"); + assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); + assert!(artifact_path(&fixture.paths, name).exists()); + } + } + + fn service_paths(root: &Path) -> ServiceSqlitePaths { + let context = RuntimeContext::resolve( + &RadrootsPathResolver::new(RadrootsPlatform::Linux, RadrootsHostEnvironment::default()), + RuntimeContextBootstrap::new( + RadrootsPathProfile::RepoLocal, + Some(root.to_path_buf()), + RuntimeContextSource::BootstrapCli, + RuntimeContextSource::BootstrapCli, + ) + .expect("bootstrap"), + ServiceId::new("myc").expect("service"), + InstanceId::new("recovery").expect("instance"), + ) + .expect("runtime context"); + ServiceSqlitePaths::from_runtime_context(&context).expect("SQLite paths") + } + + fn database_metadata(paths: &ServiceSqlitePaths) -> ServiceDatabaseMetadata { + ServiceDatabaseMetadata::new( + paths, + SourceGeneration::new([7; 32]).expect("source generation"), + NonZeroU32::new(1).expect("schema version"), + 1_700_000_000_000, + ServiceSqliteApplicationId::new(0x5244_5351).expect("application ID"), + ) + .expect("metadata") + } + + fn write_new(path: &Path, bytes: &[u8]) -> File { + let mut file = OpenOptions::new() + .read(true) + .write(true) + .create_new(true) + .mode(0o600) + .open(path) + .expect("create artifact"); + file.set_permissions(fs::Permissions::from_mode(0o600)) + .expect("set artifact mode"); + file.write_all(bytes).expect("write artifact"); + file.sync_all().expect("sync artifact"); + file + } + + fn expectation(path: &Path) -> RestoreArtifactExpectation { + let file = File::open(path).expect("open artifact"); + let status = fstat(&file).expect("artifact status"); + RestoreArtifactExpectation::new( + u64::try_from(status.st_dev).expect("device"), + status.st_ino, + u64::try_from(status.st_size).expect("length"), + hash_exact( + &file, + u64::try_from(status.st_size).expect("positive artifact length"), + ) + .expect("artifact digest"), + ) + .expect("artifact expectation") + } + + fn artifact_path(paths: &ServiceSqlitePaths, name: &str) -> PathBuf { + paths.state_database().with_file_name(name) + } + + fn sync_state_directory(paths: &ServiceSqlitePaths) { + File::open(paths.state_database().parent().expect("state directory")) + .expect("open state directory") + .sync_all() + .expect("sync state directory"); + } + + fn assert_no_recovery_evidence(paths: &ServiceSqlitePaths) { + for name in [ + STAGED_FILE_NAME, + BACKUP_FILE_NAME, + MARKER_FILE_NAME, + MARKER_NEXT_FILE_NAME, + ] { + assert!(!artifact_path(paths, name).exists(), "unexpected {name}"); + } + } +} diff --git a/crates/service_sqlite/src/restore/stage.rs b/crates/service_sqlite/src/restore/stage.rs @@ -2056,20 +2056,18 @@ mod tests { ) ); - for mode in [OpenMode::ReadWriteExisting, OpenMode::ReadOnlyInspection] { - let error = crate::open::open_existing_connection_pool( - &fixture.paths, - &fixture.identity, - &fixture.migrations, - &fixture.schema, - mode, - ServiceSqliteConnectionOptions::reviewed(), - ) - .await - .err() - .expect("unresolved recovery must reject open"); - assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery); - } + let read_only = crate::open::open_existing_connection_pool( + &fixture.paths, + &fixture.identity, + &fixture.migrations, + &fixture.schema, + OpenMode::ReadOnlyInspection, + ServiceSqliteConnectionOptions::reviewed(), + ) + .await + .err() + .expect("read-only open must not recover"); + assert_eq!(read_only.kind(), ServiceSqliteErrorKind::Recovery); let initialize = crate::initialize_database( &fixture.paths, OpenMode::Initialize, @@ -2080,6 +2078,30 @@ mod tests { .await .expect_err("unresolved recovery must reject initialization"); assert_eq!(initialize.kind(), ServiceSqliteErrorKind::Recovery); + + let writable = crate::open::open_existing_connection_pool( + &fixture.paths, + &fixture.identity, + &fixture.migrations, + &fixture.schema, + OpenMode::ReadWriteExisting, + ServiceSqliteConnectionOptions::reviewed(), + ) + .await + .expect("writable open must recover exact finalization evidence"); + writable.close().await.expect("close recovered pool"); + assert_eq!( + fs::read(fixture.paths.state_database()).expect("recovered live database"), + replacement + ); + for name in [ + BACKUP_FILE_NAME, + MARKER_FILE_NAME, + MARKER_NEXT_FILE_NAME, + STAGED_FILE_NAME, + ] { + assert!(!recovery_path(&fixture, name).exists(), "unexpected {name}"); + } reset_finalize_controls(); } diff --git a/crates/service_sqlite/tests/package_boundary.rs b/crates/service_sqlite/tests/package_boundary.rs @@ -18,6 +18,7 @@ const MIGRATION_SOURCE: &str = include_str!("../src/migration.rs"); const OPEN_SOURCE: &str = include_str!("../src/open.rs"); const RESTORE_MARKER_SOURCE: &str = include_str!("../src/restore/marker.rs"); const RESTORE_FINALIZE_SOURCE: &str = include_str!("../src/restore/finalize.rs"); +const RESTORE_RECOVER_SOURCE: &str = include_str!("../src/restore/recover.rs"); const RESTORE_ROOT_SOURCE: &str = include_str!("../src/restore/mod.rs"); const RESTORE_STAGE_SOURCE: &str = include_str!("../src/restore/stage.rs"); const STATUS_SOURCE: &str = include_str!("../src/status.rs"); @@ -179,9 +180,22 @@ fn service_sqlite_is_unpublished_lint_governed_and_dependency_bounded() { "Once `prepared` is durable, staged-artifact cleanup is disarmed", "unknown immediate outcome", "retains the old live database and final marker", - "every initialized, read-write-existing, or read-only-inspection pool open refuses", - "marker or marker-scratch as `Recovery` before opening SQLite", - "does not delete evidence, roll back, reconcile an interruption, or reopen the database", + "Read-write-existing open is the sole recovery path", + "before opening SQLite", + "Read-only inspection, initialization, and an initialized open never recover", + "reject any stage, backup, marker, or marker scratch as `Recovery` without mutation", + "Recovery uses exact topology as the durable authority", + "rolls back by removing only the exact stage and then the marker", + "recovery advances and rolls forward", + "removes the exact old backup before retiring the marker", + "repeated recovery is idempotent", + "A marker scratch is admitted only when it is the canonical one-edge successor", + "Recovery removes only the exact bound scratch inode", + "then reproduces the transition through the governed marker-advance path", + "preserves the valid current marker if the scratch pathname was replaced", + "Recovery has no await point or hidden task", + "each synchronous filesystem step and its authority checks complete", + "Finalization itself does not reconcile or reopen the database", ] { assert!( readme_words.contains(required), @@ -675,18 +689,73 @@ fn service_sqlite_is_unpublished_lint_governed_and_dependency_bounded() { "Step 068 restore finalization source contains deferred or raw authority `{forbidden}`" ); } + for required in ["refuse_unresolved_recovery", "recover_for_open"] { + assert!( + RESTORE_ROOT_SOURCE.contains(required), + "Step 068 recovery-open guard is missing `{required}`" + ); + } + assert!(OPEN_SOURCE.contains("crate::restore::refuse_unresolved_recovery")); + assert!(OPEN_SOURCE.contains("crate::restore::recover_for_open(paths, identity, authority)")); + + let restore_recover_production = RESTORE_RECOVER_SOURCE + .split_once("#[cfg(all(test, any(target_os = \"linux\", target_os = \"macos\")))]") + .map(|(production, _)| production) + .expect("restore recovery source must keep tests separated"); for required in [ - "refuse_unresolved_recovery", + "pub(crate) fn recover_for_open(", + "WriterAuthority", + "RestoreMarkerBinding::load_for_recovery", + "matches_identity(identity)", + "interrupted_transition(paths, authority)", + "promote_interrupted_transition", + "advance_for_recovery", + "RestoreRecoveryPhase::Prepared", + "RestoreRecoveryPhase::LiveRetained", + "RestoreRecoveryPhase::ReplacementInstalled", + "RenameFlags::NOREPLACE", + "verify_named_artifact(", + "hash_exact(", + "remove_exact_artifact(", + "marker.retire(paths, authority)", + "state.sqlite-wal", + "state.sqlite-shm", + "state.sqlite-journal", "ServiceSqliteErrorKind::Recovery", + ] { + assert!( + restore_recover_production.contains(required), + "Step 069 restore recovery source is missing `{required}`" + ); + } + for forbidden in [ + "pub fn recover", + "pub async fn recover", + "pub struct PendingRestore", + "sqlx::", + "rusqlite::", + "tokio::", + "spawn_blocking", + "tokio::time::timeout", + "remove_dir_all", + "SystemTime::now", + ] { + assert!( + !restore_recover_production.contains(forbidden) && !ROOT.contains(forbidden), + "Step 069 recovery exposes or implements forbidden authority `{forbidden}`" + ); + } + for required in [ + "STAGED_FILE_NAME", + "BACKUP_FILE_NAME", "MARKER_FILE_NAME", "MARKER_NEXT_FILE_NAME", ] { assert!( - RESTORE_ROOT_SOURCE.contains(required), - "Step 068 recovery-open guard is missing `{required}`" + RESTORE_RECOVER_SOURCE.contains(required), + "Step 069 refusal inventory is missing `{required}`" ); } - assert!(OPEN_SOURCE.contains("crate::restore::refuse_unresolved_recovery")); assert!(ROOT.contains( "pub use restore::{StagedServiceRestore, finalize_staged_restore, stage_verified_restore};" ));