commit 3225d43c15f5124ce7e990aaeb6817b1713b847b
parent 18620a5d289190e36353e32643587d2636d9c34e
Author: triesap <tyson@radroots.org>
Date: Tue, 11 Aug 2026 23:38:45 +0000
service-sqlite: recover interrupted restore
- reconcile exact marker topologies under writer authority
- preserve error precedence and valid markers across scratch drift
- keep non-writable opens fail-closed and mutation-free
- verify idempotent recovery across native and cross-target gates
Diffstat:
9 files changed, 1605 insertions(+), 74 deletions(-)
diff --git a/AGENTS.md b/AGENTS.md
@@ -228,10 +228,29 @@ Before editing code:
cleanup must remain disarmed after every later error so the marker never
loses a bound artifact.
Successful finalization returns no host and leaves the old live database and
- `replacement_installed` marker for open-time recovery. Until that recovery
- exists, every pool open must reject marker or marker-scratch evidence as
- `Recovery`. Finalization must not roll back, delete recovery evidence, reopen
- SQLite, or expose paths, descriptors, marker controls, or rename controls.
+ `replacement_installed` marker for the next writable open to recover. Other
+ open modes reject that evidence as `Recovery`. Finalization itself must not
+ roll back, delete recovery evidence, reopen SQLite, or expose paths,
+ descriptors, marker controls, or rename controls.
+- Interrupted restore recovery is private and automatic only for
+ read-write-existing open under exclusive `WriterAuthority`, before any
+ SQLite connection or await point. Initialize, initialized-open, and
+ read-only inspection must reject every fixed stage, backup, marker, or
+ marker-scratch artifact without mutation. Recovery must bind the marker to
+ the requested database identity, hash and revalidate exact owner-only
+ single-link artifacts, reject sidecars, and let exact topology decide the
+ sole action: roll back `prepared` while old live is still installed, then
+ roll forward once old live is durably retained. Persist every inferred phase
+ before the next destructive step; retire exact backup before marker; and
+ admit marker scratch only as a canonical topology-consistent one-edge
+ successor whose exact bound inode is removed and durably reproduced through
+ the governed marker-advance path without overwriting the valid marker.
+ Repeated recovery may
+ finish already-absent stage or backup cleanup, but every other missing,
+ replaced, linked, malformed, mismatched, or ambiguous artifact remains
+ `Recovery` evidence. Do not expose recovery controls, add a background task
+ or hidden timeout, repair without writer authority, or fold Step 070
+ integrity/status APIs and the later process failpoint harness into recovery.
- Runtime-management flows consume a sealed `RuntimeContext` for every service
instance. They must not reconstruct service paths from raw identifiers,
ambient selectors, or manager-owned roots, and registries must not persist
diff --git a/crates/service_sqlite/README.md b/crates/service_sqlite/README.md
@@ -167,11 +167,37 @@ authority until it either fails before durability or establishes recovery
evidence and continues. Once `prepared` is durable, staged-artifact cleanup is
disarmed and the bound stage remains available after every later error.
Success returns no database host, retains the old live database and final
-marker, and requires a new open. Until open-time recovery is implemented,
-every initialized, read-write-existing, or read-only-inspection pool open
-refuses a marker or marker-scratch as `Recovery` before opening SQLite.
-Finalization does not delete evidence, roll back, reconcile an interruption,
-or reopen the database; those remain the next recovery checkpoint.
+marker, and requires a new open. Read-write-existing open is the sole recovery
+path. Under exclusive writer authority and before opening SQLite, it validates
+the marker's exact service, instance, source generation, application ID, schema
+ceiling, artifact identities, lengths, digests, restrictive modes, and the
+absence of database sidecars. Read-only inspection, initialization, and an
+initialized open never recover; they reject any stage, backup, marker, or marker
+scratch as `Recovery` without mutation.
+
+Recovery uses exact topology as the durable authority. `prepared` with the old
+live database still installed rolls back by removing only the exact stage and
+then the marker. Once the exact old live inode has reached the backup name,
+recovery advances and rolls forward. A lagging `live_retained` phase installs or
+recognizes the exact replacement, advances to `replacement_installed`, then
+removes the exact old backup before retiring the marker. Interrupted rollback
+and final cleanup accept only the corresponding already-absent exact artifact,
+so repeated recovery is idempotent. Every other topology, sidecar, replacement,
+link, mode, owner, length, digest, identity, or directory-authority mismatch
+fails closed and preserves the evidence.
+
+A marker scratch is admitted only when it is the canonical one-edge successor
+of the current marker and the artifact topology already proves that successor.
+Recovery removes only the exact bound scratch inode, synchronizes that removal,
+then reproduces the transition through the governed marker-advance path. This
+preserves the valid current marker if the scratch pathname was replaced.
+Orphaned, malformed, skipped, same-phase, terminal, mismatched, or
+topology-inconsistent scratch is never deleted or reinterpreted. Recovery has no
+await point or hidden task: once a writable open is polled, each synchronous
+filesystem step and its authority checks complete before the open can be
+cancelled. A later cancelled SQLite open is retried by rereading the already
+durable, marker-free state. Finalization itself does not reconcile or reopen the
+database.
The crate owns mechanics only. Service-specific tables, SQL, repositories,
backup content policy, identity material, process lifecycle, and readiness
diff --git a/crates/service_sqlite/src/connection.rs b/crates/service_sqlite/src/connection.rs
@@ -59,6 +59,12 @@ enum ServiceSqliteHostCloseState {
impl ServiceSqliteHost {
/// Opens existing writable state and finishes every pending governed migration.
+ ///
+ /// Before opening SQLite, this path holds exclusive writer authority and
+ /// synchronously reconciles any exact interrupted-restore topology. The
+ /// recovery sequence has no await point: cancellation cannot split one
+ /// filesystem step from its authority check. If the surrounding open is
+ /// cancelled later, a retry re-reads the already durable filesystem state.
#[allow(clippy::too_many_arguments)]
pub async fn open_read_write_existing(
paths: &ServiceSqlitePaths,
diff --git a/crates/service_sqlite/src/open.rs b/crates/service_sqlite/src/open.rs
@@ -675,12 +675,15 @@ async fn open_connection_pool(
return Err(ServiceSqliteError::new(ServiceSqliteErrorKind::Integrity));
}
let recovery_guard = match (mode, authority.as_ref(), inspection_guard.as_ref()) {
- (OpenMode::Initialize | OpenMode::ReadWriteExisting, Some(authority), None) => {
+ (OpenMode::Initialize, Some(authority), None) => {
authority.validate_for(paths)?;
let result = crate::restore::refuse_unresolved_recovery(authority.directory());
authority.validate_for(paths)?;
result
}
+ (OpenMode::ReadWriteExisting, Some(authority), None) => {
+ crate::restore::recover_for_open(paths, identity, authority)
+ }
(OpenMode::ReadOnlyInspection, None, Some(inspection_guard)) => {
inspection_guard.validate_for(paths)?;
let result = crate::restore::refuse_unresolved_recovery(&inspection_guard.directory);
@@ -1589,7 +1592,7 @@ mod tests {
convert::Infallible,
fs,
num::NonZeroU32,
- os::unix::fs::{PermissionsExt, symlink},
+ os::unix::fs::{MetadataExt, PermissionsExt, symlink},
sync::{
Arc,
atomic::{AtomicUsize, Ordering as AtomicOrdering},
@@ -1603,6 +1606,8 @@ mod tests {
};
#[cfg(any(target_os = "linux", target_os = "macos"))]
use radroots_storage::event::SourceGeneration;
+ #[cfg(any(target_os = "linux", target_os = "macos"))]
+ use sha2::{Digest, Sha256};
#[cfg(any(target_os = "linux", target_os = "macos"))]
use crate::{ServiceDatabaseMetadata, ServiceSqliteApplicationId};
@@ -1652,6 +1657,18 @@ mod tests {
}
#[cfg(any(target_os = "linux", target_os = "macos"))]
+ fn restore_expectation(path: &Path) -> crate::restore::RestoreArtifactExpectation {
+ let metadata = fs::metadata(path).expect("restore artifact metadata");
+ crate::restore::RestoreArtifactExpectation::new(
+ metadata.dev(),
+ metadata.ino(),
+ metadata.len(),
+ Sha256::digest(fs::read(path).expect("restore artifact bytes")).into(),
+ )
+ .expect("restore artifact expectation")
+ }
+
+ #[cfg(any(target_os = "linux", target_os = "macos"))]
fn base_catalog() -> MigrationCatalog {
MigrationCatalog::new([]).expect("empty v1 catalog")
}
@@ -2136,6 +2153,104 @@ mod tests {
}
#[cfg(any(target_os = "linux", target_os = "macos"))]
+ #[tokio::test(flavor = "current_thread")]
+ async fn writable_open_recovers_while_read_only_and_initialize_paths_do_not_mutate() {
+ let directory = tempfile::tempdir().expect("temporary directory");
+ let policy = ServiceSqliteConnectionOptions::reviewed();
+ let (paths, identity, mut authority) =
+ initialized_authority(directory.path(), "restore-recovery").await;
+ authority
+ .release()
+ .expect("release initialization authority");
+
+ let staged = paths
+ .state_database()
+ .with_file_name(crate::restore::STAGED_FILE_NAME);
+ fs::copy(paths.state_database(), &staged).expect("copy exact restore stage");
+ fs::set_permissions(&staged, fs::Permissions::from_mode(0o600))
+ .expect("restrict staged database");
+ let authority = WriterAuthority::acquire(&paths, OpenMode::ReadWriteExisting)
+ .expect("writer authority")
+ .expect("writable mode retains authority");
+ let marker = crate::restore::RestoreRecoveryMarker::prepared(
+ &database_metadata(&paths),
+ crate::BackupManifestSha256::from_bytes([23; 32]),
+ restore_expectation(paths.state_database()),
+ restore_expectation(&staged),
+ )
+ .expect("prepared restore marker");
+ crate::restore::RestoreMarkerBinding::create(&paths, &authority, &marker)
+ .expect("persist prepared marker");
+ drop(authority);
+
+ let state_directory = paths.state_database().parent().expect("state directory");
+ let before = directory_snapshot(state_directory);
+ let read_only = open_existing_connection_pool(
+ &paths,
+ &identity,
+ &base_catalog(),
+ &base_schema_catalog(),
+ OpenMode::ReadOnlyInspection,
+ policy,
+ )
+ .await;
+ let Err(error) = read_only else {
+ panic!("read-only open must not recover");
+ };
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert_eq!(directory_snapshot(state_directory), before);
+
+ let initialize = crate::initialize_database(
+ &paths,
+ OpenMode::Initialize,
+ &database_metadata(&paths),
+ &base_schema_catalog(),
+ |_| async { Ok::<_, Infallible>(()) },
+ )
+ .await
+ .expect_err("initialize must not recover existing evidence");
+ assert_eq!(initialize.kind(), ServiceSqliteErrorKind::Recovery);
+ assert_eq!(directory_snapshot(state_directory), before);
+
+ let initialized_authority = WriterAuthority::acquire(&paths, OpenMode::ReadWriteExisting)
+ .expect("writer authority")
+ .expect("writable mode retains authority");
+ let initialized = open_initialized_connection_pool(
+ &paths,
+ &identity,
+ &base_catalog(),
+ &base_schema_catalog(),
+ policy,
+ initialized_authority,
+ )
+ .await;
+ let Err(error) = initialized else {
+ panic!("initialized open must not recover");
+ };
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert_eq!(directory_snapshot(state_directory), before);
+
+ let writable = open_existing_connection_pool(
+ &paths,
+ &identity,
+ &base_catalog(),
+ &base_schema_catalog(),
+ OpenMode::ReadWriteExisting,
+ policy,
+ )
+ .await
+ .expect("writable open recovers before SQLite");
+ assert!(!staged.exists());
+ assert!(
+ !paths
+ .state_database()
+ .with_file_name(crate::restore::MARKER_FILE_NAME)
+ .exists()
+ );
+ drop(writable.close().await);
+ }
+
+ #[cfg(any(target_os = "linux", target_os = "macos"))]
#[test]
fn connection_failure_precedence_is_exact() {
assert_eq!(
diff --git a/crates/service_sqlite/src/restore/marker.rs b/crates/service_sqlite/src/restore/marker.rs
@@ -14,7 +14,8 @@ use serde::{Deserialize, Serialize};
use sha2::{Digest, Sha256};
use crate::{
- BackupManifestSha256, ServiceDatabaseMetadata, ServiceSqliteApplicationId, ServiceSqlitePaths,
+ BackupManifestSha256, ServiceDatabaseIdentity, ServiceDatabaseMetadata,
+ ServiceSqliteApplicationId, ServiceSqlitePaths,
};
#[cfg(any(target_os = "linux", target_os = "macos"))]
@@ -24,11 +25,11 @@ const RESTORE_MARKER_SCHEMA: &str = "radroots.service-sqlite.restore-marker";
const RESTORE_MARKER_SCHEMA_VERSION: u32 = 1;
const RESTORE_MARKER_MAX_BYTES: usize = 2_048;
const RESTORE_MARKER_CHECKSUM_DOMAIN: &[u8] = b"radroots.service_sqlite.restore_marker.v1\0";
-pub(super) const LIVE_FILE_NAME: &str = radroots_runtime_paths::SERVICE_STATE_DATABASE_FILE_NAME;
-pub(super) const STAGED_FILE_NAME: &str = "state.restore-staged.sqlite";
-pub(super) const BACKUP_FILE_NAME: &str = "state.restore-backup.sqlite";
-pub(super) const MARKER_FILE_NAME: &str = "state.restore-marker.v1";
-pub(super) const MARKER_NEXT_FILE_NAME: &str = "state.restore-marker.v1.next";
+pub(crate) const LIVE_FILE_NAME: &str = radroots_runtime_paths::SERVICE_STATE_DATABASE_FILE_NAME;
+pub(crate) const STAGED_FILE_NAME: &str = "state.restore-staged.sqlite";
+pub(crate) const BACKUP_FILE_NAME: &str = "state.restore-backup.sqlite";
+pub(crate) const MARKER_FILE_NAME: &str = "state.restore-marker.v1";
+pub(crate) const MARKER_NEXT_FILE_NAME: &str = "state.restore-marker.v1.next";
/// Exact retained identity and content expected for one restore artifact.
#[derive(Clone, Copy, PartialEq, Eq)]
@@ -371,6 +372,14 @@ impl RestoreRecoveryMarker {
fn matches_paths(&self, paths: &ServiceSqlitePaths) -> bool {
self.service == *paths.service() && self.instance == *paths.instance()
}
+
+ pub(crate) fn matches_identity(&self, identity: &ServiceDatabaseIdentity) -> bool {
+ self.service == *identity.service()
+ && self.instance == *identity.instance()
+ && self.source_generation == identity.source_generation()
+ && self.application_id == identity.application_id()
+ && self.state_schema_version <= identity.supported_state_schema_version()
+ }
}
impl fmt::Debug for RestoreRecoveryMarker {
@@ -747,13 +756,94 @@ mod store {
Ok(Some(binding))
}
+ pub(crate) fn load_for_recovery(
+ paths: &ServiceSqlitePaths,
+ authority: &WriterAuthority,
+ ) -> Result<Option<Self>, ServiceSqliteError> {
+ authority.validate_for(paths)?;
+ let layout = RestoreRecoveryLayout::for_paths(paths).map_err(recovery_contract)?;
+ let directory = authority_checked(authority, paths, || {
+ authority
+ .directory()
+ .try_clone()
+ .map_err(|_| StoreFailure::Directory)
+ })?
+ .map_err(recovery_store)?;
+ let directory_identity =
+ authority_checked(authority, paths, || validate_directory(&directory))?
+ .map_err(authority_store)?;
+ let marker_file = match authority_checked(authority, paths, || {
+ open_marker_file(&directory, MARKER_FILE_NAME)
+ })? {
+ Ok(file) => file,
+ Err(StoreFailure::Missing) => {
+ authority_checked(authority, paths, || {
+ require_absent(&directory, MARKER_NEXT_FILE_NAME)
+ })?
+ .map_err(recovery_store)?;
+ return Ok(None);
+ }
+ Err(cause) => return Err(recovery_store(cause)),
+ };
+ let marker_identity =
+ authority_checked(authority, paths, || file_identity(&marker_file))?
+ .map_err(recovery_store)?;
+ let marker = authority_checked(authority, paths, || read_marker(&marker_file))?
+ .map_err(recovery_store)?;
+ if !marker.matches_paths(paths) {
+ return Err(recovery_contract(
+ RestoreMarkerContractError::InvalidIdentity,
+ ));
+ }
+ let binding = Self {
+ directory,
+ directory_identity,
+ marker_file,
+ marker_identity,
+ marker,
+ };
+ authority_checked(authority, paths, || {
+ binding.validate_inner(paths, false, ServiceSqliteErrorKind::Recovery)
+ })??;
+ if layout
+ .marker
+ .file_name()
+ .is_none_or(|name| name != MARKER_FILE_NAME)
+ {
+ return Err(recovery_contract(RestoreMarkerContractError::InvalidLayout));
+ }
+ authority.validate_for(paths)?;
+ Ok(Some(binding))
+ }
+
pub(crate) fn advance(
self,
paths: &ServiceSqlitePaths,
authority: &WriterAuthority,
next: RestoreRecoveryPhase,
) -> Result<Self, ServiceSqliteError> {
- self.advance_with_operations(paths, authority, next, &SystemStoreOperations)
+ self.advance_with_operations(
+ paths,
+ authority,
+ next,
+ &SystemStoreOperations,
+ ServiceSqliteErrorKind::Restore,
+ )
+ }
+
+ pub(crate) fn advance_for_recovery(
+ self,
+ paths: &ServiceSqlitePaths,
+ authority: &WriterAuthority,
+ next: RestoreRecoveryPhase,
+ ) -> Result<Self, ServiceSqliteError> {
+ self.advance_with_operations(
+ paths,
+ authority,
+ next,
+ &SystemStoreOperations,
+ ServiceSqliteErrorKind::Recovery,
+ )
}
fn advance_with_operations(
@@ -762,30 +852,33 @@ mod store {
authority: &WriterAuthority,
next: RestoreRecoveryPhase,
operations: &dyn StoreOperations,
+ operation_kind: ServiceSqliteErrorKind,
) -> Result<Self, ServiceSqliteError> {
- authority_checked(authority, paths, || self.validate_for_restore(paths))??;
+ authority_checked(authority, paths, || {
+ self.validate_inner(paths, true, operation_kind)
+ })??;
let current = authority_checked(authority, paths, || read_marker(&self.marker_file))?
- .map_err(restore_store)?;
+ .map_err(|cause| operation_store(operation_kind, cause))?;
if current.canonical_bytes() != self.marker.canonical_bytes() {
- return Err(restore_store(StoreFailure::Conflict));
+ return Err(operation_store(operation_kind, StoreFailure::Conflict));
}
let next_marker = self
.marker
.transitioned_to(next)
- .map_err(restore_contract)?;
+ .map_err(|cause| operation_contract(operation_kind, cause))?;
if next_marker.canonical_bytes() == self.marker.canonical_bytes() {
return Ok(self);
}
authority_checked(authority, paths, || {
require_absent(&self.directory, MARKER_NEXT_FILE_NAME)
})?
- .map_err(restore_store)?;
+ .map_err(|cause| operation_store(operation_kind, cause))?;
let (scratch, scratch_identity) = authority_checked(authority, paths, || {
let file = create_marker_file(&self.directory, MARKER_NEXT_FILE_NAME)?;
let identity = file_identity(&file)?;
Ok::<_, StoreFailure>((file, identity))
})?
- .map_err(restore_store)?;
+ .map_err(|cause| operation_store(operation_kind, cause))?;
let scratch_write = authority_checked(authority, paths, || {
write_and_sync(
&scratch,
@@ -802,10 +895,10 @@ mod store {
MARKER_NEXT_FILE_NAME,
scratch_identity,
)?;
- return Err(restore_store(cause));
+ return Err(operation_store(operation_kind, cause));
}
authority_checked(authority, paths, || {
- self.validate_inner(paths, false, ServiceSqliteErrorKind::Restore)
+ self.validate_inner(paths, false, operation_kind)
})??;
let scratch_matches = authority_checked(authority, paths, || {
let current = open_marker_file(&self.directory, MARKER_NEXT_FILE_NAME)?;
@@ -814,7 +907,7 @@ mod store {
&& file_identity(¤t)? == scratch_identity,
)
})?
- .map_err(restore_store)?;
+ .map_err(|cause| operation_store(operation_kind, cause))?;
if !scratch_matches {
cleanup_with_authority(
authority,
@@ -823,7 +916,7 @@ mod store {
MARKER_NEXT_FILE_NAME,
scratch_identity,
)?;
- return Err(restore_store(StoreFailure::Conflict));
+ return Err(operation_store(operation_kind, StoreFailure::Conflict));
}
let replacement = authority_checked(authority, paths, || {
operations
@@ -838,7 +931,7 @@ mod store {
MARKER_NEXT_FILE_NAME,
scratch_identity,
)?;
- return Err(restore_store(StoreFailure::Rename));
+ return Err(operation_store(operation_kind, StoreFailure::Rename));
}
let parent_sync = authority_checked(authority, paths, || {
operations
@@ -846,7 +939,7 @@ mod store {
.map_err(|_| StoreFailure::Sync)
})?;
if parent_sync.is_err() {
- return Err(restore_store(StoreFailure::Sync));
+ return Err(operation_store(operation_kind, StoreFailure::Sync));
}
let (marker_file, marker_identity, reread) =
authority_checked(authority, paths, || {
@@ -855,9 +948,9 @@ mod store {
let marker = read_marker(&file)?;
Ok::<_, StoreFailure>((file, identity, marker))
})?
- .map_err(restore_store)?;
+ .map_err(|cause| operation_store(operation_kind, cause))?;
if reread.canonical_bytes() != next_marker.canonical_bytes() {
- return Err(restore_store(StoreFailure::Conflict));
+ return Err(operation_store(operation_kind, StoreFailure::Conflict));
}
let binding = Self {
directory: self.directory,
@@ -866,7 +959,9 @@ mod store {
marker_identity,
marker: next_marker,
};
- authority_checked(authority, paths, || binding.validate_for_restore(paths))??;
+ authority_checked(authority, paths, || {
+ binding.validate_inner(paths, true, operation_kind)
+ })??;
Ok(binding)
}
@@ -883,6 +978,24 @@ mod store {
authority,
next,
&FailingStoreOperations { failure },
+ ServiceSqliteErrorKind::Restore,
+ )
+ }
+
+ #[cfg(test)]
+ pub(crate) fn test_advance_for_recovery_with_failure(
+ self,
+ paths: &ServiceSqlitePaths,
+ authority: &WriterAuthority,
+ next: RestoreRecoveryPhase,
+ failure: TestStoreFailure,
+ ) -> Result<Self, ServiceSqliteError> {
+ self.advance_with_operations(
+ paths,
+ authority,
+ next,
+ &FailingStoreOperations { failure },
+ ServiceSqliteErrorKind::Recovery,
)
}
@@ -890,6 +1003,113 @@ mod store {
&self.marker
}
+ pub(crate) fn interrupted_transition(
+ &self,
+ paths: &ServiceSqlitePaths,
+ authority: &WriterAuthority,
+ ) -> Result<Option<RestoreRecoveryPhase>, ServiceSqliteError> {
+ authority_checked(authority, paths, || {
+ self.validate_inner(paths, false, ServiceSqliteErrorKind::Recovery)
+ })??;
+ let scratch = match authority_checked(authority, paths, || {
+ open_marker_file(&self.directory, MARKER_NEXT_FILE_NAME)
+ })? {
+ Ok(file) => file,
+ Err(StoreFailure::Missing) => return Ok(None),
+ Err(cause) => return Err(recovery_store(cause)),
+ };
+ let scratch_marker = authority_checked(authority, paths, || read_marker(&scratch))?
+ .map_err(recovery_store)?;
+ let next = scratch_marker.phase();
+ let expected = self.marker.transitioned_to(next);
+ if next == self.marker.phase()
+ || match expected {
+ Ok(expected) => expected.canonical_bytes() != scratch_marker.canonical_bytes(),
+ Err(_) => true,
+ }
+ {
+ return Err(recovery_store(StoreFailure::Conflict));
+ }
+ Ok(Some(next))
+ }
+
+ pub(crate) fn promote_interrupted_transition(
+ self,
+ paths: &ServiceSqlitePaths,
+ authority: &WriterAuthority,
+ expected_phase: RestoreRecoveryPhase,
+ ) -> Result<Self, ServiceSqliteError> {
+ self.promote_interrupted_transition_with_hook(paths, authority, expected_phase, || {})
+ }
+
+ fn promote_interrupted_transition_with_hook(
+ self,
+ paths: &ServiceSqlitePaths,
+ authority: &WriterAuthority,
+ expected_phase: RestoreRecoveryPhase,
+ before_exact_removal: impl FnOnce(),
+ ) -> Result<Self, ServiceSqliteError> {
+ authority.validate_for(paths)?;
+ authority_checked(authority, paths, || {
+ self.validate_inner(paths, false, ServiceSqliteErrorKind::Recovery)
+ })??;
+ let scratch = authority_checked(authority, paths, || {
+ open_marker_file(&self.directory, MARKER_NEXT_FILE_NAME)
+ })?
+ .map_err(recovery_store)?;
+ let scratch_identity = authority_checked(authority, paths, || file_identity(&scratch))?
+ .map_err(recovery_store)?;
+ let scratch_marker = authority_checked(authority, paths, || read_marker(&scratch))?
+ .map_err(recovery_store)?;
+ let expected = self
+ .marker
+ .transitioned_to(expected_phase)
+ .map_err(recovery_contract)?;
+ if scratch_marker.canonical_bytes() != expected.canonical_bytes() {
+ return Err(recovery_store(StoreFailure::Conflict));
+ }
+ before_exact_removal();
+ // Preserve the valid current marker even if the scratch pathname
+ // was replaced after an interrupted advance. Remove only the
+ // exact validated scratch, then recreate the governed transition.
+ let removal = authority_checked(authority, paths, || {
+ remove_exact_and_sync(&self.directory, MARKER_NEXT_FILE_NAME, scratch_identity)
+ })?;
+ removal.map_err(recovery_store)?;
+ self.advance_for_recovery(paths, authority, expected_phase)
+ }
+
+ #[cfg(test)]
+ pub(crate) fn test_promote_interrupted_transition_after_hook(
+ self,
+ paths: &ServiceSqlitePaths,
+ authority: &WriterAuthority,
+ expected_phase: RestoreRecoveryPhase,
+ before_exact_removal: impl FnOnce(),
+ ) -> Result<Self, ServiceSqliteError> {
+ self.promote_interrupted_transition_with_hook(
+ paths,
+ authority,
+ expected_phase,
+ before_exact_removal,
+ )
+ }
+
+ pub(crate) fn retire(
+ self,
+ paths: &ServiceSqlitePaths,
+ authority: &WriterAuthority,
+ ) -> Result<(), ServiceSqliteError> {
+ authority.validate_for(paths)?;
+ authority_checked(authority, paths, || {
+ self.validate_inner(paths, true, ServiceSqliteErrorKind::Recovery)
+ })??;
+ authority_checked(authority, paths, || {
+ remove_exact_and_sync(&self.directory, MARKER_FILE_NAME, self.marker_identity)
+ })?
+ .map_err(recovery_store)
+ }
+
pub(crate) fn validate(
&self,
paths: &ServiceSqlitePaths,
@@ -1183,6 +1403,20 @@ mod store {
}
}
+ fn remove_exact_and_sync(
+ directory: &File,
+ name: &str,
+ expected: FileIdentity,
+ ) -> Result<(), StoreFailure> {
+ let current = open_marker_file(directory, name)?;
+ if file_identity(¤t)? != expected {
+ return Err(StoreFailure::Conflict);
+ }
+ unlinkat(directory, name, AtFlags::empty()).map_err(|_| StoreFailure::Conflict)?;
+ directory.sync_all().map_err(|_| StoreFailure::Sync)?;
+ require_absent(directory, name)
+ }
+
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
enum StoreFailure {
Directory,
@@ -1253,6 +1487,8 @@ mod store {
#[cfg(any(target_os = "linux", target_os = "macos"))]
pub(crate) use store::RestoreMarkerBinding;
+#[cfg(all(test, any(target_os = "linux", target_os = "macos")))]
+pub(crate) use store::TestStoreFailure;
#[cfg(test)]
mod tests {
diff --git a/crates/service_sqlite/src/restore/mod.rs b/crates/service_sqlite/src/restore/mod.rs
@@ -2,37 +2,23 @@
mod finalize;
mod marker;
+mod recover;
mod stage;
pub use finalize::finalize_staged_restore;
pub use stage::{StagedServiceRestore, stage_verified_restore};
#[cfg(any(target_os = "linux", target_os = "macos"))]
-pub(crate) fn refuse_unresolved_recovery(
- directory: &impl std::os::fd::AsFd,
-) -> Result<(), crate::ServiceSqliteError> {
- use rustix::{
- fs::{AtFlags, statat},
- io::Errno,
- };
+pub(crate) use recover::refuse_unresolved_recovery;
- for name in [marker::MARKER_FILE_NAME, marker::MARKER_NEXT_FILE_NAME] {
- match statat(directory, name, AtFlags::SYMLINK_NOFOLLOW) {
- Err(Errno::NOENT) => {}
- Ok(_) | Err(_) => {
- return Err(crate::ServiceSqliteError::new(
- crate::ServiceSqliteErrorKind::Recovery,
- ));
- }
- }
- }
- Ok(())
-}
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+pub(crate) use recover::recover_for_open;
#[allow(unused_imports)]
pub(crate) use marker::{
+ BACKUP_FILE_NAME, LIVE_FILE_NAME, MARKER_FILE_NAME, MARKER_NEXT_FILE_NAME,
RestoreArtifactExpectation, RestoreMarkerContractError, RestoreRecoveryLayout,
- RestoreRecoveryMarker, RestoreRecoveryPhase,
+ RestoreRecoveryMarker, RestoreRecoveryPhase, STAGED_FILE_NAME,
};
#[cfg(any(target_os = "linux", target_os = "macos"))]
diff --git a/crates/service_sqlite/src/restore/recover.rs b/crates/service_sqlite/src/restore/recover.rs
@@ -0,0 +1,1052 @@
+//! Deterministic writer-authorized reconciliation of interrupted restores.
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+use core::fmt;
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+use {
+ super::{
+ RestoreArtifactExpectation, RestoreMarkerBinding, RestoreRecoveryPhase,
+ marker::{
+ BACKUP_FILE_NAME, LIVE_FILE_NAME, MARKER_FILE_NAME, MARKER_NEXT_FILE_NAME,
+ STAGED_FILE_NAME,
+ },
+ },
+ crate::{
+ ServiceDatabaseIdentity, ServiceSqliteError, ServiceSqliteErrorKind, ServiceSqlitePaths,
+ WriterAuthority,
+ },
+ rustix::{
+ fs::{
+ AtFlags, FileType, Mode, OFlags, RenameFlags, fstat, openat, renameat_with, statat,
+ unlinkat,
+ },
+ io::Errno,
+ process::geteuid,
+ },
+ sha2::{Digest, Sha256},
+ std::{error::Error, fs::File, os::unix::fs::FileExt},
+};
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+const HASH_BUFFER_BYTES: usize = 64 * 1_024;
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+pub(crate) fn recover_for_open(
+ paths: &ServiceSqlitePaths,
+ identity: &ServiceDatabaseIdentity,
+ authority: &WriterAuthority,
+) -> Result<(), ServiceSqliteError> {
+ authority.validate_for(paths)?;
+ let Some(mut marker) = RestoreMarkerBinding::load_for_recovery(paths, authority)? else {
+ authority_checked(authority, paths, || {
+ refuse_unresolved_recovery(authority.directory())
+ })?;
+ return Ok(());
+ };
+ if !marker.marker().matches_identity(identity) {
+ return Err(recovery_error(RecoveryFailureKind::Intent));
+ }
+
+ let observed = authority_checked(authority, paths, || {
+ observe_artifacts(authority.directory(), marker.marker())
+ })?;
+ if let Some(next) = marker.interrupted_transition(paths, authority)? {
+ let topology_matches = match (marker.marker().phase(), next) {
+ (RestoreRecoveryPhase::Prepared, RestoreRecoveryPhase::LiveRetained) => {
+ observed.proves_live_retained()
+ }
+ (RestoreRecoveryPhase::LiveRetained, RestoreRecoveryPhase::ReplacementInstalled) => {
+ observed.proves_replacement_installed()
+ }
+ _ => false,
+ };
+ if !topology_matches {
+ return Err(recovery_error(RecoveryFailureKind::Topology));
+ }
+ marker = marker.promote_interrupted_transition(paths, authority, next)?;
+ }
+
+ loop {
+ let observed = authority_checked(authority, paths, || {
+ observe_artifacts(authority.directory(), marker.marker())
+ })?;
+ match marker.marker().phase() {
+ RestoreRecoveryPhase::Prepared if observed.can_roll_back_prepared() => {
+ if let Some(staged) = observed.staged.as_ref() {
+ authority_checked(authority, paths, || {
+ remove_exact_artifact(
+ authority.directory(),
+ STAGED_FILE_NAME,
+ staged,
+ marker.marker().staged(),
+ )
+ })?;
+ }
+ marker.retire(paths, authority)?;
+ authority_checked(authority, paths, || {
+ refuse_unresolved_recovery(authority.directory())
+ })?;
+ return Ok(());
+ }
+ RestoreRecoveryPhase::Prepared if observed.proves_live_retained() => {
+ authority_checked(authority, paths, || {
+ authority.directory().sync_all().map_err(|source| {
+ recovery_source(RecoveryFailureKind::DirectorySync, source)
+ })
+ })?;
+ marker = marker.advance_for_recovery(
+ paths,
+ authority,
+ RestoreRecoveryPhase::LiveRetained,
+ )?;
+ }
+ RestoreRecoveryPhase::LiveRetained if observed.needs_replacement_install() => {
+ let staged = observed
+ .staged
+ .as_ref()
+ .ok_or_else(|| recovery_error(RecoveryFailureKind::Topology))?;
+ authority_checked(authority, paths, || {
+ verify_named_artifact(
+ authority.directory(),
+ STAGED_FILE_NAME,
+ staged,
+ marker.marker().staged(),
+ )
+ })?;
+ authority_checked(authority, paths, || {
+ renameat_with(
+ authority.directory(),
+ STAGED_FILE_NAME,
+ authority.directory(),
+ LIVE_FILE_NAME,
+ RenameFlags::NOREPLACE,
+ )
+ .map_err(|source| {
+ recovery_source(RecoveryFailureKind::InstallReplacement, source)
+ })
+ })?;
+ authority_checked(authority, paths, || {
+ verify_named_artifact(
+ authority.directory(),
+ LIVE_FILE_NAME,
+ staged,
+ marker.marker().staged(),
+ )
+ })?;
+ authority_checked(authority, paths, || {
+ authority.directory().sync_all().map_err(|source| {
+ recovery_source(RecoveryFailureKind::DirectorySync, source)
+ })
+ })?;
+ marker = marker.advance_for_recovery(
+ paths,
+ authority,
+ RestoreRecoveryPhase::ReplacementInstalled,
+ )?;
+ }
+ RestoreRecoveryPhase::LiveRetained if observed.proves_replacement_installed() => {
+ authority_checked(authority, paths, || {
+ authority.directory().sync_all().map_err(|source| {
+ recovery_source(RecoveryFailureKind::DirectorySync, source)
+ })
+ })?;
+ marker = marker.advance_for_recovery(
+ paths,
+ authority,
+ RestoreRecoveryPhase::ReplacementInstalled,
+ )?;
+ }
+ RestoreRecoveryPhase::ReplacementInstalled
+ if observed.proves_replacement_installed_or_cleanup() =>
+ {
+ let live = match observed.live {
+ LiveArtifact::Replacement(ref file) => file,
+ _ => return Err(recovery_error(RecoveryFailureKind::Topology)),
+ };
+ authority_checked(authority, paths, || {
+ verify_named_artifact(
+ authority.directory(),
+ LIVE_FILE_NAME,
+ live,
+ observed.marker_staged,
+ )
+ })?;
+ if let Some(backup) = observed.backup.as_ref() {
+ authority_checked(authority, paths, || {
+ remove_exact_artifact(
+ authority.directory(),
+ BACKUP_FILE_NAME,
+ backup,
+ observed.marker_live,
+ )
+ })?;
+ }
+ marker.retire(paths, authority)?;
+ authority_checked(authority, paths, || {
+ refuse_unresolved_recovery(authority.directory())
+ })?;
+ authority_checked(authority, paths, || {
+ verify_named_artifact(
+ authority.directory(),
+ LIVE_FILE_NAME,
+ live,
+ observed.marker_staged,
+ )
+ })?;
+ return Ok(());
+ }
+ _ => return Err(recovery_error(RecoveryFailureKind::Topology)),
+ }
+ }
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+pub(crate) fn refuse_unresolved_recovery(
+ directory: &impl std::os::fd::AsFd,
+) -> Result<(), ServiceSqliteError> {
+ for name in [
+ STAGED_FILE_NAME,
+ BACKUP_FILE_NAME,
+ MARKER_FILE_NAME,
+ MARKER_NEXT_FILE_NAME,
+ ] {
+ match statat(directory, name, AtFlags::SYMLINK_NOFOLLOW) {
+ Err(Errno::NOENT) => {}
+ Ok(_) | Err(_) => {
+ return Err(ServiceSqliteError::new(ServiceSqliteErrorKind::Recovery));
+ }
+ }
+ }
+ Ok(())
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+struct ObservedArtifacts {
+ live: LiveArtifact,
+ staged: Option<File>,
+ backup: Option<File>,
+ marker_live: RestoreArtifactExpectation,
+ marker_staged: RestoreArtifactExpectation,
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+impl ObservedArtifacts {
+ fn can_roll_back_prepared(&self) -> bool {
+ matches!(self.live, LiveArtifact::Original) && self.backup.is_none()
+ }
+
+ fn proves_live_retained(&self) -> bool {
+ matches!(self.live, LiveArtifact::Absent) && self.staged.is_some() && self.backup.is_some()
+ }
+
+ fn needs_replacement_install(&self) -> bool {
+ self.proves_live_retained()
+ }
+
+ fn proves_replacement_installed(&self) -> bool {
+ matches!(self.live, LiveArtifact::Replacement(_))
+ && self.staged.is_none()
+ && self.backup.is_some()
+ }
+
+ fn proves_replacement_installed_or_cleanup(&self) -> bool {
+ matches!(self.live, LiveArtifact::Replacement(_)) && self.staged.is_none()
+ }
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+enum LiveArtifact {
+ Absent,
+ Original,
+ Replacement(File),
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn observe_artifacts(
+ directory: &File,
+ marker: &super::RestoreRecoveryMarker,
+) -> Result<ObservedArtifacts, ServiceSqliteError> {
+ for sidecar in [
+ "state.sqlite-wal",
+ "state.sqlite-shm",
+ "state.sqlite-journal",
+ ] {
+ require_absent(directory, sidecar)?;
+ }
+ let live = match open_optional(directory, LIVE_FILE_NAME)? {
+ None => LiveArtifact::Absent,
+ Some(file) if artifact_has_identity(&file, marker.live())? => {
+ verify_artifact(&file, marker.live())?;
+ LiveArtifact::Original
+ }
+ Some(file) if artifact_has_identity(&file, marker.staged())? => {
+ verify_artifact(&file, marker.staged())?;
+ LiveArtifact::Replacement(file)
+ }
+ Some(_) => return Err(recovery_error(RecoveryFailureKind::Artifact)),
+ };
+ let staged = observe_expected(directory, STAGED_FILE_NAME, marker.staged())?;
+ let backup = observe_expected(directory, BACKUP_FILE_NAME, marker.backup())?;
+ Ok(ObservedArtifacts {
+ live,
+ staged,
+ backup,
+ marker_live: marker.live(),
+ marker_staged: marker.staged(),
+ })
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn observe_expected(
+ directory: &File,
+ name: &str,
+ expected: RestoreArtifactExpectation,
+) -> Result<Option<File>, ServiceSqliteError> {
+ let Some(file) = open_optional(directory, name)? else {
+ return Ok(None);
+ };
+ verify_artifact(&file, expected)?;
+ Ok(Some(file))
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn open_optional(directory: &File, name: &str) -> Result<Option<File>, ServiceSqliteError> {
+ match openat(
+ directory,
+ name,
+ OFlags::RDONLY | OFlags::NOFOLLOW | OFlags::CLOEXEC | OFlags::NONBLOCK,
+ Mode::empty(),
+ ) {
+ Ok(file) => Ok(Some(File::from(file))),
+ Err(Errno::NOENT) => Ok(None),
+ Err(source) => Err(recovery_source(RecoveryFailureKind::Artifact, source)),
+ }
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn artifact_has_identity(
+ file: &File,
+ expected: RestoreArtifactExpectation,
+) -> Result<bool, ServiceSqliteError> {
+ let status =
+ fstat(file).map_err(|source| recovery_source(RecoveryFailureKind::Artifact, source))?;
+ let device =
+ u64::try_from(status.st_dev).map_err(|_| recovery_error(RecoveryFailureKind::Artifact))?;
+ Ok((device, status.st_ino) == (expected.device(), expected.inode()))
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn verify_artifact(
+ file: &File,
+ expected: RestoreArtifactExpectation,
+) -> Result<(), ServiceSqliteError> {
+ let status =
+ fstat(file).map_err(|source| recovery_source(RecoveryFailureKind::Artifact, source))?;
+ let device =
+ u64::try_from(status.st_dev).map_err(|_| recovery_error(RecoveryFailureKind::Artifact))?;
+ let length =
+ u64::try_from(status.st_size).map_err(|_| recovery_error(RecoveryFailureKind::Artifact))?;
+ if !FileType::from_raw_mode(status.st_mode).is_file()
+ || u64::from(status.st_nlink) != 1
+ || status.st_uid != geteuid().as_raw()
+ || u32::from(status.st_mode) & 0o777 != 0o600
+ || (device, status.st_ino) != (expected.device(), expected.inode())
+ || length != expected.byte_length()
+ || length == 0
+ || length > i64::MAX as u64
+ || hash_exact(file, expected.byte_length())? != expected.sha256()
+ {
+ return Err(recovery_error(RecoveryFailureKind::Artifact));
+ }
+ Ok(())
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn verify_named_artifact(
+ directory: &File,
+ name: &str,
+ held: &File,
+ expected: RestoreArtifactExpectation,
+) -> Result<(), ServiceSqliteError> {
+ verify_artifact(held, expected)?;
+ let current = open_optional(directory, name)?
+ .ok_or_else(|| recovery_error(RecoveryFailureKind::Artifact))?;
+ verify_artifact(¤t, expected)
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn remove_exact_artifact(
+ directory: &File,
+ name: &str,
+ held: &File,
+ expected: RestoreArtifactExpectation,
+) -> Result<(), ServiceSqliteError> {
+ verify_named_artifact(directory, name, held, expected)?;
+ unlinkat(directory, name, AtFlags::empty())
+ .map_err(|source| recovery_source(RecoveryFailureKind::Cleanup, source))?;
+ directory
+ .sync_all()
+ .map_err(|source| recovery_source(RecoveryFailureKind::DirectorySync, source))?;
+ require_absent(directory, name)
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn require_absent(directory: &File, name: &str) -> Result<(), ServiceSqliteError> {
+ match statat(directory, name, AtFlags::SYMLINK_NOFOLLOW) {
+ Err(Errno::NOENT) => Ok(()),
+ Ok(_) | Err(_) => Err(recovery_error(RecoveryFailureKind::Topology)),
+ }
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn hash_exact(file: &File, expected_length: u64) -> Result<[u8; 32], ServiceSqliteError> {
+ let mut hasher = Sha256::new();
+ let mut buffer = [0_u8; HASH_BUFFER_BYTES];
+ let mut offset = 0_u64;
+ while offset < expected_length {
+ let requested = usize::try_from((expected_length - offset).min(HASH_BUFFER_BYTES as u64))
+ .map_err(|_| recovery_error(RecoveryFailureKind::Hash))?;
+ let read = file
+ .read_at(&mut buffer[..requested], offset)
+ .map_err(|source| recovery_source(RecoveryFailureKind::Hash, source))?;
+ if read == 0 {
+ return Err(recovery_error(RecoveryFailureKind::Hash));
+ }
+ hasher.update(&buffer[..read]);
+ offset = offset
+ .checked_add(
+ u64::try_from(read).map_err(|_| recovery_error(RecoveryFailureKind::Hash))?,
+ )
+ .ok_or_else(|| recovery_error(RecoveryFailureKind::Hash))?;
+ }
+ let mut extra = [0_u8; 1];
+ if file
+ .read_at(&mut extra, expected_length)
+ .map_err(|source| recovery_source(RecoveryFailureKind::Hash, source))?
+ != 0
+ {
+ return Err(recovery_error(RecoveryFailureKind::Hash));
+ }
+ Ok(hasher.finalize().into())
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn authority_checked<T>(
+ authority: &WriterAuthority,
+ paths: &ServiceSqlitePaths,
+ operation: impl FnOnce() -> Result<T, ServiceSqliteError>,
+) -> Result<T, ServiceSqliteError> {
+ authority.validate_for(paths)?;
+ let result = operation();
+ authority.validate_for(paths)?;
+ result
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+#[derive(Clone, Copy, Debug, PartialEq, Eq)]
+enum RecoveryFailureKind {
+ Intent,
+ Topology,
+ Artifact,
+ Hash,
+ InstallReplacement,
+ DirectorySync,
+ Cleanup,
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+struct RecoveryFailure {
+ kind: RecoveryFailureKind,
+ source: Option<Box<dyn Error + Send + Sync + 'static>>,
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+impl fmt::Debug for RecoveryFailure {
+ fn fmt(&self, formatter: &mut fmt::Formatter<'_>) -> fmt::Result {
+ formatter
+ .debug_struct("RecoveryFailure")
+ .field("kind", &self.kind)
+ .field("source", &self.source.as_ref().map(|_| "[redacted]"))
+ .finish()
+ }
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+impl fmt::Display for RecoveryFailure {
+ fn fmt(&self, formatter: &mut fmt::Formatter<'_>) -> fmt::Result {
+ formatter.write_str(match self.kind {
+ RecoveryFailureKind::Intent => "restore recovery intent does not match",
+ RecoveryFailureKind::Topology => "restore recovery topology is ambiguous",
+ RecoveryFailureKind::Artifact => "restore recovery artifact is invalid",
+ RecoveryFailureKind::Hash => "restore recovery artifact hash failed",
+ RecoveryFailureKind::InstallReplacement => {
+ "restore recovery replacement installation failed"
+ }
+ RecoveryFailureKind::DirectorySync => "restore recovery directory durability failed",
+ RecoveryFailureKind::Cleanup => "restore recovery cleanup failed",
+ })
+ }
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+impl Error for RecoveryFailure {
+ fn source(&self) -> Option<&(dyn Error + 'static)> {
+ self.source
+ .as_deref()
+ .map(|source| source as &(dyn Error + 'static))
+ }
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn recovery_error(kind: RecoveryFailureKind) -> ServiceSqliteError {
+ ServiceSqliteError::with_source(
+ ServiceSqliteErrorKind::Recovery,
+ RecoveryFailure { kind, source: None },
+ )
+}
+
+#[cfg(any(target_os = "linux", target_os = "macos"))]
+fn recovery_source(
+ kind: RecoveryFailureKind,
+ source: impl Error + Send + Sync + 'static,
+) -> ServiceSqliteError {
+ ServiceSqliteError::with_source(
+ ServiceSqliteErrorKind::Recovery,
+ RecoveryFailure {
+ kind,
+ source: Some(Box::new(source)),
+ },
+ )
+}
+
+#[cfg(all(test, any(target_os = "linux", target_os = "macos")))]
+mod tests {
+ use std::{
+ fs::{self, File, OpenOptions},
+ io::Write,
+ num::NonZeroU32,
+ os::unix::fs::{FileTypeExt, OpenOptionsExt, PermissionsExt},
+ path::{Path, PathBuf},
+ process::Command,
+ };
+
+ use radroots_runtime_paths::{
+ InstanceId, RadrootsHostEnvironment, RadrootsPathProfile, RadrootsPathResolver,
+ RadrootsPlatform, RuntimeContext, RuntimeContextBootstrap, RuntimeContextSource, ServiceId,
+ };
+ use radroots_storage::event::SourceGeneration;
+
+ use super::*;
+ use crate::restore::RestoreRecoveryMarker;
+ use crate::{
+ BackupManifestSha256, OpenMode, ServiceDatabaseMetadata, ServiceSqliteApplicationId,
+ };
+
+ const OLD_BYTES: &[u8] = b"old-live-state";
+ const NEW_BYTES: &[u8] = b"new-restored-state";
+
+ struct Fixture {
+ _root: tempfile::TempDir,
+ paths: ServiceSqlitePaths,
+ identity: ServiceDatabaseIdentity,
+ authority: WriterAuthority,
+ }
+
+ impl Fixture {
+ fn new() -> Self {
+ let root = tempfile::tempdir().expect("temporary root");
+ let paths = service_paths(root.path());
+ let state_directory = paths.state_database().parent().expect("state directory");
+ fs::create_dir_all(state_directory).expect("create state directory");
+ fs::set_permissions(state_directory, fs::Permissions::from_mode(0o700))
+ .expect("restrict state directory");
+ write_new(paths.state_database(), OLD_BYTES);
+ write_new(&artifact_path(&paths, STAGED_FILE_NAME), NEW_BYTES);
+ let authority = WriterAuthority::acquire(&paths, OpenMode::ReadWriteExisting)
+ .expect("writer authority")
+ .expect("writable mode retains authority");
+ let metadata = database_metadata(&paths);
+ let marker = RestoreRecoveryMarker::prepared(
+ &metadata,
+ BackupManifestSha256::from_bytes([19; 32]),
+ expectation(paths.state_database()),
+ expectation(&artifact_path(&paths, STAGED_FILE_NAME)),
+ )
+ .expect("prepared marker");
+ RestoreMarkerBinding::create(&paths, &authority, &marker)
+ .expect("persist prepared marker");
+ Self {
+ _root: root,
+ paths,
+ identity: metadata.identity(),
+ authority,
+ }
+ }
+
+ fn load_marker(&self) -> RestoreMarkerBinding {
+ RestoreMarkerBinding::load_for_recovery(&self.paths, &self.authority)
+ .expect("load marker")
+ .expect("marker exists")
+ }
+
+ fn retain_live(&self, advance: bool) {
+ fs::rename(
+ self.paths.state_database(),
+ artifact_path(&self.paths, BACKUP_FILE_NAME),
+ )
+ .expect("retain live");
+ sync_state_directory(&self.paths);
+ if advance {
+ self.load_marker()
+ .advance(
+ &self.paths,
+ &self.authority,
+ RestoreRecoveryPhase::LiveRetained,
+ )
+ .expect("advance live-retained marker");
+ }
+ }
+
+ fn install_stage(&self, advance: bool) {
+ fs::rename(
+ artifact_path(&self.paths, STAGED_FILE_NAME),
+ self.paths.state_database(),
+ )
+ .expect("install stage");
+ sync_state_directory(&self.paths);
+ if advance {
+ self.load_marker()
+ .advance(
+ &self.paths,
+ &self.authority,
+ RestoreRecoveryPhase::ReplacementInstalled,
+ )
+ .expect("advance replacement marker");
+ }
+ }
+
+ fn write_scratch(&self, next: RestoreRecoveryPhase) {
+ let marker = self.load_marker();
+ let next = marker
+ .marker()
+ .transitioned_to(next)
+ .expect("legal next phase");
+ write_new(
+ &artifact_path(&self.paths, MARKER_NEXT_FILE_NAME),
+ next.canonical_bytes(),
+ );
+ sync_state_directory(&self.paths);
+ }
+
+ fn recover(&self) -> Result<(), ServiceSqliteError> {
+ recover_for_open(&self.paths, &self.identity, &self.authority)
+ }
+ }
+
+ #[test]
+ fn prepared_with_old_live_rolls_back_and_retries_idempotently() {
+ let fixture = Fixture::new();
+ fixture.recover().expect("rollback prepared restore");
+ assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), OLD_BYTES);
+ assert_no_recovery_evidence(&fixture.paths);
+ fixture.recover().expect("repeated recovery is a no-op");
+ }
+
+ #[test]
+ fn interrupted_prepared_rollback_without_stage_finishes_marker_cleanup() {
+ let fixture = Fixture::new();
+ fs::remove_file(artifact_path(&fixture.paths, STAGED_FILE_NAME)).expect("remove stage");
+ sync_state_directory(&fixture.paths);
+ fixture.recover().expect("finish interrupted rollback");
+ assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), OLD_BYTES);
+ assert_no_recovery_evidence(&fixture.paths);
+ }
+
+ #[test]
+ fn prepared_with_proven_first_rename_rolls_forward() {
+ let fixture = Fixture::new();
+ fixture.retain_live(false);
+ fixture.recover().expect("recover after first rename");
+ assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), NEW_BYTES);
+ assert_no_recovery_evidence(&fixture.paths);
+ }
+
+ #[test]
+ fn live_retained_with_proven_second_rename_rolls_forward() {
+ let fixture = Fixture::new();
+ fixture.retain_live(true);
+ fixture.install_stage(false);
+ fixture.recover().expect("recover after second rename");
+ assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), NEW_BYTES);
+ assert_no_recovery_evidence(&fixture.paths);
+ }
+
+ #[test]
+ fn replacement_installed_retires_backup_and_marker() {
+ let fixture = Fixture::new();
+ fixture.retain_live(true);
+ fixture.install_stage(true);
+ fixture.recover().expect("retire completed recovery");
+ assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), NEW_BYTES);
+ assert_no_recovery_evidence(&fixture.paths);
+ }
+
+ #[test]
+ fn replacement_cleanup_without_backup_finishes_idempotently() {
+ let fixture = Fixture::new();
+ fixture.retain_live(true);
+ fixture.install_stage(true);
+ fs::remove_file(artifact_path(&fixture.paths, BACKUP_FILE_NAME)).expect("remove backup");
+ sync_state_directory(&fixture.paths);
+ fixture.recover().expect("finish interrupted cleanup");
+ assert_eq!(fs::read(fixture.paths.state_database()).unwrap(), NEW_BYTES);
+ assert_no_recovery_evidence(&fixture.paths);
+ fixture.recover().expect("repeated cleanup is a no-op");
+ }
+
+ #[test]
+ fn topology_consistent_marker_scratch_is_promoted_at_both_edges() {
+ let first = Fixture::new();
+ first.retain_live(false);
+ first.write_scratch(RestoreRecoveryPhase::LiveRetained);
+ first.recover().expect("promote first scratch");
+ assert_eq!(fs::read(first.paths.state_database()).unwrap(), NEW_BYTES);
+ assert_no_recovery_evidence(&first.paths);
+
+ let second = Fixture::new();
+ second.retain_live(true);
+ second.install_stage(false);
+ second.write_scratch(RestoreRecoveryPhase::ReplacementInstalled);
+ second.recover().expect("promote second scratch");
+ assert_eq!(fs::read(second.paths.state_database()).unwrap(), NEW_BYTES);
+ assert_no_recovery_evidence(&second.paths);
+ }
+
+ #[test]
+ fn inferred_advance_failures_are_recovery_and_authority_keeps_precedence() {
+ let recovery = Fixture::new();
+ let error = recovery
+ .load_marker()
+ .test_advance_for_recovery_with_failure(
+ &recovery.paths,
+ &recovery.authority,
+ RestoreRecoveryPhase::LiveRetained,
+ crate::restore::marker::TestStoreFailure::ScratchSync,
+ )
+ .expect_err("recovery marker sync failure must be classified");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert!(artifact_path(&recovery.paths, MARKER_FILE_NAME).exists());
+ assert!(!artifact_path(&recovery.paths, MARKER_NEXT_FILE_NAME).exists());
+
+ let authority = Fixture::new();
+ let state_directory = authority.paths.state_database().parent().unwrap();
+ let error = authority
+ .load_marker()
+ .test_advance_for_recovery_with_failure(
+ &authority.paths,
+ &authority.authority,
+ RestoreRecoveryPhase::LiveRetained,
+ crate::restore::marker::TestStoreFailure::AuthorityDriftAndScratchSync,
+ )
+ .expect_err("authority drift must dominate the marker failure");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Authority);
+ assert!(artifact_path(&authority.paths, MARKER_FILE_NAME).exists());
+ fs::set_permissions(state_directory, fs::Permissions::from_mode(0o700))
+ .expect("restore state directory mode");
+ }
+
+ #[test]
+ fn replaced_interrupted_scratch_never_overwrites_the_valid_marker() {
+ let fixture = Fixture::new();
+ fixture.retain_live(false);
+ fixture.write_scratch(RestoreRecoveryPhase::LiveRetained);
+ let marker = fixture.load_marker();
+ let scratch = artifact_path(&fixture.paths, MARKER_NEXT_FILE_NAME);
+ let retained = scratch.with_file_name("retained-valid-marker-next");
+ let replacement_bytes = b"foreign-marker-next";
+ let scratch_for_hook = scratch.clone();
+ let error = marker
+ .test_promote_interrupted_transition_after_hook(
+ &fixture.paths,
+ &fixture.authority,
+ RestoreRecoveryPhase::LiveRetained,
+ move || {
+ fs::rename(&scratch_for_hook, &retained).expect("retain exact scratch");
+ write_new(&scratch_for_hook, replacement_bytes);
+ },
+ )
+ .expect_err("replaced scratch must fail without marker replacement");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert_eq!(fs::read(&scratch).unwrap(), replacement_bytes);
+ assert!(
+ scratch
+ .with_file_name("retained-valid-marker-next")
+ .exists()
+ );
+ let current = fixture.load_marker();
+ assert_eq!(current.marker().phase(), RestoreRecoveryPhase::Prepared);
+ }
+
+ #[test]
+ fn topology_inconsistent_scratch_and_tampered_artifacts_fail_without_cleanup() {
+ let scratch = Fixture::new();
+ scratch.write_scratch(RestoreRecoveryPhase::LiveRetained);
+ let error = scratch
+ .recover()
+ .expect_err("scratch before the first rename is ambiguous");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert!(artifact_path(&scratch.paths, MARKER_FILE_NAME).exists());
+ assert!(artifact_path(&scratch.paths, MARKER_NEXT_FILE_NAME).exists());
+ assert!(artifact_path(&scratch.paths, STAGED_FILE_NAME).exists());
+
+ let tampered = Fixture::new();
+ fs::write(
+ artifact_path(&tampered.paths, STAGED_FILE_NAME),
+ b"tampered-restored",
+ )
+ .expect("tamper stage");
+ let before = fs::read_dir(
+ tampered
+ .paths
+ .state_database()
+ .parent()
+ .expect("state directory"),
+ )
+ .unwrap()
+ .count();
+ let error = tampered
+ .recover()
+ .expect_err("tampered stage must fail closed");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert_eq!(
+ fs::read_dir(
+ tampered
+ .paths
+ .state_database()
+ .parent()
+ .expect("state directory")
+ )
+ .unwrap()
+ .count(),
+ before
+ );
+ assert!(artifact_path(&tampered.paths, MARKER_FILE_NAME).exists());
+ }
+
+ #[test]
+ fn sidecar_mode_and_identity_mismatch_preserve_recovery_evidence() {
+ let sidecar = Fixture::new();
+ write_new(&artifact_path(&sidecar.paths, "state.sqlite-wal"), b"wal");
+ let error = sidecar
+ .recover()
+ .expect_err("restore recovery rejects sidecars");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert!(artifact_path(&sidecar.paths, MARKER_FILE_NAME).exists());
+
+ let mode = Fixture::new();
+ fs::set_permissions(
+ artifact_path(&mode.paths, STAGED_FILE_NAME),
+ fs::Permissions::from_mode(0o640),
+ )
+ .expect("change stage mode");
+ let error = mode
+ .recover()
+ .expect_err("insecure artifact mode fails closed");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert!(artifact_path(&mode.paths, MARKER_FILE_NAME).exists());
+
+ let mismatch = Fixture::new();
+ let wrong = ServiceDatabaseIdentity::new(
+ &mismatch.paths,
+ SourceGeneration::new([99; 32]).expect("different generation"),
+ NonZeroU32::new(1).expect("schema version"),
+ mismatch.identity.application_id(),
+ );
+ let error = recover_for_open(&mismatch.paths, &wrong, &mismatch.authority)
+ .expect_err("marker intent mismatch fails closed");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert!(artifact_path(&mismatch.paths, MARKER_FILE_NAME).exists());
+ }
+
+ #[test]
+ fn symlink_hardlink_fifo_and_foreign_replacement_fail_without_deletion() {
+ use std::os::unix::fs::symlink;
+
+ let symlinked = Fixture::new();
+ let staged = artifact_path(&symlinked.paths, STAGED_FILE_NAME);
+ let held = staged.with_file_name("held-stage");
+ fs::rename(&staged, &held).expect("retain original stage");
+ symlink(symlinked.paths.state_database(), &staged).expect("replace with symlink");
+ let error = symlinked
+ .recover()
+ .expect_err("symlink replacement fails closed");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert!(
+ fs::symlink_metadata(&staged)
+ .unwrap()
+ .file_type()
+ .is_symlink()
+ );
+ assert!(held.exists());
+
+ let hardlinked = Fixture::new();
+ let staged = artifact_path(&hardlinked.paths, STAGED_FILE_NAME);
+ let link = staged.with_file_name("stage-hardlink");
+ fs::hard_link(&staged, &link).expect("create hard link");
+ let error = hardlinked
+ .recover()
+ .expect_err("multiple-link stage fails closed");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert!(staged.exists());
+ assert!(link.exists());
+
+ let fifo = Fixture::new();
+ let staged = artifact_path(&fifo.paths, STAGED_FILE_NAME);
+ fs::remove_file(&staged).expect("remove original stage");
+ assert!(
+ Command::new("mkfifo")
+ .arg(&staged)
+ .status()
+ .expect("run mkfifo")
+ .success()
+ );
+ fs::set_permissions(&staged, fs::Permissions::from_mode(0o600)).expect("restrict FIFO");
+ let error = fifo
+ .recover()
+ .expect_err("nonblocking FIFO admission fails closed");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert!(fs::symlink_metadata(&staged).unwrap().file_type().is_fifo());
+
+ let replaced = Fixture::new();
+ let staged = artifact_path(&replaced.paths, STAGED_FILE_NAME);
+ let original = staged.with_file_name("original-stage");
+ fs::rename(&staged, &original).expect("retain original stage");
+ write_new(&staged, b"foreign-stage");
+ let foreign = fs::read(&staged).unwrap();
+ let error = replaced
+ .recover()
+ .expect_err("foreign replacement fails closed");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert_eq!(fs::read(&staged).unwrap(), foreign);
+ assert_eq!(fs::read(&original).unwrap(), NEW_BYTES);
+ assert!(artifact_path(&replaced.paths, MARKER_FILE_NAME).exists());
+ }
+
+ #[test]
+ fn authority_drift_precedes_recovery_and_keeps_evidence() {
+ let fixture = Fixture::new();
+ let state_directory = fixture.paths.state_database().parent().unwrap();
+ fs::set_permissions(state_directory, fs::Permissions::from_mode(0o770))
+ .expect("drift state directory mode");
+ let error = fixture
+ .recover()
+ .expect_err("authority drift must stop recovery");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Authority);
+ assert!(artifact_path(&fixture.paths, MARKER_FILE_NAME).exists());
+ assert!(artifact_path(&fixture.paths, STAGED_FILE_NAME).exists());
+ fs::set_permissions(state_directory, fs::Permissions::from_mode(0o700))
+ .expect("restore state directory mode");
+ }
+
+ #[test]
+ fn orphan_stage_backup_or_scratch_without_marker_is_never_repaired() {
+ for name in [STAGED_FILE_NAME, BACKUP_FILE_NAME, MARKER_NEXT_FILE_NAME] {
+ let fixture = Fixture::new();
+ fs::remove_file(artifact_path(&fixture.paths, MARKER_FILE_NAME))
+ .expect("remove marker");
+ if name != STAGED_FILE_NAME {
+ fs::remove_file(artifact_path(&fixture.paths, STAGED_FILE_NAME))
+ .expect("remove default stage");
+ write_new(&artifact_path(&fixture.paths, name), b"orphan");
+ }
+ sync_state_directory(&fixture.paths);
+ let error = fixture
+ .recover()
+ .expect_err("orphan recovery evidence must fail closed");
+ assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
+ assert!(artifact_path(&fixture.paths, name).exists());
+ }
+ }
+
+ fn service_paths(root: &Path) -> ServiceSqlitePaths {
+ let context = RuntimeContext::resolve(
+ &RadrootsPathResolver::new(RadrootsPlatform::Linux, RadrootsHostEnvironment::default()),
+ RuntimeContextBootstrap::new(
+ RadrootsPathProfile::RepoLocal,
+ Some(root.to_path_buf()),
+ RuntimeContextSource::BootstrapCli,
+ RuntimeContextSource::BootstrapCli,
+ )
+ .expect("bootstrap"),
+ ServiceId::new("myc").expect("service"),
+ InstanceId::new("recovery").expect("instance"),
+ )
+ .expect("runtime context");
+ ServiceSqlitePaths::from_runtime_context(&context).expect("SQLite paths")
+ }
+
+ fn database_metadata(paths: &ServiceSqlitePaths) -> ServiceDatabaseMetadata {
+ ServiceDatabaseMetadata::new(
+ paths,
+ SourceGeneration::new([7; 32]).expect("source generation"),
+ NonZeroU32::new(1).expect("schema version"),
+ 1_700_000_000_000,
+ ServiceSqliteApplicationId::new(0x5244_5351).expect("application ID"),
+ )
+ .expect("metadata")
+ }
+
+ fn write_new(path: &Path, bytes: &[u8]) -> File {
+ let mut file = OpenOptions::new()
+ .read(true)
+ .write(true)
+ .create_new(true)
+ .mode(0o600)
+ .open(path)
+ .expect("create artifact");
+ file.set_permissions(fs::Permissions::from_mode(0o600))
+ .expect("set artifact mode");
+ file.write_all(bytes).expect("write artifact");
+ file.sync_all().expect("sync artifact");
+ file
+ }
+
+ fn expectation(path: &Path) -> RestoreArtifactExpectation {
+ let file = File::open(path).expect("open artifact");
+ let status = fstat(&file).expect("artifact status");
+ RestoreArtifactExpectation::new(
+ u64::try_from(status.st_dev).expect("device"),
+ status.st_ino,
+ u64::try_from(status.st_size).expect("length"),
+ hash_exact(
+ &file,
+ u64::try_from(status.st_size).expect("positive artifact length"),
+ )
+ .expect("artifact digest"),
+ )
+ .expect("artifact expectation")
+ }
+
+ fn artifact_path(paths: &ServiceSqlitePaths, name: &str) -> PathBuf {
+ paths.state_database().with_file_name(name)
+ }
+
+ fn sync_state_directory(paths: &ServiceSqlitePaths) {
+ File::open(paths.state_database().parent().expect("state directory"))
+ .expect("open state directory")
+ .sync_all()
+ .expect("sync state directory");
+ }
+
+ fn assert_no_recovery_evidence(paths: &ServiceSqlitePaths) {
+ for name in [
+ STAGED_FILE_NAME,
+ BACKUP_FILE_NAME,
+ MARKER_FILE_NAME,
+ MARKER_NEXT_FILE_NAME,
+ ] {
+ assert!(!artifact_path(paths, name).exists(), "unexpected {name}");
+ }
+ }
+}
diff --git a/crates/service_sqlite/src/restore/stage.rs b/crates/service_sqlite/src/restore/stage.rs
@@ -2056,20 +2056,18 @@ mod tests {
)
);
- for mode in [OpenMode::ReadWriteExisting, OpenMode::ReadOnlyInspection] {
- let error = crate::open::open_existing_connection_pool(
- &fixture.paths,
- &fixture.identity,
- &fixture.migrations,
- &fixture.schema,
- mode,
- ServiceSqliteConnectionOptions::reviewed(),
- )
- .await
- .err()
- .expect("unresolved recovery must reject open");
- assert_eq!(error.kind(), ServiceSqliteErrorKind::Recovery);
- }
+ let read_only = crate::open::open_existing_connection_pool(
+ &fixture.paths,
+ &fixture.identity,
+ &fixture.migrations,
+ &fixture.schema,
+ OpenMode::ReadOnlyInspection,
+ ServiceSqliteConnectionOptions::reviewed(),
+ )
+ .await
+ .err()
+ .expect("read-only open must not recover");
+ assert_eq!(read_only.kind(), ServiceSqliteErrorKind::Recovery);
let initialize = crate::initialize_database(
&fixture.paths,
OpenMode::Initialize,
@@ -2080,6 +2078,30 @@ mod tests {
.await
.expect_err("unresolved recovery must reject initialization");
assert_eq!(initialize.kind(), ServiceSqliteErrorKind::Recovery);
+
+ let writable = crate::open::open_existing_connection_pool(
+ &fixture.paths,
+ &fixture.identity,
+ &fixture.migrations,
+ &fixture.schema,
+ OpenMode::ReadWriteExisting,
+ ServiceSqliteConnectionOptions::reviewed(),
+ )
+ .await
+ .expect("writable open must recover exact finalization evidence");
+ writable.close().await.expect("close recovered pool");
+ assert_eq!(
+ fs::read(fixture.paths.state_database()).expect("recovered live database"),
+ replacement
+ );
+ for name in [
+ BACKUP_FILE_NAME,
+ MARKER_FILE_NAME,
+ MARKER_NEXT_FILE_NAME,
+ STAGED_FILE_NAME,
+ ] {
+ assert!(!recovery_path(&fixture, name).exists(), "unexpected {name}");
+ }
reset_finalize_controls();
}
diff --git a/crates/service_sqlite/tests/package_boundary.rs b/crates/service_sqlite/tests/package_boundary.rs
@@ -18,6 +18,7 @@ const MIGRATION_SOURCE: &str = include_str!("../src/migration.rs");
const OPEN_SOURCE: &str = include_str!("../src/open.rs");
const RESTORE_MARKER_SOURCE: &str = include_str!("../src/restore/marker.rs");
const RESTORE_FINALIZE_SOURCE: &str = include_str!("../src/restore/finalize.rs");
+const RESTORE_RECOVER_SOURCE: &str = include_str!("../src/restore/recover.rs");
const RESTORE_ROOT_SOURCE: &str = include_str!("../src/restore/mod.rs");
const RESTORE_STAGE_SOURCE: &str = include_str!("../src/restore/stage.rs");
const STATUS_SOURCE: &str = include_str!("../src/status.rs");
@@ -179,9 +180,22 @@ fn service_sqlite_is_unpublished_lint_governed_and_dependency_bounded() {
"Once `prepared` is durable, staged-artifact cleanup is disarmed",
"unknown immediate outcome",
"retains the old live database and final marker",
- "every initialized, read-write-existing, or read-only-inspection pool open refuses",
- "marker or marker-scratch as `Recovery` before opening SQLite",
- "does not delete evidence, roll back, reconcile an interruption, or reopen the database",
+ "Read-write-existing open is the sole recovery path",
+ "before opening SQLite",
+ "Read-only inspection, initialization, and an initialized open never recover",
+ "reject any stage, backup, marker, or marker scratch as `Recovery` without mutation",
+ "Recovery uses exact topology as the durable authority",
+ "rolls back by removing only the exact stage and then the marker",
+ "recovery advances and rolls forward",
+ "removes the exact old backup before retiring the marker",
+ "repeated recovery is idempotent",
+ "A marker scratch is admitted only when it is the canonical one-edge successor",
+ "Recovery removes only the exact bound scratch inode",
+ "then reproduces the transition through the governed marker-advance path",
+ "preserves the valid current marker if the scratch pathname was replaced",
+ "Recovery has no await point or hidden task",
+ "each synchronous filesystem step and its authority checks complete",
+ "Finalization itself does not reconcile or reopen the database",
] {
assert!(
readme_words.contains(required),
@@ -675,18 +689,73 @@ fn service_sqlite_is_unpublished_lint_governed_and_dependency_bounded() {
"Step 068 restore finalization source contains deferred or raw authority `{forbidden}`"
);
}
+ for required in ["refuse_unresolved_recovery", "recover_for_open"] {
+ assert!(
+ RESTORE_ROOT_SOURCE.contains(required),
+ "Step 068 recovery-open guard is missing `{required}`"
+ );
+ }
+ assert!(OPEN_SOURCE.contains("crate::restore::refuse_unresolved_recovery"));
+ assert!(OPEN_SOURCE.contains("crate::restore::recover_for_open(paths, identity, authority)"));
+
+ let restore_recover_production = RESTORE_RECOVER_SOURCE
+ .split_once("#[cfg(all(test, any(target_os = \"linux\", target_os = \"macos\")))]")
+ .map(|(production, _)| production)
+ .expect("restore recovery source must keep tests separated");
for required in [
- "refuse_unresolved_recovery",
+ "pub(crate) fn recover_for_open(",
+ "WriterAuthority",
+ "RestoreMarkerBinding::load_for_recovery",
+ "matches_identity(identity)",
+ "interrupted_transition(paths, authority)",
+ "promote_interrupted_transition",
+ "advance_for_recovery",
+ "RestoreRecoveryPhase::Prepared",
+ "RestoreRecoveryPhase::LiveRetained",
+ "RestoreRecoveryPhase::ReplacementInstalled",
+ "RenameFlags::NOREPLACE",
+ "verify_named_artifact(",
+ "hash_exact(",
+ "remove_exact_artifact(",
+ "marker.retire(paths, authority)",
+ "state.sqlite-wal",
+ "state.sqlite-shm",
+ "state.sqlite-journal",
"ServiceSqliteErrorKind::Recovery",
+ ] {
+ assert!(
+ restore_recover_production.contains(required),
+ "Step 069 restore recovery source is missing `{required}`"
+ );
+ }
+ for forbidden in [
+ "pub fn recover",
+ "pub async fn recover",
+ "pub struct PendingRestore",
+ "sqlx::",
+ "rusqlite::",
+ "tokio::",
+ "spawn_blocking",
+ "tokio::time::timeout",
+ "remove_dir_all",
+ "SystemTime::now",
+ ] {
+ assert!(
+ !restore_recover_production.contains(forbidden) && !ROOT.contains(forbidden),
+ "Step 069 recovery exposes or implements forbidden authority `{forbidden}`"
+ );
+ }
+ for required in [
+ "STAGED_FILE_NAME",
+ "BACKUP_FILE_NAME",
"MARKER_FILE_NAME",
"MARKER_NEXT_FILE_NAME",
] {
assert!(
- RESTORE_ROOT_SOURCE.contains(required),
- "Step 068 recovery-open guard is missing `{required}`"
+ RESTORE_RECOVER_SOURCE.contains(required),
+ "Step 069 refusal inventory is missing `{required}`"
);
}
- assert!(OPEN_SOURCE.contains("crate::restore::refuse_unresolved_recovery"));
assert!(ROOT.contains(
"pub use restore::{StagedServiceRestore, finalize_staged_restore, stage_verified_restore};"
));