README.md (23142B)
1 # radroots_service_sqlite 2 3 `radroots_service_sqlite` is the unpublished, native SQLite-mechanism crate for 4 Radroots services. It provides narrow, reusable building blocks for exclusive 5 writer authority, instance locking, versioned schema mechanics, bounded 6 transactions, immutable service-instance database identity, integrity checks, 7 backup, restore, and passive storage status. 8 9 `ServiceSqliteHost` is the only public connection host. Its SQLx pool and raw 10 connections are sealed inside the crate. Services run typed SQLx queries through 11 the borrowed `ServiceSqliteTransaction` executor supplied by 12 `ServiceSqliteHost::transaction`; transaction begin, commit, rollback, policy 13 validation, attached-database exclusion, and cancellation quarantine remain 14 runner-owned. Writable host opening finishes every pending governed migration 15 before returning, and read-only inspection opens only current migration and 16 schema state. 17 18 Create-new state uses the borrowed `ServiceSqliteInitializer` executor rather 19 than a database path or raw connection. The runner reserves and retains the 20 exact canonical file, opens SQLite only through that retained descriptor, owns 21 `BEGIN IMMEDIATE`, and uses a private memory journal so descriptor-bound 22 initialization creates no path-derived SQLite sidecar. It commits the product 23 schema, shared `radroots_service_metadata`, empty v1 `schema_migrations` 24 ledger, and exact schema-catalog verification in one transaction. The 25 initializer screens the same closed transaction-control and attachment 26 inventory as host transactions; an ignored rejection still prevents commit. 27 Callback failure or cancellation cannot return a reusable connection or 28 publish a partial database. 29 30 Interactive callers may use `ServiceSqliteHost::open_or_initialize`. One held 31 writer authority and an exclusive create decide whether the sealed initializer 32 runs or the exact existing database is opened. The existing branch never runs 33 the initializer. Callers therefore do not use pathname probes, error-text 34 matching, recursive directory creation, permission repair, or direct SQLx 35 connections to choose the bootstrap branch. Success returns an 36 `OpenedServiceDatabase` that binds the retained host to the actual verified 37 metadata from the selected branch, including the persisted source generation 38 when state already existed. 39 40 Existing databases can be admitted without a caller guessing their stored 41 source generation. `ExistingServiceDatabaseIntent` seals the canonical 42 service and instance, supported schema ceiling, and SQLite application ID; 43 `ServiceSqliteHost::open_read_write_existing_with_intent` and 44 `ServiceSqliteHost::open_read_only_inspection_with_intent` discover and verify 45 the actual immutable metadata while retaining the corresponding writer or 46 inspection authority. Success returns an `OpenedServiceDatabase`, 47 which keeps the host and verified metadata inseparable until the caller 48 consumes them together. Recovery remains fail closed and may use this intent 49 only to discover the marker-bound generation; every other identity dimension, 50 artifact binding, migration prefix, and schema catalog remains governed. 51 52 Service-controlled SQL is screened before SQLite compilation through both the 53 borrowed transaction executor and migration callback executor. The closed 54 statement-control inventory is `PRAGMA`, `ATTACH`, `DETACH`, `BEGIN`, `COMMIT`, 55 `END`, `ROLLBACK`, `SAVEPOINT`, and `RELEASE`, regardless of case, whitespace, 56 comments, multiple statements, or prepared-query entry point. Rejection is 57 sticky for the transaction, so ignoring the immediate SQL error cannot permit 58 commit. Runner-owned setup remains private, and the complete connection policy 59 is revalidated before commit and before a connection can return to the pool. 60 61 Persisted SQLite text and blob values are admitted through bounded projections 62 before Rust decoding. Service and instance identifiers, source generations, 63 migration names, checksums, build identity, schema text, and database inventory 64 values carry an exact SQLite type, reported byte length, capped byte prefix, 65 and bounded row count. Integrity diagnostics use a borrowed-byte cap over the 66 exact `PRAGMA integrity_check(1)` operation and never convert an unbounded 67 diagnostic to UTF-8. 68 69 Cancelling a host transaction before the runner enables outer commit 70 quarantines its connection and leaves no authoritative transaction effect. A 71 service-operation error is returned only after rollback is confirmed; an 72 unconfirmed rollback is reported as `RollbackFailed`. Cancelling once outer 73 commit begins yields no result and must be treated as an unknown commit outcome. 74 Both that case and `CommitOutcomeUnknown` require rereading authoritative state 75 before an idempotent retry. 76 77 Every host must be closed explicitly with `ServiceSqliteHost::close`. Close 78 permanently stops new transaction admission, drains transactions that were 79 already admitted, and is safe to call sequentially or concurrently. Writable 80 close applies the fixed `PRAGMA wal_checkpoint(TRUNCATE)` policy, requires an 81 unblocked checkpoint, closes the private checkpoint connection, and explicitly 82 releases writer authority. Read-only inspection close drains its pool and 83 releases its shared inspection guard without checkpointing or mutating the 84 database or filesystem. Cancelling close before terminal completion leaves the 85 host non-admitting and retains authority; the private connect, checkpoint, and 86 explicit connection-close driver remains host-owned so a later call resumes 87 close without losing the SQLite handle or its close proof. Once authority 88 release is proven, the stable outer result is cached for every later call. 89 Dropping a host performs no asynchronous close work and is not proof that the 90 governed checkpoint and authority-release sequence completed. 91 92 `ServiceBackupManifest` is the stable model-only v1 backup identity. It admits 93 only compact canonical UTF-8 JSON in the frozen field order, capped at 1,024 94 bytes, and computes the external manifest SHA-256 over those exact bytes. The 95 v1 member array contains exactly one `state.sqlite` member with a nonzero byte 96 length and lowercase SHA-256; service, instance, nonzero source generation, 97 state schema, and injected creation time are explicit. SQLite and foreign-key 98 integrity are exactly `ok`, and protected material is always excluded. 99 Parsing proves only the strict structural and canonical contract. It rejects 100 unknown, duplicate, null, reordered, whitespace-altered, or version-drifted 101 input; member bytes, digest, SQLite identity, and actual integrity remain the 102 separate backup-verification boundary. Constructing or parsing the manifest 103 model performs no filesystem or SQLite work. 104 105 Writable hosts provide `ServiceSqliteHost::capture_online_backup` for one 106 incremental, point-in-time SQLite capture at a time. The caller supplies an 107 injected creation time and the exact new absolute staging-directory path under 108 an existing owner-controlled parent; capture creates that directory with mode 109 `0700` and its sole `state.sqlite` member with mode `0600`. It uses SQLite's 110 online-backup API without checkpointing or copying the live source file, then 111 requires exact service metadata, bounded `integrity_check`, an empty 112 `foreign_key_check`, a singleton member inventory, SHA-256, and file, staging, 113 and parent synchronization before returning the canonical manifest in memory. 114 No manifest file, bundle identifier, credential, or protected material is 115 written to the staging directory. 116 117 Capture rejects read-only, closing, unsupported, colliding, or concurrent 118 admission before publishing a result. Dropping the capture future requests 119 cancellation; the blocking worker retains its checked-out pool admission, 120 writer authority, and exact staging identities until SQLite handles are closed 121 and cleanup completes. Host close therefore drains capture and cancellation 122 cleanup before it checkpoints or releases authority. Capture has no hidden 123 timeout: callers own any deadline by cancelling the future. A completed capture 124 is still untrusted backup input until the separate verifier binds its manifest, 125 member bytes, expected intent, application metadata, and integrity; capture 126 does not provide restore or replacement behavior. 127 128 `verify_backup_bundle` is the synchronous, task-free boundary for an untrusted 129 manifest and bundle. The caller supplies the independently protected manifest 130 SHA-256, expected service database identity, and a positive maximum state-file 131 size. Verification requires canonical manifest bytes, exact service, instance, 132 source-generation, schema, and application intent, a restrictive owner-only 133 directory containing only `state.sqlite`, the exact bounded member length and 134 digest, immutable read-only/query-only SQLite access, main-only attachment, 135 bounded application metadata, `integrity_check(1)`, and an empty foreign-key 136 check. It performs no filesystem mutation and does not create a task or hidden 137 deadline. 138 139 Success returns a non-forgeable `VerifiedServiceBackup` that retains the exact 140 verified directory and member descriptors while exposing only the canonical 141 manifest and actual database metadata. It is not restore or replacement 142 authority, and it exposes no path or raw handle. Later restore work must copy 143 from the retained member and reverify the staged copy under its own supervised 144 blocking worker and deadline; pathname verification alone is insufficient. 145 146 Restore crash recovery uses a private sealed v1 marker stored beside canonical 147 service state. Its fixed layout names the live `state.sqlite`, staged 148 `state.restore-staged.sqlite`, retained `state.restore-backup.sqlite`, durable 149 `state.restore-marker.v1`, and create-new update scratch 150 `state.restore-marker.v1.next`. The compact canonical JSON is capped at 2,048 151 bytes and binds typed database intent, the protected source-manifest digest, 152 and exact live, staged, and retained-backup device, inode, length, and SHA-256 153 expectations. A domain-separated checksum binds the canonical fields and 154 detects corruption; it is not an authenticity credential. 155 156 The only legal durable sequence is `prepared` to `live_retained` to 157 `replacement_installed`; repeating the current phase is byte-idempotent and 158 every skip, reversal, or post-install transition fails closed. Marker files are 159 descriptor-relative, no-follow, single-link, owner-owned regular files with 160 mode `0600`. Creation synchronizes the file and state directory. Advancement 161 compare-and-reloads the current bytes, writes and synchronizes the fixed 162 create-new scratch, atomically replaces the marker, synchronizes the directory, 163 and reopens the exact new bytes. Stale scratch, tamper, collision, binding 164 replacement, insecure directory, or malformed marker remains evidence and 165 fails closed; reads do not repair or remove it. 166 167 This marker checkpoint does not stage, copy, open, rename, replace, or delete a 168 database. Later restore staging must consume a retained verified backup; 169 replacement must invoke the marker sequence around its governed renames; and 170 open-time recovery must reconcile durable marker and artifact identities before 171 removing any recovery evidence. No marker type, path, raw descriptor, or store 172 operation is public API. 173 174 `stage_verified_restore` is the offline boundary between retained backup proof 175 and live-state replacement. It first validates the expected identity plus exact 176 migration and schema catalogs, then acquires exclusive writer authority. A live 177 writable or read-only host, a live WAL/shared-memory/journal sidecar, an existing 178 stage, or any marker/retained-backup evidence fails closed. The only created 179 artifact is the fixed adjacent `state.restore-staged.sqlite`, opened create-new, 180 no-follow, owner-only, and single-link with mode `0600`. 181 182 Staging copies the exact manifest-bound bytes from the verifier's retained 183 member descriptor with a fixed-size buffer and digest, synchronizes the staged 184 file, and opens SQLite only through the retained staged descriptor. It then 185 rechecks immutable application metadata, the exact applied migration prefix and 186 schema-object catalog at the backup's actual supported version, main-only 187 read-only/query-only connection policy, bounded `integrity_check(1)`, and empty 188 `foreign_key_check`. A final retained-descriptor hash and file plus state- 189 directory synchronization precede success. Live database bytes, identity, 190 permissions, and timestamps remain untouched. 191 192 Success returns a sealed non-cloneable `StagedServiceRestore` that retains 193 writer authority and the exact staging identities. Dropping it attempts an 194 identity-checked unlink and state-directory synchronization before releasing 195 authority; cleanup failure leaves staging or recovery evidence that later 196 admission rejects. Cancelling the async operation requests bounded copy 197 cancellation; any detached work retains authority and exact cleanup ownership 198 until it ends. 199 The operation has no hidden timeout. It does not create or advance a recovery 200 marker, rename or retain live state, install a replacement, or authorize reopen; 201 those operations remain the finalization and recovery checkpoints. 202 203 `finalize_staged_restore` consumes that sealed stage in an owned blocking 204 worker. The stage has already bound the exact live inode, length, and digest 205 that will be retained. Finalization revalidates both retained descriptors, 206 creates and synchronizes the `prepared` marker, and only then disarms automatic 207 stage cleanup. It renames live to `state.restore-backup.sqlite` and staged to 208 live with descriptor-relative no-replace operations. Each rename is followed 209 by exact inode and hash verification, state-directory synchronization, and the 210 corresponding marker advance to `live_retained` or 211 `replacement_installed`. 212 213 Cancellation observed before the worker's atomic commit-ownership handoff 214 leaves live state untouched and attempts exact stage cleanup. Caller-task loss 215 after that handoff has an unknown immediate outcome, including the short 216 interval before `prepared` becomes durable; the worker retains writer 217 authority until it either fails before durability or establishes recovery 218 evidence and continues. Once `prepared` is durable, staged-artifact cleanup is 219 disarmed and the bound stage remains available after every later error. 220 Success returns no database host, retains the old live database and final 221 marker, and requires a new open. Read-write-existing open is the sole recovery 222 path. Under exclusive writer authority and before opening SQLite, it validates 223 the marker's exact service, instance, source generation, application ID, schema 224 ceiling, artifact identities, lengths, digests, restrictive modes, and the 225 absence of database sidecars. Read-only inspection, initialization, and an 226 initialized open never recover; they reject any stage, backup, marker, or marker 227 scratch as `Recovery` without mutation. 228 229 Recovery uses exact topology as the durable authority. `prepared` with the old 230 live database still installed rolls back by removing only the exact stage and 231 then the marker. Once the exact old live inode has reached the backup name, 232 recovery advances and rolls forward. A lagging `live_retained` phase installs or 233 recognizes the exact replacement, advances to `replacement_installed`, then 234 removes the exact old backup before retiring the marker. Interrupted rollback 235 and final cleanup accept only the corresponding already-absent exact artifact, 236 so repeated recovery is idempotent. Every other topology, sidecar, replacement, 237 link, mode, owner, length, digest, identity, or directory-authority mismatch 238 fails closed and preserves the evidence. 239 240 A marker scratch is admitted only when it is the canonical one-edge successor 241 of the current marker and the artifact topology already proves that successor. 242 Recovery removes only the exact bound scratch inode, synchronizes that removal, 243 then reproduces the transition through the governed marker-advance path. This 244 preserves the valid current marker if the scratch pathname was replaced. 245 Orphaned, malformed, skipped, same-phase, terminal, mismatched, or 246 topology-inconsistent scratch is never deleted or reinterpreted. Recovery has no 247 await point or hidden task: once a writable open is polled, each synchronous 248 filesystem step and its authority checks complete before the open can be 249 cancelled. A later cancelled SQLite open is retried by rereading the already 250 durable, marker-free state. Finalization itself does not reconcile or reopen the 251 database. 252 253 `ServiceSqliteHost::inspect_integrity` is the explicit active operator check. 254 It is available on initialized, writable-existing, and read-only inspection 255 hosts, admits at most one check per host, and uses one deferred read transaction 256 as the SQLite snapshot. The caller injects a positive wall-clock 257 `IntegrityCheckedAtUnixMs`; the library does not read an ambient clock or 258 create a timer. The completed report contains only `verified` or `failed` for 259 SQLite integrity and foreign keys, plus at most the fixed 260 `sqlite_integrity_failed` and `foreign_key_violation` diagnostic codes in that 261 canonical order. It can be projected to the passive `StorageIntegrity` 262 vocabulary, but the library does not persist or cache the report. 263 264 The check never publishes raw SQLite diagnostics, table or row identity, 265 filesystem paths, SQL, or dependency errors. Inability to execute, decode, or 266 finish either bounded check is an `Integrity` error rather than a fabricated 267 completed result. Authority is revalidated after every await and has precedence 268 over integrity classification. The operation has no hidden timeout or task. 269 Callers own a positive monotonic deadline by dropping the future; cancellation 270 returns no report, writes nothing, quarantines the checked-out connection, and 271 leaves it in a host-owned close driver. Retry or host close explicitly awaits 272 that retained close future until the prior SQLite worker terminates before any 273 new check or authority release. A retry uses a newly injected wall-clock time. 274 The strict backup and restore integrity verifier remains a separate fail-closed 275 boundary. 276 277 State-filesystem capacity inspection is an explicit synchronous input for 278 doctor checks and authoritative admission. `MinimumFreeBytes` must be supplied 279 and is constrained to `1..=i64::MAX`; it has no default. The value 280 `268435456` is the exact governed configuration and test vector, not an 281 implicit universal threshold. On Linux and macOS the platform adapter opens the 282 owner-owned state directory that is not group/other writable without following 283 links, retains and revalidates its identity, and uses `fstatvfs` to measure 284 bytes available to the unprivileged service user. The current native 285 qualification matrix is macOS aarch64 and Linux x86_64; other platforms fail 286 closed and successful compilation outside that matrix is not support evidence. 287 288 A successful immutable snapshot is `ready` when available bytes are greater 289 than or equal to the configured minimum and `low_disk` when they are below it. 290 Low disk rejects or pauses new authoritative admission; measurement failure is 291 a typed unavailable error and is never fabricated as low-disk evidence. A 292 consumer may cache the successful snapshot and later project low disk to the 293 stable `database_low_disk` readiness reason. `/readyz` remains passive and must 294 read only that caller-owned cached state; it never invokes the capacity 295 adapter. The measurement is advisory rather than a space reservation and does 296 not guarantee a later write. 297 298 Capacity inspection is host-independent and performs no database open, pool 299 operation, SQLite query, filesystem mutation, ambient time read, timer, task, 300 or hidden sampling. Service configuration, threshold defaults, cache refresh, 301 status persistence, admission wiring, and route projection remain consumer 302 responsibilities. 303 304 Durability fault injection is a private test-only mechanism. A closed 305 instance-scoped controller can arm exactly one named before/after boundary and 306 returns one injected error the first time that boundary is reached; later hits 307 are no-ops. The complete inventory covers database initialization, runner-owned 308 transaction begin and commit, online-backup creation/copy/synchronization, 309 restore-marker creation and advancement, both restore rename/synchronization 310 steps, and explicit host drain/checkpoint/connection-close/authority-release. 311 An ordinary controller has zero behavior, and no failpoint type or selector is 312 exported from the crate root. There is no process-global failpoint state, 313 environment or configuration selector, Cargo feature, hidden task, timer, 314 panic, or process-exit behavior. These deterministic in-process edges qualify 315 error ordering, rollback, cleanup, recovery evidence, and one-shot retry 316 semantics. 317 318 Process-crash qualification remains test-only and reuses Cargo's private 319 library and integration-test binaries; the crate ships no helper executable or 320 signal handler. A parent sends one bounded temporary root over stdin, waits for 321 a fixed stdout readiness token from an occurrence-aware failpoint barrier, and 322 then issues `SIGKILL`. The suite proves cross-process writer contention and 323 lock release plus five restore boundaries: orphan-stage refusal before a 324 durable marker, prepared rollback, interrupted marker-scratch promotion, 325 installed-replacement recovery, and terminal-marker cleanup. A permissive 326 child umask cannot broaden the fixed `0700` state directory or `0600` lock, 327 database, stage, backup, marker, and marker-scratch artifacts. Linux execution 328 on x86_64 is required for OS-level qualification; macOS aarch64 execution on 329 the current machine is developer evidence. No other platform or architecture 330 is an active qualification gate. These tests exercise process death at named 331 durable edges and do not claim abrupt power-loss or storage-device durability 332 behavior. 333 334 The crate owns mechanics only. Service-specific tables, SQL, repositories, 335 backup content policy, identity material, process lifecycle, and readiness 336 policy remain with the consuming service. The crate does not provide callers 337 with raw database authority. 338 339 Publication is disabled. The package is not part of the public Radroots crate 340 release closure. 341 342 ## Public API and package boundary 343 344 This unpublished `private_runtime`, `package_private` crate owns only 345 service-neutral SQLite mechanics. Its built-in persistent objects are the 346 shared `radroots_service_metadata` and `schema_migrations` tables plus their 347 governed immutability triggers. Every service-owned table, index, trigger, 348 migration statement, and schema policy is supplied through the caller-owned 349 migration and schema catalogs; no product identifier or product table belongs 350 in this package. 351 352 The crate-root exports are frozen in the reviewed 353 [service-SQLite API baseline](../../contracts/api_baselines/radroots_service_sqlite.txt). 354 Raw pools, pooled or direct connections, transaction-control handles, and 355 dependency re-exports are forbidden. The deliberate narrow exception is the 356 `sqlx::Executor` implementation for borrowed 357 `&mut ServiceSqliteInitializer<'_>` and `&mut ServiceSqliteTransaction<'_>` 358 values. They permit compile-time typed queries. The crate retains connection 359 ownership and sole begin, commit, rollback, policy, and cancellation authority.