diff --git a/crates/openshell-driver-podman/README.md b/crates/openshell-driver-podman/README.md index abf2147513..8590f22c46 100644 --- a/crates/openshell-driver-podman/README.md +++ b/crates/openshell-driver-podman/README.md @@ -45,6 +45,25 @@ workspace. The workload starts before the supervisor so its user namespace exist when the supervisor joins it; a stopped supervisor resolves that namespace again on its next start. +Lifecycle changes and workload containment share a per-sandbox mutex across +all clones of the driver. Restart holds it until both container starts finish +(or fail). Watch events, initial reconciliation, and running-sandbox get/list +inspections wait for that operation and re-inspect before deciding whether to +stop a workload. An exited supervisor during the workload-first restart window +therefore cannot stop the new workload. Genuine supervisor loss still stops the +workload once the lifecycle operation finishes. Cancellation drops the guard; +there is no persistent restart exemption to clear before retrying. This +coordination is local to one driver instance and its clones. Periodic resource +admission reconciliation also holds the gate across fresh ownership checks, +validation, and both containment stops. Reused container IDs are not treated as +run identities: containment uses current state read under the gate. Exit-event +fences separately protect published observations from a previous run. + +Independent driver processes and external Podman commands do not share this +gate. Cancelling a future releases the local guard but cannot retract an HTTP +mutation already accepted by Podman; fencing those remote operations requires +a separate runtime ownership or cancellation-completion contract. + The runtime must pass the sandbox's unprivileged enforcement probe, including nested seccomp notification and Landlock. Unsupported runtime defaults fail closed; do not switch to an unconfined profile or add capabilities. diff --git a/crates/openshell-driver-podman/src/driver.rs b/crates/openshell-driver-podman/src/driver.rs index d3023cddd8..c4ecf9cbb3 100644 --- a/crates/openshell-driver-podman/src/driver.rs +++ b/crates/openshell-driver-podman/src/driver.rs @@ -847,6 +847,26 @@ impl PodmanComputeDriver { return; }; for entry in entries.iter().filter(|entry| entry.state == "running") { + let Some(sandbox_id) = entry.labels.get(LABEL_SANDBOX_ID) else { + continue; + }; + let _operation = self.lifecycle_event_fences.lock(sandbox_id).await; + // The list predates the guard. Never carry its running state or a + // grant denial across a concurrent lifecycle operation. + let Ok(current) = self.client.inspect_container(&entry.id).await else { + continue; + }; + if !current.state.running + || current.id != entry.id + || current.config.labels.get(LABEL_SANDBOX_ID) != Some(sandbox_id) + || current + .config + .labels + .get(crate::isolation::LABEL_ROLE) + .is_none_or(|role| role != "sandbox") + { + continue; + } if let Err(ComputeDriverError::Precondition(reason)) = self.admit_container_resources(&entry.id).await { @@ -889,6 +909,7 @@ impl PodmanComputeDriver { // Validate the composed container name early, before creating any // resources (volume), so we don't leave orphans when the name is // invalid. + let _operation = self.lifecycle_event_fences.lock(&sandbox.id).await; let name = validated_container_name(sandbox)?; let validated = self.validated_sandbox_create(sandbox).await?; @@ -1395,6 +1416,7 @@ impl PodmanComputeDriver { )] pub async fn stop_sandbox(&self, sandbox_id: &str) -> Result<(), ComputeDriverError> { let span_status = openshell_otel::ErrorStatusGuard::current(); + let _operation = self.lifecycle_event_fences.lock(sandbox_id).await; let container = self.find_container(sandbox_id).await?; let supervisor = crate::isolation::supervisor_name(sandbox_id); match self @@ -1470,6 +1492,7 @@ impl PodmanComputeDriver { encoded_authentication: &[u8], ) -> Result<(), ComputeDriverError> { let span_status = openshell_otel::ErrorStatusGuard::current(); + let _operation = self.lifecycle_event_fences.lock(sandbox_id).await; let generation = openshell_core::sandbox_generation::SandboxGenerationId::parse( generation_id.to_string(), ) @@ -1596,6 +1619,7 @@ impl PodmanComputeDriver { )] pub async fn delete_sandbox(&self, sandbox_id: &str) -> Result { let span_status = openshell_otel::ErrorStatusGuard::current(); + let _operation = self.lifecycle_event_fences.lock(sandbox_id).await; if sandbox_id.is_empty() { return Err(ComputeDriverError::Precondition( "sandbox id is required".into(), @@ -1723,11 +1747,13 @@ impl PodmanComputeDriver { return Ok(None); }; if entry.state == "running" { - Ok(watcher::inspect_workload(&self.client, &entry.id) - .await - .ok() - .and_then(|inspect| driver_sandbox_from_inspect(&inspect)) - .or_else(|| driver_sandbox_from_list_entry(entry))) + Ok( + watcher::inspect_workload(&self.client, &entry.id, &self.lifecycle_event_fences) + .await + .ok() + .and_then(|inspect| driver_sandbox_from_inspect(&inspect)) + .or_else(|| driver_sandbox_from_list_entry(entry)), + ) } else { Ok(driver_sandbox_from_list_entry(entry)) } @@ -1748,7 +1774,13 @@ impl PodmanComputeDriver { for entry in &entries { if entry.state == "running" { // Running containers need inspect for health check status. - match watcher::inspect_workload(&self.client, &entry.id).await { + match watcher::inspect_workload( + &self.client, + &entry.id, + &self.lifecycle_event_fences, + ) + .await + { Ok(inspect) => { if let Some(sandbox) = driver_sandbox_from_inspect(&inspect) { sandboxes.push(sandbox); @@ -2164,6 +2196,137 @@ mod tests { assert!(requests[1].contains(&crate::isolation::supervisor_name("sandbox-1"))); } + // The stub pauses supervisor start after workload start has completed, + // while a clone performs the same reconciliation used by the watcher. + // This tests the real driver API over HTTP, without claiming live Podman qualification. + #[tokio::test] + async fn restart_and_reconciliation_do_not_stop_the_new_workload() { + restart_overlaps_reconciliation(false, false).await; + } + + #[tokio::test] + async fn cancelled_restart_releases_containment_and_allows_retry() { + restart_overlaps_reconciliation(true, false).await; + } + + #[tokio::test] + async fn failed_supervisor_start_releases_containment_and_allows_retry() { + restart_overlaps_reconciliation(false, true).await; + } + + async fn restart_overlaps_reconciliation(cancel: bool, fail: bool) { + use tokio::sync::Notify; + let reached = Arc::new(Notify::new()); + let release = Arc::new(Notify::new()); + let running = r#"{"Id":"ctr-1","Name":"sandbox","State":{"Status":"running","Running":true},"Config":{"Labels":{"openshell.ai/sandbox-id":"sandbox-1","openshell.ai/isolation-role":"sandbox"}}}"#; + let exited = r#"{"Id":"ctr-1","Name":"sandbox","State":{"Status":"exited","Running":false},"Config":{}}"#; + let mut responses = vec![ + StubResponse::new(StatusCode::OK, r#"[{"Id":"ctr-1","State":"stopped"}]"#), + StubResponse::new(StatusCode::OK, exited), + ]; + let mut restart = restart_responses(); + *restart.last_mut().unwrap() = StubResponse::new( + if fail { + StatusCode::INTERNAL_SERVER_ERROR + } else { + StatusCode::NO_CONTENT + }, + "", + ) + .with_gate(reached.clone(), release.clone()); + responses.extend(restart); + // Identification may run during restart; containment must wait. + responses.push(StubResponse::new(StatusCode::OK, running)); + if fail { + responses.push(StubResponse::new(StatusCode::NO_CONTENT, "")); // start rollback + } + responses.push(StubResponse::new(StatusCode::OK, running)); // fresh snapshot + responses.push(StubResponse::new(StatusCode::OK, if cancel || fail { + r#"{"Id":"supervisor","Name":"supervisor","State":{"Status":"exited","Running":false},"Config":{}}"# + } else { + r#"{"Id":"supervisor","Name":"supervisor","State":{"Status":"running","Running":true,"Health":{"Status":"healthy"}},"Config":{}}"# + })); + if cancel || fail { + responses.push(StubResponse::new(StatusCode::NO_CONTENT, "")); + responses.push(StubResponse::new(StatusCode::OK, exited)); + // A subsequent start through a clone must not inherit a stuck fence. + responses.push(StubResponse::new( + StatusCode::OK, + r#"[{"Id":"ctr-1","State":"stopped"}]"#, + )); + responses.push(StubResponse::new(StatusCode::OK, exited)); + responses.extend(restart_responses()); + } + let (socket, requests, handle) = spawn_podman_stub("restart-watch", responses); + let driver = test_driver(socket); + let starter = driver.clone(); + let start = tokio::spawn(async move { + starter + .start_sandbox( + "sandbox-1", + "generation-1", + &encoded_launch_authentication(), + ) + .await + }); + tokio::time::timeout(Duration::from_secs(5), reached.notified()) + .await + .unwrap(); + let reader = driver.clone(); + let mut inspect = Box::pin(watcher::inspect_workload( + &reader.client, + "ctr-1", + &reader.lifecycle_event_fences, + )); + // Poll through the identification request into the blocked operation lock. + assert!( + tokio::time::timeout(Duration::from_millis(100), &mut inspect) + .await + .is_err() + ); + assert!( + !requests + .lock() + .unwrap() + .iter() + .any(|r| r.contains("/ctr-1/stop")) + ); + if cancel { + start.abort(); + assert!(start.await.unwrap_err().is_cancelled()); + release.notify_one(); + } else { + release.notify_one(); + assert_eq!(start.await.unwrap().is_err(), fail); + } + let inspected = tokio::time::timeout(Duration::from_secs(5), inspect) + .await + .unwrap() + .unwrap(); + assert_eq!(inspected.state.running, !(cancel || fail)); + if cancel || fail { + driver + .clone() + .start_sandbox( + "sandbox-1", + "generation-1", + &encoded_launch_authentication(), + ) + .await + .unwrap(); + } + tokio::time::timeout(Duration::from_secs(5), handle) + .await + .unwrap() + .unwrap(); + let requests = requests.lock().unwrap(); + let stops = requests + .iter() + .filter(|r| r.contains("/ctr-1/stop")) + .count(); + assert_eq!(stops, usize::from(cancel) + 2 * usize::from(fail)); + } + #[tokio::test] async fn stop_and_start_target_the_existing_container() { let (stop_socket, stop_requests, stop_handle) = spawn_podman_stub( @@ -3033,6 +3196,82 @@ mod tests { let _ = fs::remove_file(socket); } + #[tokio::test] + async fn admission_denial_holds_lifecycle_guard_through_both_stops() { + let reached = Arc::new(tokio::sync::Notify::new()); + let release = Arc::new(tokio::sync::Notify::new()); + let (socket, requests, handle) = spawn_podman_stub("admission-gate", vec![ + StubResponse::new(StatusCode::OK, r#"[{"Id":"ctr-1","State":"running","Labels":{"openshell.ai/sandbox-id":"sandbox-1"}}]"#), + StubResponse::new(StatusCode::OK, r#"{"Id":"ctr-1","Name":"sandbox","State":{"Status":"running","Running":true},"Config":{"Labels":{"openshell.ai/sandbox-id":"sandbox-1","openshell.ai/isolation-role":"sandbox"}}}"#), + // Missing admission provenance is a confirmed denial. + StubResponse::new(StatusCode::OK, r#"{"Id":"ctr-1","Name":"sandbox","State":{"Status":"running","Running":true},"Config":{}}"#).with_gate(reached.clone(), release.clone()), + StubResponse::new(StatusCode::NO_CONTENT, ""), + StubResponse::new(StatusCode::NO_CONTENT, ""), + ]); + let driver = PodmanComputeDriver::for_tests(PodmanComputeConfig { + socket_path: Some(socket), + ..Default::default() + }); + let reconciler = driver.clone(); + let reconcile = + tokio::spawn(async move { reconciler.reconcile_resource_admission().await }); + tokio::time::timeout(Duration::from_secs(5), reached.notified()) + .await + .unwrap(); + let mut later_operation = Box::pin(driver.lifecycle_event_fences.lock("sandbox-1")); + assert!( + tokio::time::timeout(Duration::from_millis(10), &mut later_operation) + .await + .is_err() + ); + release.notify_one(); + tokio::time::timeout(Duration::from_secs(5), reconcile) + .await + .unwrap() + .unwrap(); + let guard = tokio::time::timeout(Duration::from_secs(5), later_operation) + .await + .unwrap(); + handle.await.unwrap(); + assert!(requests.lock().unwrap()[3].contains("/ctr-1/stop?timeout=0")); + assert!( + requests.lock().unwrap()[4].contains("/openshell-supervisor-sandbox-1/stop?timeout=0") + ); + drop(guard); + } + + #[tokio::test] + async fn admission_rechecks_stale_running_entries_and_ownership() { + for (state, labels) in [ + ( + serde_json::json!({"Status":"exited","Running":false}), + serde_json::json!({"openshell.ai/sandbox-id":"sandbox-1","openshell.ai/isolation-role":"sandbox"}), + ), + ( + serde_json::json!({"Status":"running","Running":true}), + serde_json::json!({"openshell.ai/sandbox-id":"replacement","openshell.ai/isolation-role":"sandbox"}), + ), + ] { + let (socket, requests, handle) = spawn_podman_stub("admission-stale", vec![ + StubResponse::new(StatusCode::OK, r#"[{"Id":"ctr-1","State":"running","Labels":{"openshell.ai/sandbox-id":"sandbox-1"}}]"#), + StubResponse::new(StatusCode::OK, serde_json::json!({"Id":"ctr-1","Name":"sandbox","State":state,"Config":{"Labels":labels}}).to_string()), + ]); + let driver = PodmanComputeDriver::for_tests(PodmanComputeConfig { + socket_path: Some(socket), + ..Default::default() + }); + driver.reconcile_resource_admission().await; + handle.await.unwrap(); + assert!( + requests + .lock() + .unwrap() + .iter() + .all(|r| r.starts_with("GET ")) + ); + } + } + #[tokio::test] async fn admission_requires_explicit_volume_labels_and_workspace_match() { for (labels, allowed) in [ diff --git a/crates/openshell-driver-podman/src/test_utils.rs b/crates/openshell-driver-podman/src/test_utils.rs index 25cbcb4ac2..e324acae25 100644 --- a/crates/openshell-driver-podman/src/test_utils.rs +++ b/crates/openshell-driver-podman/src/test_utils.rs @@ -23,6 +23,7 @@ pub struct StubResponse { pub body: Bytes, pub delay: Duration, pub archive_members: Option>, + pub gate: Option<(Arc, Arc)>, } impl StubResponse { @@ -32,9 +33,19 @@ impl StubResponse { body: body.into(), delay: Duration::ZERO, archive_members: None, + gate: None, } } + pub fn with_gate( + mut self, + reached: Arc, + release: Arc, + ) -> Self { + self.gate = Some((reached, release)); + self + } + pub fn with_delay(mut self, delay: Duration) -> Self { self.delay = delay; self @@ -88,56 +99,66 @@ pub fn spawn_podman_stub( let log_for_task = request_log.clone(); let queue_for_task = response_queue; let handle = tokio::spawn(async move { + let mut connections = tokio::task::JoinSet::new(); for _ in 0..expected { let (stream, _) = listener.accept().await.expect("test stub should accept"); let log = log_for_task.clone(); let queue = queue_for_task.clone(); - let result = http1::Builder::new() - .serve_connection( - TokioIo::new(stream), - service_fn(move |req: hyper::Request| { - let log = log.clone(); - let queue = queue.clone(); - async move { - let path = req.uri().path_and_query().map_or_else( - || req.uri().path().to_string(), - |pq| pq.as_str().to_string(), - ); - log.lock() - .expect("request log lock should not be poisoned") - .push(format!("{} {}", req.method(), path)); - let response = queue - .lock() - .expect("response queue lock should not be poisoned") - .pop_front() - .expect("stub response should exist"); - if let Some(expected_members) = &response.archive_members { - assert_eq!(req.method(), hyper::Method::PUT); - let body = req.into_body().collect().await.unwrap().to_bytes(); - let mut archive = tar::Archive::new(body.as_ref()); - let members: Vec<_> = archive - .entries() - .unwrap() - .map(|entry| entry.unwrap().path().unwrap().into_owned()) - .collect(); - assert_eq!(&members, expected_members); + connections.spawn(async move { + let result = http1::Builder::new() + .serve_connection( + TokioIo::new(stream), + service_fn(move |req: hyper::Request| { + let log = log.clone(); + let queue = queue.clone(); + async move { + let path = req.uri().path_and_query().map_or_else( + || req.uri().path().to_string(), + |pq| pq.as_str().to_string(), + ); + log.lock() + .expect("request log lock should not be poisoned") + .push(format!("{} {}", req.method(), path)); + let response = queue + .lock() + .expect("response queue lock should not be poisoned") + .pop_front() + .expect("stub response should exist"); + if let Some(expected_members) = &response.archive_members { + assert_eq!(req.method(), hyper::Method::PUT); + let body = req.into_body().collect().await.unwrap().to_bytes(); + let mut archive = tar::Archive::new(body.as_ref()); + let members: Vec<_> = archive + .entries() + .unwrap() + .map(|entry| entry.unwrap().path().unwrap().into_owned()) + .collect(); + assert_eq!(&members, expected_members); + } + if let Some((reached, release)) = &response.gate { + reached.notify_one(); + release.notified().await; + } + tokio::time::sleep(response.delay).await; + Ok::<_, Infallible>( + hyper::Response::builder() + .status(response.status) + .body(Full::new(response.body)) + .expect("stub response should build"), + ) } - tokio::time::sleep(response.delay).await; - Ok::<_, Infallible>( - hyper::Response::builder() - .status(response.status) - .body(Full::new(response.body)) - .expect("stub response should build"), - ) - } - }), - ) - .await; - // The one-shot test client can close the Unix socket after the - // response, which Hyper reports as a shutdown error. Let the - // request log assertions below decide whether the stub served - // the expected API calls. - let _ = result; + }), + ) + .await; + // The one-shot test client can close the Unix socket after the + // response, which Hyper reports as a shutdown error. Let the + // request log assertions below decide whether the stub served + // the expected API calls. + let _ = result; + }); + } + while let Some(result) = connections.join_next().await { + result.unwrap(); } let _ = std::fs::remove_file(&socket_path_for_task); }); diff --git a/crates/openshell-driver-podman/src/watcher.rs b/crates/openshell-driver-podman/src/watcher.rs index d7364d4f6c..02e0f302eb 100644 --- a/crates/openshell-driver-podman/src/watcher.rs +++ b/crates/openshell-driver-podman/src/watcher.rs @@ -18,7 +18,7 @@ use openshell_core::proto::compute::v1::{ }; use std::collections::HashMap; use std::pin::Pin; -use std::sync::{Arc, Mutex}; +use std::sync::{Arc, Mutex, Weak}; use tokio::sync::mpsc; use tokio_stream::wrappers::ReceiverStream; use tracing::{debug, info, warn}; @@ -34,7 +34,7 @@ use openshell_core::driver_utils::{ pub type WatchStream = Pin> + Send>>; -/// Per-sandbox container exit timestamps that fence state changes from an earlier run. +/// Shared lifecycle coordination and exit timestamps for each sandbox. /// /// Podman can deliver a container's `die` or `stop` event after the stop API /// has returned. If a restart is already in progress, inspecting the container @@ -43,9 +43,32 @@ pub type WatchStream = #[derive(Clone, Debug, Default)] pub struct LifecycleEventFences { previous_finished_at: Arc>>, + operations: Arc>>>>, } impl LifecycleEventFences { + /// Serialize lifecycle mutations and containment for one sandbox. Weak entries + /// avoid retaining deleted sandboxes; never remove an entry while a guard or + /// waiter exists, including when clearing the previous-exit timestamp. + pub async fn lock(&self, sandbox_id: &str) -> tokio::sync::OwnedMutexGuard<()> { + let operation = { + let mut operations = self + .operations + .lock() + .unwrap_or_else(std::sync::PoisonError::into_inner); + operations.retain(|_, operation| operation.strong_count() > 0); + operations + .get(sandbox_id) + .and_then(Weak::upgrade) + .unwrap_or_else(|| { + let operation = Arc::new(tokio::sync::Mutex::new(())); + operations.insert(sandbox_id.to_string(), Arc::downgrade(&operation)); + operation + }) + }; + operation.lock_owned().await + } + pub fn record_previous_exit(&self, sandbox_id: &str, finished_at: Option<&str>) { let mut fences = self .previous_finished_at @@ -148,7 +171,7 @@ pub async fn start_watch( // health check status — matching the same condition derivation used // for live events. if entry.state == "running" { - match inspect_workload(&client, &entry.id).await { + match inspect_workload(&client, &entry.id, &lifecycle_event_fences).await { Ok(inspect) => { if let Some(sandbox) = driver_sandbox_from_inspect(&inspect) { if tx.send(Ok(sandbox_event(sandbox))).await.is_err() { @@ -264,7 +287,7 @@ async fn map_podman_event( .await .ok()?; let workload = workloads.first()?; - return inspect_workload(client, &workload.id) + return inspect_workload(client, &workload.id, lifecycle_event_fences) .await .ok() .and_then(|inspect| driver_sandbox_from_inspect(&inspect)) @@ -275,7 +298,7 @@ async fn map_podman_event( "remove" => Some(deleted_event(sandbox_id.clone())), "create" | "start" | "stop" | "die" | "health_status" => { // Inspect the container to get current state. - match inspect_workload(client, container_id).await { + match inspect_workload(client, container_id, lifecycle_event_fences).await { Ok(inspect) => { if lifecycle_event_fences.matches_previous_exit( event, @@ -358,6 +381,7 @@ async fn map_podman_event( pub async fn inspect_workload( client: &PodmanClient, id: &str, + lifecycle_event_fences: &LifecycleEventFences, ) -> Result { let mut workload = client.inspect_container(id).await?; if workload @@ -371,8 +395,26 @@ pub async fn inspect_workload( let Some(sandbox_id) = workload.config.labels.get(LABEL_SANDBOX_ID) else { return Ok(workload); }; + let sandbox_id = sandbox_id.clone(); + let workload_id = workload.id.clone(); + let _operation = lifecycle_event_fences.lock(&sandbox_id).await; + // The first inspect only identifies the sandbox. Re-read after acquiring + // the guard: a restart may have completed while this inspection waited. + workload = client.inspect_container(id).await?; + if workload.id != workload_id + || workload.config.labels.get(LABEL_SANDBOX_ID) != Some(&sandbox_id) + || workload + .config + .labels + .get(crate::isolation::LABEL_ROLE) + .is_none_or(|role| role != "sandbox") + { + return Err(PodmanApiError::Conflict( + "workload ownership changed during reconciliation".into(), + )); + } let supervisor = client - .inspect_container(&crate::isolation::supervisor_name(sandbox_id)) + .inspect_container(&crate::isolation::supervisor_name(&sandbox_id)) .await; if workload.state.running { match supervisor { @@ -616,6 +658,10 @@ mod tests { let (path, requests, handle) = spawn_podman_stub( "lost-supervisor", vec![ + StubResponse::new( + StatusCode::OK, + r#"{"Id":"workload","Name":"workload","State":{"Status":"running","Running":true},"Config":{"Labels":{"openshell.ai/sandbox-id":"test","openshell.ai/isolation-role":"sandbox"}}}"#, + ), StubResponse::new( StatusCode::OK, r#"{"Id":"workload","Name":"workload","State":{"Status":"running","Running":true},"Config":{"Labels":{"openshell.ai/sandbox-id":"test","openshell.ai/isolation-role":"sandbox"}}}"#, @@ -629,7 +675,51 @@ mod tests { ], ); let client = PodmanClient::new(path.clone()); - let inspected = inspect_workload(&client, "workload").await.unwrap(); + let inspected = inspect_workload(&client, "workload", &LifecycleEventFences::default()) + .await + .unwrap(); + assert!(!inspected.state.running); + handle.await.unwrap(); + assert!( + requests + .lock() + .unwrap() + .iter() + .any(|request| request.ends_with("/libpod/containers/workload/stop?timeout=0")) + ); + let _ = std::fs::remove_file(path); + } + + #[tokio::test] + async fn exited_supervisor_still_stops_workload_during_reconciliation() { + use crate::test_utils::{StubResponse, spawn_podman_stub}; + use hyper::StatusCode; + let (path, requests, handle) = spawn_podman_stub( + "exited-supervisor", + vec![ + StubResponse::new( + StatusCode::OK, + r#"{"Id":"workload","Name":"workload","State":{"Status":"running","Running":true},"Config":{"Labels":{"openshell.ai/sandbox-id":"test","openshell.ai/isolation-role":"sandbox"}}}"#, + ), + StubResponse::new( + StatusCode::OK, + r#"{"Id":"workload","Name":"workload","State":{"Status":"running","Running":true},"Config":{"Labels":{"openshell.ai/sandbox-id":"test","openshell.ai/isolation-role":"sandbox"}}}"#, + ), + StubResponse::new( + StatusCode::OK, + r#"{"Id":"supervisor","Name":"supervisor","State":{"Status":"exited","Running":false},"Config":{}}"#, + ), + StubResponse::new(StatusCode::NO_CONTENT, ""), + StubResponse::new( + StatusCode::OK, + r#"{"Id":"workload","Name":"workload","State":{"Status":"exited","Running":false},"Config":{}}"#, + ), + ], + ); + let client = PodmanClient::new(path.clone()); + let inspected = inspect_workload(&client, "workload", &LifecycleEventFences::default()) + .await + .unwrap(); assert!(!inspected.state.running); handle.await.unwrap(); assert!( @@ -653,6 +743,10 @@ mod tests { StatusCode::OK, r#"{"Id":"workload","Name":"workload","State":{"Status":"running","Running":true},"Config":{"Labels":{"openshell.ai/sandbox-id":"test","openshell.ai/isolation-role":"sandbox"}}}"#, ), + StubResponse::new( + StatusCode::OK, + r#"{"Id":"workload","Name":"workload","State":{"Status":"running","Running":true},"Config":{"Labels":{"openshell.ai/sandbox-id":"test","openshell.ai/isolation-role":"sandbox"}}}"#, + ), StubResponse::new( StatusCode::OK, r#"{"Id":"supervisor","Name":"supervisor","State":{"Status":"configured","Running":false},"Config":{}}"#, @@ -660,11 +754,13 @@ mod tests { ], ); let client = PodmanClient::new(path.clone()); - let inspected = inspect_workload(&client, "workload").await.unwrap(); + let inspected = inspect_workload(&client, "workload", &LifecycleEventFences::default()) + .await + .unwrap(); assert!(inspected.state.running); assert_eq!(inspected.state.health.unwrap().status, "starting"); handle.await.unwrap(); - assert_eq!(requests.lock().unwrap().len(), 2); + assert_eq!(requests.lock().unwrap().len(), 3); let _ = std::fs::remove_file(path); } @@ -680,6 +776,90 @@ mod tests { } } + #[tokio::test] + async fn clearing_exit_fences_preserves_active_operations_across_clones() { + let fences = LifecycleEventFences::default(); + let operation = fences.lock("sandbox-1").await; + fences.record_previous_exit("sandbox-1", Some("previous-exit")); + fences.remove("sandbox-1"); + let clone = fences.clone(); + let mut waiting = Box::pin(clone.lock("sandbox-1")); + assert!( + tokio::time::timeout(std::time::Duration::from_millis(10), &mut waiting) + .await + .is_err() + ); + // One sandbox's lifecycle must not suspend containment for another. + let other = + tokio::time::timeout(std::time::Duration::from_secs(1), clone.lock("sandbox-2")) + .await + .unwrap(); + drop(other); + drop(operation); + let resumed = tokio::time::timeout(std::time::Duration::from_secs(1), waiting) + .await + .unwrap(); + drop(resumed); + } + + #[tokio::test] + async fn delayed_workload_and_supervisor_events_reinspect_after_restart() { + use crate::test_utils::{StubResponse, spawn_podman_stub}; + use hyper::StatusCode; + for supervisor_event in [false, true] { + let stale = r#"{"Id":"container-1","Name":"sandbox","State":{"Status":"exited","Running":false,"FinishedAt":"2026-08-12T16:39:13Z"},"Config":{"Labels":{"openshell.ai/sandbox-id":"sandbox-1","openshell.ai/sandbox-name":"sandbox","openshell.ai/sandbox-workspace":"default","openshell.ai/isolation-role":"sandbox"}}}"#; + let running = r#"{"Id":"container-1","Name":"sandbox","State":{"Status":"running","Running":true},"Config":{"Labels":{"openshell.ai/sandbox-id":"sandbox-1","openshell.ai/sandbox-name":"sandbox","openshell.ai/sandbox-workspace":"default","openshell.ai/isolation-role":"sandbox"}}}"#; + let mut responses = Vec::new(); + if supervisor_event { + responses.push(StubResponse::new( + StatusCode::OK, + r#"[{"Id":"container-1","State":"running"}]"#, + )); + } + responses.extend([ + StubResponse::new(StatusCode::OK, stale), + StubResponse::new(StatusCode::OK, running), + StubResponse::new(StatusCode::OK, r#"{"Id":"supervisor","Name":"supervisor","State":{"Status":"running","Running":true,"Health":{"Status":"healthy"}},"Config":{}}"#), + ]); + let (socket, requests, handle) = spawn_podman_stub("delayed-event", responses); + let client = PodmanClient::new(socket); + let fences = LifecycleEventFences::default(); + fences.record_previous_exit("sandbox-1", Some("2026-08-12T16:39:13Z")); + let operation = fences.lock("sandbox-1").await; + let mut event = podman_event("die", "sandbox-1", 199); + if supervisor_event { + event + .actor + .attributes + .insert(crate::isolation::LABEL_ROLE.into(), "supervisor".into()); + } + let mut mapped = Box::pin(map_podman_event(&event, &client, &fences)); + assert!( + tokio::time::timeout(std::time::Duration::from_millis(20), &mut mapped) + .await + .is_err() + ); + drop(operation); + let result = tokio::time::timeout(std::time::Duration::from_secs(5), mapped) + .await + .unwrap() + .unwrap(); + let Some(watch_sandboxes_event::Payload::Sandbox(snapshot)) = result.payload else { + panic!("expected sandbox event") + }; + let snapshot = snapshot.sandbox.unwrap(); + assert_eq!(snapshot.status.unwrap().conditions[0].status, "True"); + handle.await.unwrap(); + assert!( + requests + .lock() + .unwrap() + .iter() + .all(|r| r.starts_with("GET ")) + ); + } + } + #[test] fn lifecycle_fence_rejects_delayed_stop_events_from_before_restart() { let fences = LifecycleEventFences::default(); diff --git a/docs/how-it-works/sandboxes/runtimes.mdx b/docs/how-it-works/sandboxes/runtimes.mdx index 505f2cc69b..9f9942a8e6 100644 --- a/docs/how-it-works/sandboxes/runtimes.mdx +++ b/docs/how-it-works/sandboxes/runtimes.mdx @@ -178,6 +178,11 @@ On macOS, set `host_gateway_ip` only if your Podman machine uses a non-standard For networks that require a corporate proxy, set `https_proxy`, `no_proxy`, and related `proxy_*` options. See the [Gateway Configuration File](/how-it-works/gateways/configuration) reference. +Podman starts the workload before its companion supervisor to establish their +shared user namespace. During restart, health reconciliation waits for both +container starts to finish. After restart completes or fails, a running workload +with a missing or exited supervisor is stopped. + ### Podman Mounts Podman supports the same `volume`, `tmpfs`, and `bind` mounts as [Docker](#docker-mounts), using the `podman` key in driver config, and the same bind-mount warning applies. Podman `volume` mounts do not support `subpath`. diff --git a/skills/debug-openshell-cluster/SKILL.md b/skills/debug-openshell-cluster/SKILL.md index ab774e3a88..4d24c49e4e 100644 --- a/skills/debug-openshell-cluster/SKILL.md +++ b/skills/debug-openshell-cluster/SKILL.md @@ -369,6 +369,12 @@ Common findings: - Sandbox image missing or pull denied: verify image reference and registry credentials. - Sandbox fails before readiness with an identity-resolution error: inspect the image's OCI `USER` and matching `/etc/passwd` and `/etc/group` entries, or explicitly set both process identity fields in policy. Numeric workload identities `1` through `4294967294` are accepted; root, the invalid identity sentinel, and missing identities are rejected. - Supervisor cannot connect: check its gateway endpoint and gateway logs. +- During Podman restart, the workload starts before the supervisor to establish + their shared user namespace. Reconciliation waits for the lifecycle operation + to finish before treating an exited supervisor as lost. If restart fails or + is cancelled, a running workload with a missing or exited supervisor is still + stopped. Capture both container states and start/stop events when diagnosing + a restart failure; one intermediate inspect does not establish supervisor loss. - Inspect both Podman containers for the sandbox: the `sandbox` isolation role must have network mode `none`; the `supervisor` role owns the gateway session and egress. Both run non-root with all capabilities dropped. Check the private