Quikcast Help

Explicit server drain lifecycle review gate

The small drain milestone implements authenticated, one-way admission shutdown while preserving already-owned streaming work. It does not claim standalone feature completeness or production qualification.

1. Lifecycle design

The lifecycle is accepting → draining → shutting_down/terminated. Lifecycle is a concrete shared runtime component used by continuous mounts and HLS producer HTTP admission. A OnceLock<DrainStarted> atomically publishes the one-way transition and its immutable wall-clock start/monotonic clock together. Duplicate callers cannot replace the value or restore acceptance. Fast admission checks require no global mutex, network operation, task spawn, or await. The first transition produces a structured log.

Drain state is independent of the existing shutdown cancellation token. It never cancels source generations, listener tokens, public HTTP bodies, or HLS storage. No background drain task, timer, automatic exit, source replacement, or resume endpoint is introduced. Per-mount lifecycle visibility becomes draining; source/listener connection state remains unchanged until its own normal termination.

2. Endpoint contract

POST /api/server/drain runs on the existing administration listener and requires the existing Bearer credential. It has no parameters or request body. A successful first or duplicate request returns 200 with the same ServerInfo model as GET /api/server. Errors retain the native envelope: 401 unauthorized, 400 invalid_request, or 405 method_not_allowed with Allow: POST.

GET /api/server reports accepting, draining, or shutting_down. New drain_started_at and drain_age_ms fields are null before explicit drain, then report original Unix epoch milliseconds and monotonic elapsed milliseconds. Duplicate calls preserve the start time. Shutdown state takes precedence, so a delayed drain cannot restore accepting state. The complete curl contract is in management-api.md.

Management inspection, stats, mount/source/listener/HLS inspection, and administrative source/listener disconnect remain usable during drain. Existing source metadata updates remain usable. There is no API version change or generic action endpoint.

3. Continuous admission semantics and races

New continuous source/listener admissions fail with typed Error::ServerDraining, HTTP 503, and public body Server draining\n. Rejection is distinct from capacity exhaustion and listener limits.

Admission checks occur early, before generation/listener scarce resources are acquired, and again at the final domain commit boundary. For sources, the final check is immediately before publishing the generation under the mount admission lock. For listeners, it is immediately before committing counters and the record under the per-mount registry lock. The successful final atomic accepting observation is the admission linearization point.

If drain is authoritative before that observation, all provisional generation references, local/global semaphore permits, and record allocations are dropped normally. No generation is published, no listener record inserted, no total admission counter incremented, and no generation sequence consumed. If the accepting observation wins the race, the operation may complete its commit, response, and source startup after the drain response. A connection accepted by TCP or a partially read request is not yet admitted. An already-claimed source can start after drain and continue normally.

This contract needs no giant global admission mutex. Existing mount-local locking and owned permits retain their roles. Ordinary source ingest/body delivery never checks drain; its normal lease, source-generation cancellation, timeout, administrative disconnect, and shutdown paths remain authoritative. Continuous HEAD requests inspect headers without listener admission and remain available.

Transport connections and bounded request handshakes continue to be accepted because management, health, metrics, and HLS playback must remain available. Their existing global capacity/timeout ceilings remain in force. Drain therefore does not free occupied listener slots or reserve a separate administration capacity.

4. HLS producer behavior: option A

Every new non-GET producer request checks drain immediately after producer authentication and before upload permits, allocation charges, body reads, or mutations. All new producer mutations are blocked uniformly: stream/rendition creation, segment upload, playlist publication, master publication, ending, and deleting. They receive 503 Server draining\n through the existing producer HTTP error model. Producer GET inspection remains available.

The HLS admission boundary is that request-level accepting observation. An admitted producer request may finish after drain, including a body upload and its segment acceptance. A producer does not acquire permission for future requests merely by having created a stream earlier. This preserves a coherent option A contract without adding producer session ownership.

Public HLS GET/HEAD/playlists/segment serving and native HLS inspection remain available. HLS maintenance and existing grace/retention rules continue. An in-flight accepted upload can create a retained staging segment, but a new request cannot publish it once draining; producers should complete publication before requesting drain if they want it advertised. Quikcast neither rewrites nor extends an existing playlist to compensate. Published objects remain immutable.

5. Shutdown interaction

SIGTERM, Ctrl+C, or the existing shutdown token triggers the unchanged supervised shutdown path: close listening descriptors, cancel owned connections and source generations, join bounded tasks, apply the configured deadline/abort fallback, and close HLS state after connection ownership is reaped.

That path works from either accepting or draining and need not reinitialize drain. Direct shutdown from accepting is still valid; drain_started_at stays null if no explicit drain occurred. Starting drain alone never schedules or initiates shutdown. Operators choose when to signal shutdown, and existing traffic may continue indefinitely until then.

6. Health and observability

Liveness stays HTTP 200 during drain. Readiness becomes HTTP 503 with Server draining\n; the existing ready Prometheus gauge becomes 0. Neither endpoint scans streaming state or requires active sources/listeners/HLS. Readiness stays unavailable during process shutdown.

A fixed drain_rejections_total counter distinguishes new workload rejected because of drain, and /api/stats exposes drain_rejections. The existing admission rejection/error counters also account for those rejected source/listener/HLS requests. Health/readiness/inspection polls do not count as drain rejections. No dynamic labels or separate metrics framework are added. /api/server supplies the explicit state/start/age instead of an additional drain gauge.

7. Tests and validation

Eight new tests extend the existing 61-test baseline:

  • One-way publication and eight barrier-synchronized concurrent drain requests, preserving the original timestamp.

  • Authenticated endpoint transition, missing auth, wrong method, invalid query/body, duplicate POST, stable drain timestamp, inspection during drain, and shutdown state precedence.

  • A source admitted before transition can start and publish afterward; post-drain source/listener attempts fail without consuming permits; source administration still works.

  • Source admission paused after generation reservation and before commit: drain wins, the source rejects, reservation returns, no generation is active, and generation sequence/counters remain unchanged.

  • Listener admission paused after global/local reservations and before commit: drain wins, both permits return, only the existing listener remains, and existing source/listener cancellation stays untouched. Administrative disconnect then uses normal cleanup.

  • Real-socket streaming survives drain with exact-byte delivery to two existing listeners. TCP connections with incomplete headers before drain reject when admission completes afterward. Raw/framed source and listener rejection use 503; management/detail/stat inspection and administrative controls remain usable. Readiness changes, liveness remains healthy, rejection counters exclude health probes.

  • Shutdown while already draining closes and joins active source/listener transports, retains shutdown accounting, and leaves no active leases/audio allocation.

  • HLS option A allows an observed admitted partial upload to finish after drain; all nine new producer mutation shapes reject; producer GET, all four native HLS inspection routes, playlist playback, and segment reads continue with unchanged revision/state. Upload counts return to zero and shutdown releases retained HLS state.

Concurrency tests use explicit barriers at test-only commit checkpoints. Socket tests observe actual upload/lease metrics and exact media delivery with bounded deadlines and cooperative yields; arbitrary sleeps do not choose the race result. Commit checkpoints are absent from production builds. Existing generation safety, codecs/ICY, media corruption/isolation, immutable HLS publication, slow-consumer, property, and memory-ceiling checks remain intact.

Final validation: 69 tests passed (49 unit/property/concurrency tests and 20 streaming/HTTP integration tests), zero failures. cargo clippy --all-targets -- -D warnings, cargo fmt --check, and git diff --check passed. Test execution used localhost socket access. Manual encoder/player/proxy sessions were not repeated.

8. Remaining standalone feature gaps

Required before a feature-completeness decision

Confirm operator requirements against the currently implemented feature set: static configuration/mounts, restart-based secret/limit updates, external TLS termination, available encoder/listener formats, and HLS option A during drain. A focused manual operator walkthrough remains useful evidence: authenticate, inspect, drain, observe readiness removal, continue playback, disconnect intentionally, and signal shutdown. Drain is now implemented and should no longer appear as an unimplemented core operational primitive.

No additional feature is automatically declared required without the agreed standalone scope. Drain does not constitute production qualification.

Useful but optional

Configuration reload/credential rotation, build provenance, a small operator CLI, deployment examples for private administration/external TLS, or additional controls justified by a real workflow. Resume from drain, automatic drain deadlines, and session-aware HLS option B would require an explicit operational need and a separate lifecycle design. Native TLS, PROXY protocol, additional codecs, and HLS containers remain requirement-driven options.

Intentionally deferred

Relays, fallback mounts, takeover/replacement/priority, failover, mount chaining/master-slave behavior, integration, broad optimization, and production qualification. No work on these began. A future continuous relay must remain continuous delivery and must not automatically produce HLS.

Not appropriate for Quikcast

RadioPlatform-specific IDs/authentication/configuration/persistence, account/session/RBAC systems, speculative Icecast administration/configuration parity, transcoding/remuxing continuous streams into HLS, and external media process orchestration.

The next decision remains a focused operator workflow review and explicit selection of any remaining standalone feature. Stop here; do not automatically proceed into optimization, relays, integration, or production qualification.

08 October 2026