- open_hw_accel marks the device unavailable when the decoder OPEN fails
(cuvidCreateDecoder OOM at 4K) too, not just device-context creation:
without it every subsequent decoder session retried CUDA and flooded
the log per open.
- Decoder gains hardware_decoding(); the oak-render decode-session LRU
evicts hardware sessions first (each pins a GPU surface pool — ~100 MB
at 4K), so a full cache cannot exhaust video memory before the next
open.
- Worker pool count now factors GPU vram: per-worker budget = 1 GiB
(1080p peak) scaled by pixel ratio + 256 MiB idle floor, 10% reserve
of free vram; applied when hardware decoding is on (nvidia-smi query,
None otherwise falls back to the RAM/CPU policy).
- Dynamic pool resize: ProcessDispatcher::set_target_workers grows or
retires workers; retiring ones stop claiming, drain their in-flight
batch (future playback frames included), then exit naturally on the
shutdown signal — no mid-work kill (30 s deadline only as a hung-
decoder last resort). A retiring worker that dies re-queues its frames
to surviving workers. Resizes are throttled to 2 s (a resolution burst
merges; only the latest target applies) so 1080p<->4K flaps cannot
thrash process spawns.
- RenderManager::set_workspace_size announces the sequence resolution;
RealEngine calls it from refresh_sequence_info.
- Integration test: shrink 3->1 mid-wave (all frames complete, retired
workers exit naturally) then regrow 1->3 and render a fresh wave.