Two user-mandated amendments:
- Decode must be GPU wherever possible and share the render GPU's
memory: hardware surfaces (NV12/P010) are imported as GPU textures
via the platform interop paths (DMA-BUF / DXGI / IOSurface /
CUDA-Vulkan), av_hwframe_transfer_data is never executed on the hw
path, and CPU decode + staging upload demotes to fallback only.
FFmpeg hwaccel first (the hwdecode.rs device model already builds
the device contexts; upstream Olive has no hw decode at all, so the
reference for this part is FFmpeg + the existing crate), hand-written
GPU decode strictly second. YUV->RGB becomes a built-in GPU pass
replacing CPU swscale. Milestone M5 becomes the GPU-decode
zero-copy track with HW_TRANSFERS zero as its acceptance counter.
- The Job graph becomes a real adjacency structure (no linear table,
no 2D array, possibly not fully connected) with a fixed pair of
virtual GraphInput/GraphOutput nodes per graph: connected by
default, undeletable, un-duplicable, shown in the node editor.
resolve is a Kahn-style BFS from the input node — multi-input joins
wait for every input, multi-output fans out, the order is
deterministic and graph-explicit, cycles error out, unreachable
nodes never run — until every branch converges at the output node.
M0 splits into M0a (Job enum + single-loop match) and M0b (virtual
endpoints + BFS + node-editor display).