ba1143e7a3584a92e201e91c211d4354458103f8
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ba1143e7a3 |
render: the M4 playback prefetch — dependency window, priorities and backpressure
docs/zh/plans/render-pipeline-threads.md M4: the thread pipeline now
keeps its decode thread ahead of the render thread and the app's
playback window consumes in-process frames.
- Render queue: priority-ordered by JobSchedule.priority (Seek >
Playback > Background, FIFO within a class), so interactive frames
jump playback exports/autocache. Seek posts may over-admit the bound:
priority only reorders queued jobs, so a full queue of background work
must not park the UI thread until an export frame finishes.
- Decode queue: rendezvous Requests are served ahead of queued
Prefetches (a frame the renderer needs never waits behind speculative
decodes); Sync barriers stay FIFO. The queue is a bounded
Mutex+Condvar structure, preserving the request backpressure and the
wait_idle contract.
- Playback read-ahead: a Playback job's footage decode requests are
derived from its montage/footage spec on post (same media time, size
and force_format.unwrap_or(F32) as the eval) and queued immediately,
so frame N+1 decodes while frame N runs its GPU passes.
- App window: PreviewWindow slots are generalized to
PreviewSlot::{Shm, Video}; the pipeline's in-process TicketPayload is
cached and consumed by cpu_frame exactly like a worker slot.
PipelineBackend::preview_window_capacity reports the render-queue
headroom, so playback posts are capped to what the queue can take;
cancel_preview_frame drops queued frames the playhead has passed,
matched on the full (sequence, frame, version) key so one monitor's
window never drops the other sequence's same-numbered frame.
- Tests: decode-queue preemption/FIFO, render-queue ordering, request
derivation, and deterministic end-to-end M4 tests: a prefetch that
must be reused by the render request (LRU hit, single decode — the
read-ahead claim is falsifiable), a parked-render-thread priority test
where a full queue of background work still lets a Seek over-admit and
run first, and a sequence-aware cancel test. The playback prefetch
smoke asserts prefetches == distinct decodes == frames; it does not
claim zero heap copies (Frame.data is deep-copied at the eval-cache
and service-LRU boundaries today).
- bench_playback gains a pipeline mode with CPU (self+children) and
first-frame latency; both backends now produce F32 frames so the
comparison is like-for-like. The §3.4 backfill records the numbers:
at the proxy size the pipeline is faster with a lower first frame; at
1080p peak throughput is below the multi-worker pool, but that is an
artifact of the decode still being CPU software (M5), not a case for
pooling decode threads — GPU decode is a single device/queue and the
zero-copy import shares one GPU memory pool, so the single decode
thread stays the target shape.
|
||
|
|
fec6e9dba7 |
render: the M3 OFX host — one oak-worker --ofx-host process for every plugin job
docs/zh/plans/render-pipeline-threads.md M3 (design 3.2): OpenFX crash isolation moves from "every worker hosts plugins" to a single dedicated host process, served over NDJSON + shared memory. - oak-worker --ofx-host mode (src/ofx_host.rs): loads every plugin once, resolves jobs by the cross-process-stable OFX identifier, and renders through the same in-process executor the workers used to install. - oak-render/ofxhost.rs: the single-host client. The render manager creates and installs it for the Pipeline backend (lazy spawn on the first plugin job); eval::process_plugin_job prefers it and falls back to the in-process executor otherwise, so the process backend keeps its current behavior until M4. - Data plane: input/output FrameSlotPool pairs (the handshake's input_* fields are used for the first time). Named clips and the source frame are written to input slots after the explicit CPU readback; the plugin output returns through an output slot. Pool size/capacity grow by a host restart when a job needs more (safe: submissions are serialized and one job is in flight). - Crash loop: reader EOF fails the in-flight submit, which respawns the host and re-posts the same job (frames are read back once); after three consecutive crashes the client is permanently dead and the evaluator falls back to a purple frame. The dead child is reaped immediately, and a submit mutex enforces the one-job-in-flight contract. - Progress/cancel: the host flushes plugin_progress immediately (live progress), and reads stdin on its own thread so plugin_cancel takes effect mid-render at the plugin's next progressUpdate; the sticky flag resets at progressStart and request_plugin_cancel_all broadcasts to both the worker pool and the host. - JobSpec::Plugin / PluginJobPayload carry the plugin type_id (stable across processes); `--ofx-crash-once` / `--ofx-crash-always` are the deterministic crash hooks, matching the worker's env hooks. - Tests: wire round-trips; host unit tests (crash budget, cancel-flag reset through the factory, source mapping); oak-worker integration tests against the real host + bundled test plugin (render + progress, crash respawn and re-post, three-crash give-up, mid-render cancel on the new slow variant, concurrent submits); eval's purple fallback. |
||
|
|
48e99e56b7 |
render: the M2 GPU zero-copy pipeline — wgpu 29, shared gpui device, GPU color LUTs
docs/zh/plans/render-pipeline-threads.md M2: the graph's textures stay on the GPU from evaluation through presentation, and presentation runs on the UI's own wgpu device. - wgpu 25 -> 29 (naga 29) across the engine, unifying it with gpui_wgpu so engine textures are directly sampleable by the presenter (a single wgpu remains in the lockfile). - GpuContext::adopt/install_shared: the app registers the window's device at startup and the render thread renders on it; texture_handle hands the raw Arc<wgpu::Texture> to SurfaceSource::Texture - zero-copy present on Linux/FreeBSD. The shared slot replaces an engine context that has not touched the GPU yet (startup-order guard) and refuses once it has. - Texture::Gpu shares a GpuLease so clones release the registry token exactly once; the compositor, transitions and adjustment sweeps keep GPU textures end to end (no per-clip readbacks; GPU clears for black/generated frames). - Color management stays on the GPU: the output node + display ICC chain is baked into a 65^3 3D LUT with the exact CPU reference and applied by the present WGSL pass (manual trilinear); ColorTransformJob bakes its OCIO processor the same way. Neither path skips color management. - The explicit readback boundaries accept GPU textures: export encoder, CLI, worker shm, disk cache; CPU OpenFX already read back. - M5 dependency: the YUV->RGB GPU pass (BT.601/709/2020 x limited/full) matches colormath::yuv444p16_to_rgb_f32. - Acceptance: gpu_transfer_counters; single-clip and layered (multi-track + transition + adjustment) playback tests assert zero GPU->CPU readbacks, and the app test asserts adopted-device present is zero-copy. GPU tests hard-fail when OAK_REQUIRE_GPU is set (CI lavapipe) instead of skipping silently. |
||
|
|
29204d1f63 |
node: virtual graph endpoints and the Kahn-order BFS sweep (M0b core)
Per docs/zh/plans/render-pipeline-threads.md §3.8: - oak-node/nodes/graphendpoints.rs: the GraphInput/GraphOutput virtual node pair — factory-registered but hidden from every create menu, duplicate refused, real value() semantics (the input forwards its feed_in row, the output publishes its tex_in as the frame). The input endpoint also declares a connectable feed_in port (documented deviation: footage/generator sources have no connectable inputs, so the walk needs a feeder anchor). - graph.rs: ensure_endpoints/endpoints/is_endpoint — idempotent, identified by type id, default input->output edge only while the output's tex_in is free; remove_node refuses endpoints. - project.rs + serializer.rs: every project graph carries the pair; a legacy file without endpoints migrates on load (roundtrip and legacy-migration tests, re-save is idempotent). - traverser.rs: eval_graph_bfs — the endpoint-to-endpoint Kahn sweep. Live set = (input's forward cone U its feeder cone) INTERSECT (output's backward cone); multi-input nodes dequeue at zero in-degree over the live subgraph; deterministic ascending-id ready order (Graph::edges is a BTreeSet, so insertion order is unrecoverable — documented); time-shifted upstreams pull through the shared DFS memo (walk_dfs, factored out of evaluate); un-orderable remainder reports a named cycle; missing endpoints / unreachable output are errors. Eight BFS tests cover the plan's acceptance bullets. - oak-render: bfs_endpoint_sweep_renders_footage_through_position — real clip through a real Position node via the sweep, shifted pixels asserted against a reference decode. - Endpoint names localized in all eight i18n packs; storage/structure tests updated for the two extra nodes. |
||
|
|
3c05f7107b |
docs: harden the pipeline plan — GPU decode zero-copy, Job-graph BFS
Two user-mandated amendments: - Decode must be GPU wherever possible and share the render GPU's memory: hardware surfaces (NV12/P010) are imported as GPU textures via the platform interop paths (DMA-BUF / DXGI / IOSurface / CUDA-Vulkan), av_hwframe_transfer_data is never executed on the hw path, and CPU decode + staging upload demotes to fallback only. FFmpeg hwaccel first (the hwdecode.rs device model already builds the device contexts; upstream Olive has no hw decode at all, so the reference for this part is FFmpeg + the existing crate), hand-written GPU decode strictly second. YUV->RGB becomes a built-in GPU pass replacing CPU swscale. Milestone M5 becomes the GPU-decode zero-copy track with HW_TRANSFERS zero as its acceptance counter. - The Job graph becomes a real adjacency structure (no linear table, no 2D array, possibly not fully connected) with a fixed pair of virtual GraphInput/GraphOutput nodes per graph: connected by default, undeletable, un-duplicable, shown in the node editor. resolve is a Kahn-style BFS from the input node — multi-input joins wait for every input, multi-output fans out, the order is deterministic and graph-explicit, cycles error out, unreachable nodes never run — until every branch converges at the output node. M0 splits into M0a (Job enum + single-loop match) and M0b (virtual endpoints + BFS + node-editor display). |
||
|
|
87ed45ffee |
docs: plan the decode/render thread pipeline with GPU zero-copy
Task book for the render-pipeline rearchitecture: one decode thread and one render thread feeding queue-linked stages with the main process presenting (the GPU's single queue makes the multi-process backend dead weight), one dedicated OpenFX host process with bounded respawn, GPU-resident frames end-to-end except at the CPU OFX/export/cache boundaries, a per-backend interop table, and the resolve rewrite to a single-loop match over a completed Job enum (CacheJob included) following upstream Olive's NodeTraverser::ResolveJobs. Milestones M0-M5 with the thread backend kept behind an OAK_PIPELINE fallback switch. |