- Autocache range jobs now post at Background priority
(submit_video_background): they used to go through the Seek path and,
after the M4 seek over-admission, jumped ahead of playback and past the
render-queue bound. The interactive single-frame preview keeps Seek.
- Job.cancelled: the arena installs the slot's cancel atom, and
execute_job finishes a cancelled job with Error::State before running
the producer — a cancel no longer burns a full render/GPU pass only to
discard the result. Exactly-once delivery is unchanged.
- DECODE_LRU_CAP 8 -> 2: the decode service's LRU is a hand-off buffer,
not the cache of record (the eval-side decoded_frames LRU is); the
double-cache footprint at 1080p F32 drops by ~6 frames. A hand-off miss
is served from the eval cache without a new decode.
- Tests: sequence-aware preview cancel, over-admitted seek ordering,
deterministic prefetch LRU reuse, cancelled-job skip, autocache
priority. docs §3.4 backfilled with the A/B/C audit outcomes.
docs/zh/plans/render-pipeline-threads.md M4: the thread pipeline now
keeps its decode thread ahead of the render thread and the app's
playback window consumes in-process frames.
- Render queue: priority-ordered by JobSchedule.priority (Seek >
Playback > Background, FIFO within a class), so interactive frames
jump playback exports/autocache. Seek posts may over-admit the bound:
priority only reorders queued jobs, so a full queue of background work
must not park the UI thread until an export frame finishes.
- Decode queue: rendezvous Requests are served ahead of queued
Prefetches (a frame the renderer needs never waits behind speculative
decodes); Sync barriers stay FIFO. The queue is a bounded
Mutex+Condvar structure, preserving the request backpressure and the
wait_idle contract.
- Playback read-ahead: a Playback job's footage decode requests are
derived from its montage/footage spec on post (same media time, size
and force_format.unwrap_or(F32) as the eval) and queued immediately,
so frame N+1 decodes while frame N runs its GPU passes.
- App window: PreviewWindow slots are generalized to
PreviewSlot::{Shm, Video}; the pipeline's in-process TicketPayload is
cached and consumed by cpu_frame exactly like a worker slot.
PipelineBackend::preview_window_capacity reports the render-queue
headroom, so playback posts are capped to what the queue can take;
cancel_preview_frame drops queued frames the playhead has passed,
matched on the full (sequence, frame, version) key so one monitor's
window never drops the other sequence's same-numbered frame.
- Tests: decode-queue preemption/FIFO, render-queue ordering, request
derivation, and deterministic end-to-end M4 tests: a prefetch that
must be reused by the render request (LRU hit, single decode — the
read-ahead claim is falsifiable), a parked-render-thread priority test
where a full queue of background work still lets a Seek over-admit and
run first, and a sequence-aware cancel test. The playback prefetch
smoke asserts prefetches == distinct decodes == frames; it does not
claim zero heap copies (Frame.data is deep-copied at the eval-cache
and service-LRU boundaries today).
- bench_playback gains a pipeline mode with CPU (self+children) and
first-frame latency; both backends now produce F32 frames so the
comparison is like-for-like. The §3.4 backfill records the numbers:
at the proxy size the pipeline is faster with a lower first frame; at
1080p peak throughput is below the multi-worker pool, but that is an
artifact of the decode still being CPU software (M5), not a case for
pooling decode threads — GPU decode is a single device/queue and the
zero-copy import shares one GPU memory pool, so the single decode
thread stays the target shape.
docs/zh/plans/render-pipeline-threads.md M2: the graph's textures stay
on the GPU from evaluation through presentation, and presentation runs
on the UI's own wgpu device.
- wgpu 25 -> 29 (naga 29) across the engine, unifying it with
gpui_wgpu so engine textures are directly sampleable by the presenter
(a single wgpu remains in the lockfile).
- GpuContext::adopt/install_shared: the app registers the window's
device at startup and the render thread renders on it;
texture_handle hands the raw Arc<wgpu::Texture> to
SurfaceSource::Texture - zero-copy present on Linux/FreeBSD. The
shared slot replaces an engine context that has not touched the GPU
yet (startup-order guard) and refuses once it has.
- Texture::Gpu shares a GpuLease so clones release the registry token
exactly once; the compositor, transitions and adjustment sweeps keep
GPU textures end to end (no per-clip readbacks; GPU clears for
black/generated frames).
- Color management stays on the GPU: the output node + display ICC
chain is baked into a 65^3 3D LUT with the exact CPU reference and
applied by the present WGSL pass (manual trilinear);
ColorTransformJob bakes its OCIO processor the same way. Neither
path skips color management.
- The explicit readback boundaries accept GPU textures: export
encoder, CLI, worker shm, disk cache; CPU OpenFX already read back.
- M5 dependency: the YUV->RGB GPU pass (BT.601/709/2020 x
limited/full) matches colormath::yuv444p16_to_rgb_f32.
- Acceptance: gpu_transfer_counters; single-clip and layered
(multi-track + transition + adjustment) playback tests assert zero
GPU->CPU readbacks, and the app test asserts adopted-device present
is zero-copy. GPU tests hard-fail when OAK_REQUIRE_GPU is set (CI
lavapipe) instead of skipping silently.
docs/zh/plans/render-pipeline-threads.md M1: an in-process
alternative to the worker-process pool, behind OAK_PIPELINE=threads
(processes stays the default and is fully retained).
- pipeline.rs: PipelineBackend implements JobDispatch over a single
render thread draining a bounded FIFO (cap 8; blocking post with
condvar backpressure and a one-ahead exception for re-posts from
the render thread itself; shutdown drains with Error::State like
the inline dispatcher). The DecodeService is a single decode
thread behind a bounded command queue with a real LRU (tick-based
eviction), rendezvous requests (None on shutdown -> the caller
decodes inline), prefetch gated on render-queue room, and a Sync
barrier; it installs into a process-wide slot that eval's footage
path consults per frame (no service -> the synchronous decode it
always was).
- The manager gains RenderBackendChoice::Pipeline; init() reads
OAK_PIPELINE (threads -> pipeline, anything else -> the process
pool), audio stays deliberately inline.
- Present mapping: the UI thread consumes through the ticket
completion, unchanged — no fourth thread is invented.
- Tests: decode-service unit tests (rendezvous, LRU hit/eviction,
error propagation, backpressure gate) plus a six-case integration
suite matrixed over inline vs pipeline — consecutive-frame and
out-of-order seek pixel equality asserted byte for byte, with
decode counters proving the service (not the caller) did the
codec work.