Files
oak-editor/crates/oak-render/tests
Mike-Solar ba1143e7a3 render: the M4 playback prefetch — dependency window, priorities and backpressure
docs/zh/plans/render-pipeline-threads.md M4: the thread pipeline now
keeps its decode thread ahead of the render thread and the app's
playback window consumes in-process frames.

- Render queue: priority-ordered by JobSchedule.priority (Seek >
  Playback > Background, FIFO within a class), so interactive frames
  jump playback exports/autocache. Seek posts may over-admit the bound:
  priority only reorders queued jobs, so a full queue of background work
  must not park the UI thread until an export frame finishes.
- Decode queue: rendezvous Requests are served ahead of queued
  Prefetches (a frame the renderer needs never waits behind speculative
  decodes); Sync barriers stay FIFO. The queue is a bounded
  Mutex+Condvar structure, preserving the request backpressure and the
  wait_idle contract.
- Playback read-ahead: a Playback job's footage decode requests are
  derived from its montage/footage spec on post (same media time, size
  and force_format.unwrap_or(F32) as the eval) and queued immediately,
  so frame N+1 decodes while frame N runs its GPU passes.
- App window: PreviewWindow slots are generalized to
  PreviewSlot::{Shm, Video}; the pipeline's in-process TicketPayload is
  cached and consumed by cpu_frame exactly like a worker slot.
  PipelineBackend::preview_window_capacity reports the render-queue
  headroom, so playback posts are capped to what the queue can take;
  cancel_preview_frame drops queued frames the playhead has passed,
  matched on the full (sequence, frame, version) key so one monitor's
  window never drops the other sequence's same-numbered frame.
- Tests: decode-queue preemption/FIFO, render-queue ordering, request
  derivation, and deterministic end-to-end M4 tests: a prefetch that
  must be reused by the render request (LRU hit, single decode — the
  read-ahead claim is falsifiable), a parked-render-thread priority test
  where a full queue of background work still lets a Seek over-admit and
  run first, and a sequence-aware cancel test. The playback prefetch
  smoke asserts prefetches == distinct decodes == frames; it does not
  claim zero heap copies (Frame.data is deep-copied at the eval-cache
  and service-LRU boundaries today).
- bench_playback gains a pipeline mode with CPU (self+children) and
  first-frame latency; both backends now produce F32 frames so the
  comparison is like-for-like. The §3.4 backfill records the numbers:
  at the proxy size the pipeline is faster with a lower first frame; at
  1080p peak throughput is below the multi-worker pool, but that is an
  artifact of the decode still being CPU software (M5), not a case for
  pooling decode threads — GPU decode is a single device/queue and the
  zero-copy import shares one GPU memory pool, so the single decode
  thread stays the target shape.
2026-09-15 17:25:01 +08:00
..