render: the M4 playback prefetch — dependency window, priorities and backpressure
docs/zh/plans/render-pipeline-threads.md M4: the thread pipeline now
keeps its decode thread ahead of the render thread and the app's
playback window consumes in-process frames.
- Render queue: priority-ordered by JobSchedule.priority (Seek >
Playback > Background, FIFO within a class), so interactive frames
jump playback exports/autocache. Seek posts may over-admit the bound:
priority only reorders queued jobs, so a full queue of background work
must not park the UI thread until an export frame finishes.
- Decode queue: rendezvous Requests are served ahead of queued
Prefetches (a frame the renderer needs never waits behind speculative
decodes); Sync barriers stay FIFO. The queue is a bounded
Mutex+Condvar structure, preserving the request backpressure and the
wait_idle contract.
- Playback read-ahead: a Playback job's footage decode requests are
derived from its montage/footage spec on post (same media time, size
and force_format.unwrap_or(F32) as the eval) and queued immediately,
so frame N+1 decodes while frame N runs its GPU passes.
- App window: PreviewWindow slots are generalized to
PreviewSlot::{Shm, Video}; the pipeline's in-process TicketPayload is
cached and consumed by cpu_frame exactly like a worker slot.
PipelineBackend::preview_window_capacity reports the render-queue
headroom, so playback posts are capped to what the queue can take;
cancel_preview_frame drops queued frames the playhead has passed,
matched on the full (sequence, frame, version) key so one monitor's
window never drops the other sequence's same-numbered frame.
- Tests: decode-queue preemption/FIFO, render-queue ordering, request
derivation, and deterministic end-to-end M4 tests: a prefetch that
must be reused by the render request (LRU hit, single decode — the
read-ahead claim is falsifiable), a parked-render-thread priority test
where a full queue of background work still lets a Seek over-admit and
run first, and a sequence-aware cancel test. The playback prefetch
smoke asserts prefetches == distinct decodes == frames; it does not
claim zero heap copies (Frame.data is deep-copied at the eval-cache
and service-LRU boundaries today).
- bench_playback gains a pipeline mode with CPU (self+children) and
first-frame latency; both backends now produce F32 frames so the
comparison is like-for-like. The §3.4 backfill records the numbers:
at the proxy size the pipeline is faster with a lower first frame; at
1080p peak throughput is below the multi-worker pool, but that is an
artifact of the decode still being CPU software (M5), not a case for
pooling decode threads — GPU decode is a single device/queue and the
zero-copy import shares one GPU memory pool, so the single decode
thread stays the target shape.
This commit is contained in:
@@ -15,32 +15,41 @@
|
||||
// along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
|
||||
//! Real-footage playback benchmark: renders `N` sequential frames of a
|
||||
//! real media file through the real oak-worker pool at the app's preview
|
||||
//! proxy size, mimicking the playback pre-render window (Playback
|
||||
//! priority, interleaved claiming, immediate slot release). Reports
|
||||
//! throughput and completion latency so worker-side decode/render
|
||||
//! hotspots can be measured end to end.
|
||||
//! real media file through the real oak-worker pool **or** the M1/M4
|
||||
//! thread pipeline at the app's preview proxy size, mimicking the
|
||||
//! playback pre-render window (Playback priority, immediate release).
|
||||
//! Reports throughput, completion latency, first-frame latency and CPU
|
||||
//! time (self + children) so the M4 acceptance comparison between the two
|
||||
//! backends is reproducible.
|
||||
//!
|
||||
//! Run from the repo root:
|
||||
//!
|
||||
//! ```sh
|
||||
//! cargo run --release -p oakrender --example bench_playback -- <media> [frames] [workers] [long_edge]
|
||||
//! cargo run --release -p oakrender --example bench_playback -- <media> [frames] [workers] [long_edge] [processes|pipeline]
|
||||
//! ```
|
||||
//!
|
||||
//! `frames` defaults to 240, `workers` to the adaptive policy, and
|
||||
//! `long_edge` to 480 (the app's preview proxy size). To profile a
|
||||
//! worker while this runs: `pgrep oak-worker | head -1 | xargs sample 10`.
|
||||
//! `frames` defaults to 240, `workers` to the adaptive policy, `long_edge`
|
||||
//! to 480 (the app's preview proxy size) and the backend to `processes`.
|
||||
//! Set `OAK_BENCH_GENERATE=1` to synthesize a 1080p/25 fps 10 s clip at
|
||||
//! `<media>` when the file does not exist.
|
||||
|
||||
use std::path::PathBuf;
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
use oak_core::Rational;
|
||||
use oak_render::ipc::SLOT_FORMAT_BGRA8;
|
||||
use oak_core::{PixelFormat, Rational};
|
||||
use oak_render::procpool::{DispatcherConfig, ProcessDispatcher};
|
||||
use oak_render::ticket::{TicketPayload, TicketResult, VideoTicketParams};
|
||||
use oak_render::ticket::{
|
||||
Completion, Producer, TicketPayload, TicketResult, VideoTicketParams,
|
||||
};
|
||||
use oak_render::worker::{Job, JobDispatch, JobSchedule};
|
||||
|
||||
/// The comparison frame format: both backends must produce the same
|
||||
/// pixels for the numbers to be comparable (the app's 8-bit preview path
|
||||
/// uses BGRA8, but the process pool would then quantize — the F32 slots
|
||||
/// are the like-for-like path, and what the 10-bit preview uses).
|
||||
const BENCH_FORMAT: PixelFormat = PixelFormat::F32;
|
||||
|
||||
/// Locate the oak-worker binary (see bench_process).
|
||||
fn worker_bin() -> PathBuf {
|
||||
if let Ok(p) = std::env::var("OAK_WORKER_BIN") {
|
||||
@@ -59,66 +68,121 @@ fn worker_bin() -> PathBuf {
|
||||
PathBuf::from("oak-worker")
|
||||
}
|
||||
|
||||
fn main() {
|
||||
let media = std::env::args()
|
||||
.nth(1)
|
||||
.unwrap_or_else(|| "tests/demo.mp4".to_string());
|
||||
let frames: usize = std::env::args()
|
||||
.nth(2)
|
||||
.and_then(|s| s.parse().ok())
|
||||
.unwrap_or(240);
|
||||
let workers: Option<usize> = std::env::args().nth(3).and_then(|s| s.parse().ok());
|
||||
let long_edge: i32 = std::env::args()
|
||||
.nth(4)
|
||||
.and_then(|s| s.parse().ok())
|
||||
.unwrap_or(480);
|
||||
/// `(user, system)` CPU seconds of this process and its children.
|
||||
fn cpu_times() -> (f64, f64) {
|
||||
fn rusage(who: i32) -> (f64, f64) {
|
||||
let mut usage: libc::rusage = unsafe { std::mem::zeroed() };
|
||||
if unsafe { libc::getrusage(who, &mut usage) } != 0 {
|
||||
return (0.0, 0.0);
|
||||
}
|
||||
let seconds = |tv: libc::timeval| tv.tv_sec as f64 + tv.tv_usec as f64 / 1e6;
|
||||
(seconds(usage.ru_utime), seconds(usage.ru_stime))
|
||||
}
|
||||
let self_times = rusage(libc::RUSAGE_SELF);
|
||||
let children = rusage(libc::RUSAGE_CHILDREN);
|
||||
(
|
||||
self_times.0 + children.0,
|
||||
self_times.1 + children.1,
|
||||
)
|
||||
}
|
||||
|
||||
// The app's preview proxy size: the sequence's aspect scaled to the
|
||||
// long edge (demo.mp4 is 16:9 1080p).
|
||||
let (width, height) = ((long_edge as f64 * 16.0 / 9.0).round() as i32, long_edge);
|
||||
/// One footage ticket over the whole timeline.
|
||||
fn footage_params(media: &str, time: Rational, width: i32, height: i32) -> VideoTicketParams {
|
||||
VideoTicketParams {
|
||||
viewer: 1,
|
||||
project: String::new(),
|
||||
time,
|
||||
force_size: Some((width, height)),
|
||||
force_format: Some(BENCH_FORMAT),
|
||||
cache: None,
|
||||
cache_dir: None,
|
||||
cache_id: None,
|
||||
cache_timebase: None,
|
||||
footage: Some((media.to_string(), 0)),
|
||||
montage: Vec::new(),
|
||||
adjustments: Vec::new(),
|
||||
}
|
||||
}
|
||||
|
||||
fn report(entries: &[(i64, Instant, Instant)], start: Instant, elapsed: Duration, cpu: (f64, f64)) {
|
||||
let completed = entries.len();
|
||||
let throughput = completed as f64 / elapsed.as_secs_f64();
|
||||
let mut latencies: Vec<f64> = entries
|
||||
.iter()
|
||||
.map(|(_, submit, done)| (*done - *submit).as_secs_f64() * 1000.0)
|
||||
.collect();
|
||||
latencies.sort_by(|a, b| a.partial_cmp(b).unwrap());
|
||||
let first = entries
|
||||
.iter()
|
||||
.map(|(_, _, done)| (*done - start).as_secs_f64() * 1000.0)
|
||||
.fold(f64::INFINITY, f64::min);
|
||||
|
||||
let report = |name: &str, value: String| println!("{name:<38} {value}");
|
||||
report("frames completed", completed.to_string());
|
||||
report("total wall time", format!("{:.2} s", elapsed.as_secs_f64()));
|
||||
report(
|
||||
"throughput",
|
||||
format!(
|
||||
"{throughput:.1} fps ({:.1} ms/frame)",
|
||||
1000.0 / throughput.max(f64::EPSILON)
|
||||
),
|
||||
);
|
||||
report("first-frame latency", format!("{first:.1} ms"));
|
||||
if !latencies.is_empty() {
|
||||
let mean = latencies.iter().sum::<f64>() / latencies.len() as f64;
|
||||
report("completion latency mean", format!("{mean:.1} ms"));
|
||||
report(
|
||||
"completion latency p50/p95/max",
|
||||
format!(
|
||||
"{:.1} / {:.1} / {:.1} ms",
|
||||
latencies[latencies.len() / 2],
|
||||
latencies[((latencies.len() as f64 * 0.95) as usize).min(latencies.len() - 1)],
|
||||
latencies.last().unwrap()
|
||||
),
|
||||
);
|
||||
}
|
||||
report(
|
||||
"cpu user + sys",
|
||||
format!("{:.2} + {:.2} s", cpu.0, cpu.1),
|
||||
);
|
||||
report(
|
||||
"main-heap frame copies",
|
||||
oak_render::procpool::main_heap_frame_copies().to_string(),
|
||||
);
|
||||
}
|
||||
|
||||
/// Process-pool playback: post every frame at Playback priority and pump.
|
||||
fn run_processes(media: &str, frames: usize, width: i32, height: i32, workers: Option<usize>) {
|
||||
let config = DispatcherConfig {
|
||||
worker_bin: Some(worker_bin()),
|
||||
workers: workers.unwrap_or(0),
|
||||
slots_per_worker: 8,
|
||||
width,
|
||||
height,
|
||||
slot_format: SLOT_FORMAT_BGRA8,
|
||||
slot_format: BENCH_FORMAT as i32,
|
||||
batch_size: 0,
|
||||
graph_snapshot: None,
|
||||
handshake_timeout_ms: 30_000,
|
||||
};
|
||||
let dispatcher = ProcessDispatcher::new(config).expect("dispatcher config");
|
||||
dispatcher.start().expect("workers start + handshake");
|
||||
let worker_count = dispatcher.worker_count();
|
||||
println!("oak-worker pool: {worker_count} worker(s), {frames} x {width}x{height} BGRA8 frames of {media}");
|
||||
println!(
|
||||
"oak-worker pool: {} worker(s), {frames} x {width}x{height} {BENCH_FORMAT:?} frames of {media}",
|
||||
dispatcher.worker_count()
|
||||
);
|
||||
|
||||
// One completion record per frame: (ticket/frame, submit, completion).
|
||||
let cpu_start = cpu_times();
|
||||
let results = Arc::new(Mutex::new(Vec::<(i64, Instant, Instant)>::new()));
|
||||
let start = Instant::now();
|
||||
for i in 0..frames {
|
||||
let results = results.clone();
|
||||
let dc = dispatcher.clone();
|
||||
let frame = i as i64;
|
||||
let media_clone = media.clone();
|
||||
let media_clone = media.to_string();
|
||||
let job = Job {
|
||||
node_identity: 1,
|
||||
time: Rational::new(frame, 25),
|
||||
params: Arc::new(VideoTicketParams {
|
||||
viewer: 1,
|
||||
project: String::new(),
|
||||
time: Rational::new(frame, 25),
|
||||
force_size: Some((width, height)),
|
||||
force_format: None,
|
||||
cache: None,
|
||||
cache_dir: None,
|
||||
cache_id: None,
|
||||
cache_timebase: None,
|
||||
// A single footage clip covers the whole timeline.
|
||||
footage: Some((media_clone, 0)),
|
||||
montage: Vec::new(),
|
||||
adjustments: Vec::new(),
|
||||
}),
|
||||
params: Arc::new(footage_params(&media_clone, Rational::new(frame, 25), width, height)),
|
||||
audio: None,
|
||||
produce: Arc::new(|_, _| {
|
||||
Err(oak_render::error::Error::Failed(
|
||||
@@ -143,7 +207,6 @@ fn main() {
|
||||
eprintln!("frame {frame} failed: {e}");
|
||||
}
|
||||
}),
|
||||
// Playback priority, the pre-render window's schedule.
|
||||
schedule: JobSchedule::playback(frame, frame, 0),
|
||||
};
|
||||
if !dispatcher.post(job) {
|
||||
@@ -152,7 +215,6 @@ fn main() {
|
||||
}
|
||||
}
|
||||
|
||||
// Pump until every completion has landed.
|
||||
let deadline = Instant::now() + Duration::from_secs(300);
|
||||
loop {
|
||||
dispatcher.poll();
|
||||
@@ -167,44 +229,123 @@ fn main() {
|
||||
std::thread::sleep(Duration::from_millis(2));
|
||||
}
|
||||
let elapsed = start.elapsed();
|
||||
|
||||
let entries: Vec<(i64, Instant, Instant)> =
|
||||
results.lock().unwrap_or_else(|e| e.into_inner()).drain(..).collect();
|
||||
let completed = entries.len();
|
||||
let throughput = completed as f64 / elapsed.as_secs_f64();
|
||||
|
||||
// Per-frame completion latency (submit -> done), an end-to-end proxy
|
||||
// for the worker's per-frame render cost under load.
|
||||
let mut latencies: Vec<f64> = entries
|
||||
.iter()
|
||||
.map(|(_, submit, done)| (*done - *submit).as_secs_f64() * 1000.0)
|
||||
.collect();
|
||||
latencies.sort_by(|a, b| a.partial_cmp(b).unwrap());
|
||||
|
||||
let report = |name: &str, value: String| println!("{name:<38} {value}");
|
||||
report("frames completed", completed.to_string());
|
||||
report("total wall time", format!("{:.2} s", elapsed.as_secs_f64()));
|
||||
report(
|
||||
"throughput",
|
||||
format!("{throughput:.1} fps ({:.1} ms/frame)", 1000.0 / throughput.max(f64::EPSILON)),
|
||||
);
|
||||
if !latencies.is_empty() {
|
||||
let mean = latencies.iter().sum::<f64>() / latencies.len() as f64;
|
||||
report("completion latency mean", format!("{mean:.1} ms"));
|
||||
report(
|
||||
"completion latency p50/p95/max",
|
||||
format!(
|
||||
"{:.1} / {:.1} / {:.1} ms",
|
||||
latencies[latencies.len() / 2],
|
||||
latencies[((latencies.len() as f64 * 0.95) as usize).min(latencies.len() - 1)],
|
||||
latencies.last().unwrap()
|
||||
),
|
||||
);
|
||||
}
|
||||
report(
|
||||
"main-heap frame copies",
|
||||
oak_render::procpool::main_heap_frame_copies().to_string(),
|
||||
);
|
||||
|
||||
// Children (the worker pool) are only accounted at wait(): shut the
|
||||
// pool down before reading RUSAGE_CHILDREN, then report.
|
||||
dispatcher.shutdown();
|
||||
let cpu_end = cpu_times();
|
||||
let entries: Vec<(i64, Instant, Instant)> = results
|
||||
.lock()
|
||||
.unwrap_or_else(|e| e.into_inner())
|
||||
.drain(..)
|
||||
.collect();
|
||||
report(
|
||||
&entries,
|
||||
start,
|
||||
elapsed,
|
||||
(cpu_end.0 - cpu_start.0, cpu_end.1 - cpu_start.1),
|
||||
);
|
||||
}
|
||||
|
||||
/// M4 thread-pipeline playback: the same Playback jobs on the single
|
||||
/// render/decode threads; the pipeline prefetches each frame's decode on
|
||||
/// post and orders the queue by priority.
|
||||
fn run_pipeline(media: &str, frames: usize, width: i32, height: i32) {
|
||||
let backend = oak_render::pipeline::PipelineBackend::new().expect("pipeline start");
|
||||
println!("thread pipeline: 1 render + 1 decode thread, {frames} x {width}x{height} F32 frames of {media}");
|
||||
|
||||
let cpu_start = cpu_times();
|
||||
let results = Arc::new(Mutex::new(Vec::<(i64, Instant, Instant)>::new()));
|
||||
let start = Instant::now();
|
||||
for i in 0..frames {
|
||||
let frame = i as i64;
|
||||
let results = results.clone();
|
||||
let submitted = Instant::now();
|
||||
let done: Completion = Box::new(move |result: TicketResult| {
|
||||
if let Ok(TicketPayload::Video(_)) = result {
|
||||
results
|
||||
.lock()
|
||||
.unwrap_or_else(|e| e.into_inner())
|
||||
.push((frame, submitted, Instant::now()));
|
||||
}
|
||||
});
|
||||
let producer: Producer =
|
||||
Arc::new(|time, params| {
|
||||
oak_render::eval::render_produced_frame(time, params)
|
||||
.map(TicketPayload::Video)
|
||||
});
|
||||
let job = Job {
|
||||
node_identity: 1,
|
||||
time: Rational::new(frame, 25),
|
||||
params: Arc::new(footage_params(media, Rational::new(frame, 25), width, height)),
|
||||
audio: None,
|
||||
produce: producer,
|
||||
done,
|
||||
schedule: JobSchedule::playback(frame, frame, 0),
|
||||
};
|
||||
// The blocking post is the pipeline's backpressure: once the render
|
||||
// queue is full the submitter waits (the app's window is capped by
|
||||
// `preview_window_capacity`).
|
||||
if !backend.post(job) {
|
||||
eprintln!("post refused at frame {frame}");
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
let deadline = Instant::now() + Duration::from_secs(300);
|
||||
loop {
|
||||
let done = results.lock().unwrap_or_else(|e| e.into_inner()).len();
|
||||
if done >= frames {
|
||||
break;
|
||||
}
|
||||
if Instant::now() > deadline {
|
||||
eprintln!("timeout: {done}/{frames} completions");
|
||||
break;
|
||||
}
|
||||
std::thread::sleep(Duration::from_millis(2));
|
||||
}
|
||||
let elapsed = start.elapsed();
|
||||
let cpu_end = cpu_times();
|
||||
let entries: Vec<(i64, Instant, Instant)> = results
|
||||
.lock()
|
||||
.unwrap_or_else(|e| e.into_inner())
|
||||
.drain(..)
|
||||
.collect();
|
||||
report(
|
||||
&entries,
|
||||
start,
|
||||
elapsed,
|
||||
(cpu_end.0 - cpu_start.0, cpu_end.1 - cpu_start.1),
|
||||
);
|
||||
backend.shutdown();
|
||||
}
|
||||
|
||||
fn main() {
|
||||
let media = std::env::args()
|
||||
.nth(1)
|
||||
.unwrap_or_else(|| "tests/demo.mp4".to_string());
|
||||
let frames: usize = std::env::args()
|
||||
.nth(2)
|
||||
.and_then(|s| s.parse().ok())
|
||||
.unwrap_or(240);
|
||||
let workers: Option<usize> = std::env::args().nth(3).and_then(|s| s.parse().ok());
|
||||
let long_edge: i32 = std::env::args()
|
||||
.nth(4)
|
||||
.and_then(|s| s.parse().ok())
|
||||
.unwrap_or(480);
|
||||
let mode = std::env::args().nth(5).unwrap_or_else(|| "processes".to_string());
|
||||
|
||||
// The app's preview proxy size: the sequence's aspect scaled to the
|
||||
// long edge (demo.mp4 is 16:9 1080p).
|
||||
let (width, height) = ((long_edge as f64 * 16.0 / 9.0).round() as i32, long_edge);
|
||||
|
||||
if !std::path::Path::new(&media).exists() && std::env::var_os("OAK_BENCH_GENERATE").is_some() {
|
||||
oak_codec::testmedia::write_test_clip(std::path::Path::new(&media), 1920, 1080, 250, 25)
|
||||
.expect("generate the 1080p benchmark clip");
|
||||
println!("generated 1080p/25 fps test media at {media}");
|
||||
}
|
||||
|
||||
match mode.as_str() {
|
||||
"pipeline" => run_pipeline(&media, frames, width, height),
|
||||
_ => run_processes(&media, frames, width, height, workers),
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user