render: the M2 GPU zero-copy pipeline — wgpu 29, shared gpui device, GPU color LUTs

docs/zh/plans/render-pipeline-threads.md M2: the graph's textures stay
on the GPU from evaluation through presentation, and presentation runs
on the UI's own wgpu device.

- wgpu 25 -> 29 (naga 29) across the engine, unifying it with
  gpui_wgpu so engine textures are directly sampleable by the presenter
  (a single wgpu remains in the lockfile).
- GpuContext::adopt/install_shared: the app registers the window's
  device at startup and the render thread renders on it;
  texture_handle hands the raw Arc<wgpu::Texture> to
  SurfaceSource::Texture - zero-copy present on Linux/FreeBSD. The
  shared slot replaces an engine context that has not touched the GPU
  yet (startup-order guard) and refuses once it has.
- Texture::Gpu shares a GpuLease so clones release the registry token
  exactly once; the compositor, transitions and adjustment sweeps keep
  GPU textures end to end (no per-clip readbacks; GPU clears for
  black/generated frames).
- Color management stays on the GPU: the output node + display ICC
  chain is baked into a 65^3 3D LUT with the exact CPU reference and
  applied by the present WGSL pass (manual trilinear);
  ColorTransformJob bakes its OCIO processor the same way. Neither
  path skips color management.
- The explicit readback boundaries accept GPU textures: export
  encoder, CLI, worker shm, disk cache; CPU OpenFX already read back.
- M5 dependency: the YUV->RGB GPU pass (BT.601/709/2020 x
  limited/full) matches colormath::yuv444p16_to_rgb_f32.
- Acceptance: gpu_transfer_counters; single-clip and layered
  (multi-track + transition + adjustment) playback tests assert zero
  GPU->CPU readbacks, and the app test asserts adopted-device present
  is zero-copy. GPU tests hard-fail when OAK_REQUIRE_GPU is set (CI
  lavapipe) instead of skipping silently.
This commit is contained in:
2026-09-12 20:52:17 +08:00
parent a5b0b2a1b1
commit 48e99e56b7
29 changed files with 3014 additions and 782 deletions
+14 -1
View File
@@ -673,12 +673,25 @@ pub fn render_frame(
data: frame.data.clone(),
})
}
Ok(TicketPayload::Video(texture @ oak_core::texture::Texture::Gpu { .. })) => {
// M2: the thread pipeline renders all-GPU; the CLI writes CPU
// pixels, so this is an explicit readback boundary.
let frame = texture
.to_frame()
.map_err(|e| format!("render readback: {e:?}"))?;
Ok(RenderedFrame {
width: frame.width,
height: frame.height,
format: frame.format as i32,
linesize: frame.linesize_bytes() as i32,
data: frame.data,
})
}
Ok(TicketPayload::ShmFrame(frame)) => {
let out = shm_to_rendered_frame(frame);
m.release_frame(frame);
Ok(out)
}
Ok(TicketPayload::Video(_)) => Err("render produced a non-CPU frame".to_string()),
_ => Err("render produced no video frame".to_string()),
}
}