Commit Graph
19 Commits
Author SHA1 Message Date
Mike-Solar 7e87ec135e render: the M4 audit follow-ups — autocache priority, cancel-in-flight, decode LRU
- Autocache range jobs now post at Background priority
  (submit_video_background): they used to go through the Seek path and,
  after the M4 seek over-admission, jumped ahead of playback and past the
  render-queue bound. The interactive single-frame preview keeps Seek.
- Job.cancelled: the arena installs the slot's cancel atom, and
  execute_job finishes a cancelled job with Error::State before running
  the producer — a cancel no longer burns a full render/GPU pass only to
  discard the result. Exactly-once delivery is unchanged.
- DECODE_LRU_CAP 8 -> 2: the decode service's LRU is a hand-off buffer,
  not the cache of record (the eval-side decoded_frames LRU is); the
  double-cache footprint at 1080p F32 drops by ~6 frames. A hand-off miss
  is served from the eval cache without a new decode.
- Tests: sequence-aware preview cancel, over-admitted seek ordering,
  deterministic prefetch LRU reuse, cancelled-job skip, autocache
  priority. docs §3.4 backfilled with the A/B/C audit outcomes.
2026-09-15 19:29:17 +08:00
Mike-Solar ba1143e7a3 render: the M4 playback prefetch — dependency window, priorities and backpressure
docs/zh/plans/render-pipeline-threads.md M4: the thread pipeline now
keeps its decode thread ahead of the render thread and the app's
playback window consumes in-process frames.

- Render queue: priority-ordered by JobSchedule.priority (Seek >
  Playback > Background, FIFO within a class), so interactive frames
  jump playback exports/autocache. Seek posts may over-admit the bound:
  priority only reorders queued jobs, so a full queue of background work
  must not park the UI thread until an export frame finishes.
- Decode queue: rendezvous Requests are served ahead of queued
  Prefetches (a frame the renderer needs never waits behind speculative
  decodes); Sync barriers stay FIFO. The queue is a bounded
  Mutex+Condvar structure, preserving the request backpressure and the
  wait_idle contract.
- Playback read-ahead: a Playback job's footage decode requests are
  derived from its montage/footage spec on post (same media time, size
  and force_format.unwrap_or(F32) as the eval) and queued immediately,
  so frame N+1 decodes while frame N runs its GPU passes.
- App window: PreviewWindow slots are generalized to
  PreviewSlot::{Shm, Video}; the pipeline's in-process TicketPayload is
  cached and consumed by cpu_frame exactly like a worker slot.
  PipelineBackend::preview_window_capacity reports the render-queue
  headroom, so playback posts are capped to what the queue can take;
  cancel_preview_frame drops queued frames the playhead has passed,
  matched on the full (sequence, frame, version) key so one monitor's
  window never drops the other sequence's same-numbered frame.
- Tests: decode-queue preemption/FIFO, render-queue ordering, request
  derivation, and deterministic end-to-end M4 tests: a prefetch that
  must be reused by the render request (LRU hit, single decode — the
  read-ahead claim is falsifiable), a parked-render-thread priority test
  where a full queue of background work still lets a Seek over-admit and
  run first, and a sequence-aware cancel test. The playback prefetch
  smoke asserts prefetches == distinct decodes == frames; it does not
  claim zero heap copies (Frame.data is deep-copied at the eval-cache
  and service-LRU boundaries today).
- bench_playback gains a pipeline mode with CPU (self+children) and
  first-frame latency; both backends now produce F32 frames so the
  comparison is like-for-like. The §3.4 backfill records the numbers:
  at the proxy size the pipeline is faster with a lower first frame; at
  1080p peak throughput is below the multi-worker pool, but that is an
  artifact of the decode still being CPU software (M5), not a case for
  pooling decode threads — GPU decode is a single device/queue and the
  zero-copy import shares one GPU memory pool, so the single decode
  thread stays the target shape.
2026-09-15 17:25:01 +08:00
Mike-Solar 48e99e56b7 render: the M2 GPU zero-copy pipeline — wgpu 29, shared gpui device, GPU color LUTs
docs/zh/plans/render-pipeline-threads.md M2: the graph's textures stay
on the GPU from evaluation through presentation, and presentation runs
on the UI's own wgpu device.

- wgpu 25 -> 29 (naga 29) across the engine, unifying it with
  gpui_wgpu so engine textures are directly sampleable by the presenter
  (a single wgpu remains in the lockfile).
- GpuContext::adopt/install_shared: the app registers the window's
  device at startup and the render thread renders on it;
  texture_handle hands the raw Arc<wgpu::Texture> to
  SurfaceSource::Texture - zero-copy present on Linux/FreeBSD. The
  shared slot replaces an engine context that has not touched the GPU
  yet (startup-order guard) and refuses once it has.
- Texture::Gpu shares a GpuLease so clones release the registry token
  exactly once; the compositor, transitions and adjustment sweeps keep
  GPU textures end to end (no per-clip readbacks; GPU clears for
  black/generated frames).
- Color management stays on the GPU: the output node + display ICC
  chain is baked into a 65^3 3D LUT with the exact CPU reference and
  applied by the present WGSL pass (manual trilinear);
  ColorTransformJob bakes its OCIO processor the same way. Neither
  path skips color management.
- The explicit readback boundaries accept GPU textures: export
  encoder, CLI, worker shm, disk cache; CPU OpenFX already read back.
- M5 dependency: the YUV->RGB GPU pass (BT.601/709/2020 x
  limited/full) matches colormath::yuv444p16_to_rgb_f32.
- Acceptance: gpu_transfer_counters; single-clip and layered
  (multi-track + transition + adjustment) playback tests assert zero
  GPU->CPU readbacks, and the app test asserts adopted-device present
  is zero-copy. GPU tests hard-fail when OAK_REQUIRE_GPU is set (CI
  lavapipe) instead of skipping silently.
2026-09-12 20:52:17 +08:00
Mike-Solar a5b0b2a1b1 render: the M1 thread pipeline — one render thread, one decode thread
docs/zh/plans/render-pipeline-threads.md M1: an in-process
alternative to the worker-process pool, behind OAK_PIPELINE=threads
(processes stays the default and is fully retained).

- pipeline.rs: PipelineBackend implements JobDispatch over a single
  render thread draining a bounded FIFO (cap 8; blocking post with
  condvar backpressure and a one-ahead exception for re-posts from
  the render thread itself; shutdown drains with Error::State like
  the inline dispatcher). The DecodeService is a single decode
  thread behind a bounded command queue with a real LRU (tick-based
  eviction), rendezvous requests (None on shutdown -> the caller
  decodes inline), prefetch gated on render-queue room, and a Sync
  barrier; it installs into a process-wide slot that eval's footage
  path consults per frame (no service -> the synchronous decode it
  always was).
- The manager gains RenderBackendChoice::Pipeline; init() reads
  OAK_PIPELINE (threads -> pipeline, anything else -> the process
  pool), audio stays deliberately inline.
- Present mapping: the UI thread consumes through the ticket
  completion, unchanged — no fourth thread is invented.
- Tests: decode-service unit tests (rendezvous, LRU hit/eviction,
  error propagation, backpressure gate) plus a six-case integration
  suite matrixed over inline vs pipeline — consecutive-frame and
  out-of-order seek pixel equality asserted byte for byte, with
  decode counters proving the service (not the caller) did the
  codec work.
2026-09-11 19:37:27 +08:00
Mike-Solar 4f0f5cbba6 workspace: zero compiler warnings across all targets
254 warnings (320 counting replayed-cache re-emitters) cleaned:
unused mut/imports/variables, irrefutable if-lets and unreachable
patterns, dead code removed or annotated #[allow(dead_code)] with
the reason (C++ parity value sets, cfg(test) helpers, public API
reservations), drop(&ref) no-ops removed, fn-pointer identity via
std::ptr::fn_addr_eq, the test-stubs feature declared in
oak-node's manifest, missing docs filled. Every unused-Result site
was judged individually: meaningful errors propagate, intentional
ignores are let _ = with a note.

Two pre-existing latent bugs are documented in place, behavior
preserved: app.rs's timeline-tool observer and dialogs.rs's format
subscription both drop the returned Subscription immediately, so
they never fire.
2026-09-11 16:38:44 +08:00
Mike-Solar 2817ac286c timeline: drag generators onto the timeline, drop transitions at clip edges
Generator effects (bars, checkerboard) can be dragged from the library
onto the timeline, where they land as a standalone five-second clip
built from the node factory; the inspector shows the generator's
parameters as the clip's own chain.

Transitions are no longer junction-only. The render planner accepts a
transition with at least one wired neighbor and blends the missing
side against transparent black, so head transitions fade in from black
and tail transitions fade out to black. add_transition_at_edge creates
those single-sided blocks (wired to just the IN or OUT block), the
default-transition command covers both ends of a lone clip, and an
effect drag dropped near a clip edge routes to the nearest seam or
edge within a one-second window.
2026-09-11 08:49:08 +08:00
Mike-Solar 16364d414a nodes: mask off-frame samples to transparent in the distort shaders
Translating, rotating or warping content past the frame edge used to
smear the clamped edge row/column across the vacated region. The
transform, position, swirl, ripple and wave shaders now multiply the
sample by an in-bounds mask so off-frame pixels come out transparent
(and composite as black when nothing sits below). Tile deliberately
keeps its wrapping lookup.
2026-09-11 08:48:48 +08:00
Mike-Solar 4a2614b3fc timeline: adjustment layers and first-class transitions
Adjustment layers (docs/zh/plans/adjustment-layers-and-transitions.md):
a new timeline block type whose effect chain grades the composite of
every video track below it, over its own range (spanning clips or a
slice of one). The graph path flushes the lower tracks at the block's
track boundary and sweeps the composite through the chain via a
transient texture-source node; the montage path mirrors it with
AdjustmentSpan tickets (wire-compatible), so worker previews and
exports agree. An empty-area context menu creates one; the block
trims/moves/deletes like a clip, with undo everywhere.

Transitions: seam blocks come alive - cross dissolve/fade/wipe/slide
evaluate both neighbors through the graph path with progress from the
transition's own range (never the whole clip). Ctrl+Shift+D or the clip
menu inserts a default transition; the gpui wedges render and drag to
resize offsets undoably, and TransitionRemoveCommand now restores
offsets and edges on undo. The transitionfx node form runs the same
shaders on an adjustment layer with progress_in auto-filled from the
layer's span (explicit value wins).

Also: every built-in effect name and parameter name is now
translatable (360 node.* keys per locale, zh-CN fully translated, two
coverage tests guard future gaps); the new nodes register in
nodes/mod.rs with the factory smoke table updated; textfootage and
adjustment-layer i18n keys included.
2026-09-10 22:03:15 +08:00
Mike-Solar 1e0d48578e nodes: text becomes structured footage with a real text engine
Text is no longer a hand-written HTML effect (docs in
docs/zh/plans/text-footage-redesign.md):
- textv3 gains structured inputs - plain text, font family/size, font
  color, outline (enable/color/width), glow (enable/color/radius); the
  legacy text_in HTML is hidden and auto-migrated to plain text on load.
- Outline and glow render as GPU post-process chains (dilate/blur +
  colorize under the text); the font color tints the raster
  premultiplied. Plain text now rasterizes even with both passes off
  (previously a null deferred job), fixing a use-after-free where the
  handle was lifted out of an owning Option<NodeValue> before addref.
- A cosmic-text backend (the lockfile's 0.19) installs at engine
  startup through the textbackend hooks and feeds the font-family combo.
- The project panel gains 添加文本素材 next to 新建序列: a text entry
  in the bin that drops onto the timeline as a clip (one undo row), its
  parameters shown as structured fields in the inspector (multiline
  text area, no HTML anywhere). text3 is hidden from the effect add
  menus; legacy text3 chains keep evaluating.
2026-09-10 22:02:57 +08:00
Mike-Solar a7916aa93d nodes: OpenFX-Misc cleanroom GPU ports, grouped and collapsible in the library
21 built-in effects reimplemented as native GPU nodes from the
OpenFX-Misc algorithm references (cleanroom, docs in
docs/zh/plans/ofx-misc-gpu-cleanroom.md):
- Color: Color Correct, Gamma, Saturation, Invert, Clamp, Grade
- Matrix/morphology: Color Matrix, Edge Detect, Dilate, Erode
- Blur: Directional Blur, Sharpen (unsharp mask)
- Merge: Dissolve, Key Mix, Premultiply, Unpremultiply
- Geometry/generators: Position, Mirror, Checkerboard, Color Bars, Ramp

Every node carries unit tests plus GPU pixel tests (28 cases over five
ofxmisc_* suites). The effect library groups built-ins by category
(color/filter/distort/keying/generator/math/general) with collapsible
group headers persisted to the config; the inspector's add menu groups
the same way. Registration wiring and the factory smoke table land with
the adjustment/transition wave sharing the same files.
2026-09-10 22:02:38 +08:00
Mike-Solar fc90604245 render: anchor resolution_in to the sequence resolution
Preview renders at proxy size while a paused frame renders full-res, so
anchoring resolution_in to the render target made every sequence-pixel
effect (shape size/pos, transform offsets, corner pin points, drop
shadow distance) change apparent size whenever the transport stopped.
Pre-fill resolution_in from the sequence's video params (C++ inserts
the NodeGlobals square resolution at job-build time), covering the
nested generator job inside a merge as well; a node that inserted its
own resolution_in keeps it.
2026-09-10 16:53:55 +08:00
Mike-Solar 4f404f8cc6 tests: replace placebo assertions with real behavior checks 2026-09-09 16:33:35 +08:00
Mike-Solar 4babbf5de8 core: merge oak-common into oak-core
CI / Build & test (Linux) (push) Successful in 24m6s
CI / Build & test (Windows) (push) Successful in 31m14s
oak-common is gone; its modules (configstore, xmlutils, ocioutils,
oiioutils, colormath, colortransform, videoparams, ffmpegutils, ...)
now live in oak-core alongside the value types. The render value/GPU
types moved too: backend (wgpu context + DisplayRenderer), color
(ColorProcessor over ocio-rs), texture, frame, and the commonutil
config helpers.

Fix-ups to make the merged tree build and pass tests:

- oak-core Cargo.toml: wgpu back to 25 (the moved backend code is
  written against that API generation); add the toml/quick-xml/image
  deps oak-common carried.
- lib.rs: drop the duplicate 'pub mod error;'.
- error.rs: unified OAKCORE_* codes; restore Error::new() and
  From<OcioError> from oak-common's error type.
- backend.rs/color.rs: oak_core::/oak_render:: self-references
  rewritten to crate::; the shaderfx-dependent GPU effect test moved
  to oak-render's shaderfx tests (shaderfx depends on oak-node and
  cannot live in oak-core).
- oak-render's error module re-exports oak_core::error::{Error,
  Result}; the OAKRENDER_* codes stay as the public-code contract.
- oak-node jobs.rs: ColorProcessor imported from oak_core::color.
- Integration tests repointed at oak_core::{texture, frame, backend,
  color, colormath}.
- the display-ICC regression test treats an empty OAK_DISPLAY_ICC as
  unset, matching displayicc::env_override_icc.
2026-09-03 17:42:20 +08:00
Mike-Solar 5937e557a7 app: project-explorer new-sequence button, 4K presets, keying node fix, color-picker canvases
- Project explorer header gains a 新建序列 button (opens the existing
  new-sequence dialog, seeded like the menu action).
- Sequence presets gain 4K UHD (3840x2160@25) and 4K DCI (4096x2160@24);
  the sequence-properties dialog re-selects them on reopen. Format fields
  are also seedable from a probed footage format (the drop flow).
- Chroma Key (and Color Difference Key) value() now box a ShaderJobPayload
  like Despill: the old OCIO-processor gate pushed nothing (the processor
  is never populated without the render bridge), so the traverser handed
  the clip NodeValue::None and the rendered frame lost the clip. The
  renderer resolves the OCIO stub at compile time from OCIO_SHADER_STUBS.
  End-to-end graph test: green key on a green frame keys out, red key
  keeps it.
- The OFX color picker's SV palette / hue bar / preview / swatch canvases
  get size_full(): the bare canvases collapsed to zero height in the
  block layout, so the palette painted nothing (the reported 色板没显示).
  Regression test clicks the palette center and expects mid s/v.
2026-08-30 21:58:26 +08:00
Mike-Solar fdb5caabd5 color: non-sRGB preview, per-monitor display ICC, pipeline hardening
Preview now follows the project output colorspace end to end: the
display chain derives its content space from the project's OutputColorSpec
instead of a hardcoded sRGB name, self-managed ICC transforms go through
an XYZ D65 interchange stage (OCIO cie_xyz_d65_interchange) for non-sRGB
targets, and the platform layer declares the content colorspace (gpui
submodule bump). macOS defaults to OS-managed (fixes wide-gamut UI
oversaturation); Windows ACM warns once on non-sRGB targets.

Multi-monitor: the display ICC is looked up per the window's current
screen (macOS display id, Windows per-monitor DC, X11 RandR output
profile) with a throttled poll that invalidates frame caches on moves.

Pipeline precision: 10-bit+ sources fall back to YUV444P16LE + a Rust
matrix conversion when swscale lacks F32 output (no more 8-bit
truncation); BT.709/2020 SDR decodes with BT.1886 gamma 2.4 instead of
the sRGB EOTF; working-space compositing no longer clamps RGB to [0,1]
(alpha still clamped); the output node clamps to the target gamut;
frames without colorimetry metadata convert with BT.709 defaults
(warned once) instead of passing through; scopes read the
output-colorspace signal on both F32 paths.

Also: only emit rerun-if-changed for .env when it exists (a missing file
made every build fully dirty).
2026-08-29 00:24:15 +08:00
Mike-Solar c03f1ec604 render: stack higher-numbered video tracks on top
Match the timeline UI (V_max drawn topmost): composite tracks from
V1 up to V_max so the highest-numbered track is composited last, in
both the montage path and direct graph evaluation.
2026-08-27 15:20:18 +08:00
Mike-Solar c51a349070 render: translate node GLSL shaders to WGSL and run them as wgpu passes
CI / Build & test (Linux) (push) Failing after 16m57s
CI / Build & test (Windows) (push) Successful in 31m46s
2026-08-27 07:07:09 +08:00
Mike-Solar fa3951344b render: fix 4K playback memory growth and decode-to-target-size
Root causes found for the 4K stalls and the second-footage memory
blowup (audit + code review):

- ticket bookkeeping leaked unbounded: the procpool ticket table and
  the arena slot map only ever grew (50-100 tickets/sec during
  playback, each pinning montage params and shm region views).
  Completed/cancelled/superseded/crashed entries are now removed, and
  the arena reaps fire-and-forget tickets once finished; the sync poll
  path reaps via a terminal result() read. InFlight duplicate submits
  now answer State immediately instead of sitting in the map forever.
- decode ran a full-resolution swscale to F32 RGBA (~132 MB at 4K)
  plus a second full-res copy before downscaling to the 480px proxy:
  RetrieveVideoParams.target_size lets swscale convert AND resize in
  one pass (bilinear, matching the old Rust resampler), so a 4K
  preview frame costs ~1 MB instead of ~260 MB of churn. This applies
  to proxy AND full-res requests alike.
- per-process decoder cache was unbounded (each session pins an FFmpeg
  context + 2 native decoded frames): LRU-capped at 16, eviction drops
  the map entry (in-flight renders keep their Arc; Drop releases
  FFmpeg).
- playback window completions were not generation-gated: a stale
  render from before an edit landed in the rebuilt window (wrong frame
  displayed, fresh request blocked). Stale completions now return
  their shm slot credit instead.
- async audio prefetch used the polling ticket submit without ever
  polling: switched to the fire-and-forget submit so entries reap.
2026-08-26 01:31:10 +08:00
Mike-Solar 244d5e860f workspace: kebab-case crates, app under crates/oak-app, shared versions
CI / Build & test (Windows) (push) Failing after 7s
All crates take the oak-* kebab-case naming (oak-audio, oak-codec,
oak-common, oak-core, oak-ffmpeg-link, oak-node, oak-otio, oak-plugin,
oak-render, oak-storage, oak-task, oak-timeline, oak-undo), with the
lib identifiers rewritten (oakrender:: -> oak_render::, oakcore_rs:: ->
oak_core::, ...) across all 226 referencing files.

The GUI application moves from the workspace root into
crates/oak-app/: src/, build.rs (paths fixed for the new location) and
tests/ travel with it, the root Cargo.toml becomes workspace-only
([workspace] + workspace.package + profiles), and the app package
inherits the workspace version. The screenshots example becomes a
standalone crate examples/simple_player/ with its own Cargo.toml.

Every crate now inherits the single workspace version
(version.workspace = true), and the workflows' crate paths and the
build docs follow the renames.

Validated with a clean cargo check --workspace.
2026-08-22 16:58:37 +08:00