Commit Graph
22 Commits
Author SHA1 Message Date
Mike-Solar 6153a2ac33 tests: fix three stale assertions
CI / Build & test (Windows) (push) Failing after 13m15s
CI / Build & test (Linux) (push) Successful in 24m22s
- multicamnode: type_id() assertions called the method on dyn
  NodeBehavior, which method resolution routed to std::any::Any's
  TypeId::of — qualify via NodeBehavior::type_id so the trait method
  (the &str node type id) is compared
- procpool: the per-worker GPU budget grew a security headroom (×2)
  for the CUDA-OOM flood; the two budget tests now assert against the
  real formula (2 GiB + 256 MiB at 1080p24, 2 workers at 4K/24 GiB)
- cli info fixture: the fixture runs 29.97 fps; the assertion expected
  30/1 (stale from the older fixture)
2026-09-02 12:34:39 +08:00
Mike-Solar be42620e22 multicam: throttle angle refresh, shrink pre-render window
- MulticamPanel: playback angle refresh runs one cycle per 3 ticks
  instead of re-requesting every source every tick — each angle decode
  is a keyframe-scanning FFmpeg seek that stole worker capacity from
  the main viewer
- PreRender frames default 120 -> 12: a window larger than what the
  pool can render in real time queues far ahead of the playhead, so
  the painted frame lags seconds behind (playback frozen); a smaller
  window keeps the backlog bounded
- render_graph_frame: OAK_PERF clip-level timing
2026-09-01 21:52:56 +08:00
Mike-Solar 7f41570596 render: stop evicting the only hardware decoder session on every open
The per-process 'hardware session == 1, evict before every new open'
guard was turning every frame's open into a decoder re-open: after
inserting the fresh session, the NEXT frame's pre-open eviction dropped
it again, so no request ever hit the cache ([open] (request) on every
single frame, zero CACHED hits). Every decode then cost a full FFmpeg
session open (~0.5 s) + a keyframe-scanning seek — playback could never
keep up (the 'main viewer barely moves' report; [perf] showed 1.6-2.2 s
per frame).

Hardware VRAM is bounded by the LRU cap (hardware-first eviction at
MAX_CACHED_DECODERS) and the gpu-vram worker-count policy; the eager
pre-open eviction was the regression.
2026-09-01 21:46:01 +08:00
Mike-Solar 1b0a15192e multicam: wizard + AFV + node-graph preview + performance
CI / Build & test (Linux) (push) Failing after 24s
CI / Build & test (Windows) (push) Failing after 3m38s
- wizard: angle multi-select, sync modes, auto-align, create sequence;
  keeps the host sequence current (host clip = the multicam clip)
- build_multicam_sequence: per-angle clip -> SOURCES_INPUT[element],
  array slot growth, audio angle tracks for AFV
- AFV: host linked audio clip follows the switched source (one undo),
  muted host audio track disables it
- node-graph preview: inline producer renders through the traverser
  (viewer/project-matched graph frames); sequence viewer uses the graph
- multicam node value(): element-tagged row keys (sources_in[i]),
  reads the current source; build_row keys array inputs by element
- angle grid: viewer=0 (single-track montage, not whole-graph)
- project explorer: rename (dialog) + delete (undoable) real items
- performance: decoded-frame LRU, per-process NVDEC quota (1 session,
  evict before open), GPU composite fail-once fallback, snapshot
  upload debounced on the engine tick, worker vram budget headroom
- timeline clip: multicam overlay via ClipDecorator
- wizard menu item moved to Sequence menu
2026-09-01 20:27:59 +08:00
Mike-Solar 8de9af705e render: query GPU vram on AMD/Intel Linux too, document the UMA/Windows fallbacks
gpu_vram_bytes() now chains per vendor/platform:
- NVIDIA everywhere: nvidia-smi (shipped by the NVIDIA driver on every
  OS) — the NVDEC path's primary device.
- AMD/Intel on Linux: the DRM mem_info_vram_total/used sysfs attributes
  (amdgpu, i915, Xe). Free = total - used; the first non-zero card wins
  (an iGPU without dedicated vram reports 0 and is skipped). The walk is
  now testable via an injectable sysfs root; a card missing the attrs is
  skipped, never aborts the walk.
- Apple Silicon: unified memory — no separate vram exists; the RAM/4
  budget IS the correct bound for decode surfaces and render targets, so
  no query (an explicit vram budget would double-count the same pool).
- Windows AMD/Intel: no portable CLI; DXGI QueryVideoMemoryInfo is the
  real API but wgpu 25 does not expose it. Falling back to the RAM
  policy is safe-side (under-sized pool loses throughput, never OOMs the
  device).

Tests: the sysfs walk (fixture with a missing-attr card, a 0-total
iGPU and a discrete winner) and the per-worker budget scaling (1080p
baseline, 4K ~4x, 60 fps over-provision).
2026-08-30 22:03:05 +08:00
Mike-Solar 9352fca4a9 codec/render: silence hw-decode failures, vram-aware dynamic worker pool
- open_hw_accel marks the device unavailable when the decoder OPEN fails
  (cuvidCreateDecoder OOM at 4K) too, not just device-context creation:
  without it every subsequent decoder session retried CUDA and flooded
  the log per open.
- Decoder gains hardware_decoding(); the oak-render decode-session LRU
  evicts hardware sessions first (each pins a GPU surface pool — ~100 MB
  at 4K), so a full cache cannot exhaust video memory before the next
  open.
- Worker pool count now factors GPU vram: per-worker budget = 1 GiB
  (1080p peak) scaled by pixel ratio + 256 MiB idle floor, 10% reserve
  of free vram; applied when hardware decoding is on (nvidia-smi query,
  None otherwise falls back to the RAM/CPU policy).
- Dynamic pool resize: ProcessDispatcher::set_target_workers grows or
  retires workers; retiring ones stop claiming, drain their in-flight
  batch (future playback frames included), then exit naturally on the
  shutdown signal — no mid-work kill (30 s deadline only as a hung-
  decoder last resort). A retiring worker that dies re-queues its frames
  to surviving workers. Resizes are throttled to 2 s (a resolution burst
  merges; only the latest target applies) so 1080p<->4K flaps cannot
  thrash process spawns.
- RenderManager::set_workspace_size announces the sequence resolution;
  RealEngine calls it from refresh_sequence_info.
- Integration test: shrink 3->1 mid-wave (all frames complete, retired
  workers exit naturally) then regrow 1->3 and render a fresh wave.
2026-08-30 21:59:06 +08:00
Mike-Solar 5937e557a7 app: project-explorer new-sequence button, 4K presets, keying node fix, color-picker canvases
- Project explorer header gains a 新建序列 button (opens the existing
  new-sequence dialog, seeded like the menu action).
- Sequence presets gain 4K UHD (3840x2160@25) and 4K DCI (4096x2160@24);
  the sequence-properties dialog re-selects them on reopen. Format fields
  are also seedable from a probed footage format (the drop flow).
- Chroma Key (and Color Difference Key) value() now box a ShaderJobPayload
  like Despill: the old OCIO-processor gate pushed nothing (the processor
  is never populated without the render bridge), so the traverser handed
  the clip NodeValue::None and the rendered frame lost the clip. The
  renderer resolves the OCIO stub at compile time from OCIO_SHADER_STUBS.
  End-to-end graph test: green key on a green frame keys out, red key
  keeps it.
- The OFX color picker's SV palette / hue bar / preview / swatch canvases
  get size_full(): the bare canvases collapsed to zero height in the
  block layout, so the palette painted nothing (the reported 色板没显示).
  Regression test clicks the palette center and expects mid s/v.
2026-08-30 21:58:26 +08:00
Mike-Solar 9dab9efc38 codec: carry codec-frame overflow across audio chunk boundaries
retrieve_audio_to decodes whole codec frames but copies only the part
inside the requested chunk; the tail of the frame crossing the chunk end
(up to 1023 samples for AAC) was consumed by the decoder and lost, so
the next chunk started with a hole. On the playback grid (1920 samples
at 25fps/48kHz) the hole cycles 128..896 samples and hits 7 of 8 chunk
starts -- the heavy stutter/noise heard during playback.

Keep the overflow (decode and resampler-flush tails) in a per-session
carry buffer and serve it at the start of the next contiguous chunk;
clear it on seek, format change and chunk failure. Also make seek()
actually drop the cached resampler as its comment claimed.

Verified sample-exact: chunked decode of a 440Hz tone now matches a
one-shot decode bit for bit, and chunked renders of real media line up
with the ffmpeg CLI reference at correlation 1.0 / drift 0.

Adds a regression test (playback_sized_chunks_match_oneshot_sample_exact)
with a new tone fixture -- demo.mp4's audio is -90dB digital silence and
cannot expose the holes -- plus a render_audio_wav example used to
diagnose chunk-boundary artifacts offline.
2026-08-29 19:51:59 +08:00
Mike-Solar e128a6aca1 render: clamp mixed audio to [-1,1] before it reaches the device
Overlapping clips sum linearly in mix_audio_montage and can exceed full
scale (two hot clips reach +/-2; gain > 1 would too); the cpal sink
forwarded samples unclamped, so overlaps clipped at the DAC. Clamp the
accumulator after the montage mix (both the heap and shm-slot paths
share mix_audio_montage) and document it on render_audio_samples.
2026-08-29 19:51:25 +08:00
Mike-Solar cbd4ba2442 optimize: video and audio render
CI / Build & test (Linux) (push) Canceled after 0s
CI / Build & test (Windows) (push) Canceled after 0s
2026-08-29 05:01:55 +08:00
Mike-Solar fdb5caabd5 color: non-sRGB preview, per-monitor display ICC, pipeline hardening
Preview now follows the project output colorspace end to end: the
display chain derives its content space from the project's OutputColorSpec
instead of a hardcoded sRGB name, self-managed ICC transforms go through
an XYZ D65 interchange stage (OCIO cie_xyz_d65_interchange) for non-sRGB
targets, and the platform layer declares the content colorspace (gpui
submodule bump). macOS defaults to OS-managed (fixes wide-gamut UI
oversaturation); Windows ACM warns once on non-sRGB targets.

Multi-monitor: the display ICC is looked up per the window's current
screen (macOS display id, Windows per-monitor DC, X11 RandR output
profile) with a throttled poll that invalidates frame caches on moves.

Pipeline precision: 10-bit+ sources fall back to YUV444P16LE + a Rust
matrix conversion when swscale lacks F32 output (no more 8-bit
truncation); BT.709/2020 SDR decodes with BT.1886 gamma 2.4 instead of
the sRGB EOTF; working-space compositing no longer clamps RGB to [0,1]
(alpha still clamped); the output node clamps to the target gamut;
frames without colorimetry metadata convert with BT.709 defaults
(warned once) instead of passing through; scopes read the
output-colorspace signal on both F32 paths.

Also: only emit rerun-if-changed for .env when it exists (a missing file
made every build fully dirty).
2026-08-29 00:24:15 +08:00
Mike-Solar 9c0269bcbb render: cache worker frames and decouple audio from the video workers
- oak-worker gains an LRU frame cache (default 64 MiB) keyed by the
  render-deterministic spec subset; repeat frames (paused frames,
  scrubs over rendered ranges, re-renders after an effect change) are
  memcpy-cheap, which is what made adding an OFX plugin -- and the
  in-flight batches after removing one -- stall the UI
- audio dispatch on the Processes backend now mixes inline on the UI
  tick instead of queueing behind video batches in the worker pool, so
  video stalls no longer starve the ~100ms cpal output buffer into
  silence
2026-08-27 20:35:50 +08:00
Mike-Solar 5d21f83e1e render/app: display bit depth option (10-bit default, 8-bit optional)
Preferences gains a display-bit-depth combo stored in the config store;
at startup the app forwards it to gpui_wgpu via OAK_DISPLAY_BIT_DEPTH,
which prefers Rgb10a2Unorm for 10-bit presentation. Takes effect after
restart (noted in the dialog); i18n in all eight packs.
2026-08-27 19:08:55 +08:00
Mike-Solar c03f1ec604 render: stack higher-numbered video tracks on top
Match the timeline UI (V_max drawn topmost): composite tracks from
V1 up to V_max so the highest-numbered track is composited last, in
both the montage path and direct graph evaluation.
2026-08-27 15:20:18 +08:00
Mike-Solar f7352ae19d render: fall back to CPU when the adapter cannot render Rgba32Float
CI / Build & test (Linux) (push) Successful in 19m8s
CI / Build & test (Windows) (push) Successful in 32m35s
2026-08-27 13:52:03 +08:00
Mike-Solar c51a349070 render: translate node GLSL shaders to WGSL and run them as wgpu passes
CI / Build & test (Linux) (push) Failing after 16m57s
CI / Build & test (Windows) (push) Successful in 31m46s
2026-08-27 07:07:09 +08:00
Mike-Solar fa3951344b render: fix 4K playback memory growth and decode-to-target-size
Root causes found for the 4K stalls and the second-footage memory
blowup (audit + code review):

- ticket bookkeeping leaked unbounded: the procpool ticket table and
  the arena slot map only ever grew (50-100 tickets/sec during
  playback, each pinning montage params and shm region views).
  Completed/cancelled/superseded/crashed entries are now removed, and
  the arena reaps fire-and-forget tickets once finished; the sync poll
  path reaps via a terminal result() read. InFlight duplicate submits
  now answer State immediately instead of sitting in the map forever.
- decode ran a full-resolution swscale to F32 RGBA (~132 MB at 4K)
  plus a second full-res copy before downscaling to the 480px proxy:
  RetrieveVideoParams.target_size lets swscale convert AND resize in
  one pass (bilinear, matching the old Rust resampler), so a 4K
  preview frame costs ~1 MB instead of ~260 MB of churn. This applies
  to proxy AND full-res requests alike.
- per-process decoder cache was unbounded (each session pins an FFmpeg
  context + 2 native decoded frames): LRU-capped at 16, eviction drops
  the map entry (in-flight renders keep their Arc; Drop releases
  FFmpeg).
- playback window completions were not generation-gated: a stale
  render from before an edit landed in the rebuilt window (wrong frame
  displayed, fresh request blocked). Stale completions now return
  their shm slot credit instead.
- async audio prefetch used the polling ticket submit without ever
  polling: switched to the fire-and-forget submit so entries reap.
2026-08-26 01:31:10 +08:00
Mike-Solar f2aab8ce15 render: apply clip effect stacks in the montage path
CI / Build & test (Windows) (push) Failing after 16m26s
CI / Build & test (Linux) (push) Successful in 19m12s
Adding an effect to a clip did nothing: the sequence render is
flattened into a montage (decode + composite), and MontageClip carried
no effect data at all.

- MontageClip gains an ordered effect stack (type id / enabled /
  effect input / parameter values); protocol v2 carries it as an
  additive wire field (older peers default to an empty stack).
- renderops::video_montage fills the stack from the effect chain
  (the footage source node — the chain end without an effect input —
  is dropped; the montage decodes the footage itself). Export
  (oak-task) and the multicam single-track montage fill it too.
- The worker applies the stack between decode and composite: built-in
  Opacity gets a CPU evaluator (C++ opacity.frag parity — whole vec4,
  alpha included, unity pass-through); everything else dispatches as an
  OFX plugin job through a new instance-factory slot (oak-plugin
  lazily creates + caches one instance per identifier per render
  process) with the montage's parameters injected. Disabled effects
  bypass (the C++ traverser pushes the effect input through). Unknown
  types warn once per type id and pass through — no silent no-ops.

Not covered (explicitly): Transform/Crop and the other ~30 built-in
effects have no CPU evaluator in oak-render (they pass through with a
warning), keyframed parameter animation, audio effect chains, and the
CLI's simplified montage.

Acceptance: a real 50% Opacity on real media quarters the rendered
pixels both in-process (renderops test) and through a real worker
process over IPC + shared memory (procpool_integration test);
disabling restores the plain render byte-for-byte.
2026-08-25 04:49:00 +08:00
Mike-Solar 476114cec0 app: project properties dialog (File > Project Properties…)
The menu item was a placeholder print; it now opens a real dialog (the
C++ ProjectPropertiesDialog):

- Per-project OCIO config override with a 浏览… picker: validated on OK
  (an invalid config keeps the dialog open with the error shown, like
  the C++ accept()), persisted in the project settings, applied to the
  display color pipeline on accept and on project open, and reverted to
  the app default when the project closes. oak-render gains
  set_up_default_config_from for the explicit-path load.
- Disk-cache location (default / alongside the project / custom path):
  persisted through the OVE serializer (cachesetting/customcachepath
  round-trip the settings map, clamped on load) and honored by the
  thumbnail writer — the first live consumer of Project::cache_path.
- PathField gains an enabled state (the custom path field follows the
  combo selection).

The C++ color tab's Default Input Color Space and Reference Space
combos are intentionally absent: the Rust render pipeline has no
consumer for them today (decode performs no input transfer conversion),
so showing them would be dead settings.

Tests: dialog opens, OK applies the cache location, an invalid OCIO
config keeps the dialog open with the error row. i18n keys for all
eight packs.
2026-08-25 00:45:04 +08:00
Mike-Solar f7b5996032 render: fix the BGRA8 display transform + dispatcher diagnostics
CI / Build & test (Windows) (push) Failing after 17m13s
CI / Build & test (Linux) (push) Successful in 19m6s
convert_bgra8 applied packed u8 pixels to the default (F32-finalized)
OCIO CPU processor, which rejects them with a bit-depth mismatch; the
callers swallow the error, so the display-ICC transform was silently
inert on the shm preview path. ocio-rs exposes no Uint8-finalized CPU
processor, so the conversion now detours through F32. Verified against
the machine's actual display profile with the new
display_icc_bgra8_never_outputs_black test (OAK_DISPLAY_ICC).

Also:
- procpool_integration: audio tickets are Seek priority and claimable
  by any worker now, so the shard-spread assertion goes (rendering on a
  live worker is what matters).
- OAK_DEBUG_VIEWER=1: the program viewer logs frame pushes and dumps
  the displayed frame to /tmp/oak_viewer_frame.ppm (the black-screen
  investigation tooling).
2026-08-24 23:26:09 +08:00
Mike-Solar 30ef02803d render: fix the interactive-seek deadlock and seek starvation
Three compounding bugs froze the UI when dragging the playhead after
playback:

1. Self-deadlock on preview_windows: supply_preview_window /
   cancel_preview_windows / cancel_preview_window called
   cancel_preview_sequence / cancel_preview_frame while HOLDING the
   preview_windows mutex; those calls fire completions synchronously and
   the completion locks preview_windows again. Caught by sampling the
   hung process: UI thread in cancel_preview_sequence -> TicketSlot::
   finish -> completion -> Mutex::lock. Cancels/releases are now
   collected under the lock and fired after it is dropped.

2. Seek starvation by shard pinning: a Seek request's scheduler frame
   is its ticket id, pinning it to worker (id mod W). The playback
   window fills every worker's slots (window slots are only released by
   UI-thread consumption), so the seek's pinned worker could have zero
   free slots while the UI thread blocked on the seek — permanent
   starvation. Seeks (interactive frame / real-time audio) are now
   claimable by ANY worker; the no-stealing shard rule stays for
   Playback frames (adjacent frames finish together).

3. No per-worker reserve: the global preview_window_capacity reserve is
   pool-wide accounting, but exhaustion happens per worker. Playback /
   Background claims now leave one credit unused per worker; Seek
   claims may use the last slot (they complete on the worker without
   UI involvement).

Also: RealEngine::drop cancels the preview windows — ShmFrameRef has no
self-release, so every dropped engine leaked its window's slots from
the shared pool, starving later windows (surfaced as the full-suite
playback_window_supplies_playhead_frames failure once the new probe
test shifted the test schedule). new_sequence_has_default_two_video_
two_audio_tracks now takes the engine test lock (it asserts on the
global undo stack; running lock-free raced parallel undo histories).

New regression probe interactive_seek_renders_without_hanging: play 30
ticks (window fills and holds shm slots), pause, seek, synchronously
render — must not hang. Scheduler tests updated for the reserve and
seek-any-worker contract. OAK_DEBUG_DISPATCH=1 enables the dispatcher
starvation/pool diagnostics used to track this down.
2026-08-24 02:11:00 +08:00
Mike-Solar 244d5e860f workspace: kebab-case crates, app under crates/oak-app, shared versions
CI / Build & test (Windows) (push) Failing after 7s
All crates take the oak-* kebab-case naming (oak-audio, oak-codec,
oak-common, oak-core, oak-ffmpeg-link, oak-node, oak-otio, oak-plugin,
oak-render, oak-storage, oak-task, oak-timeline, oak-undo), with the
lib identifiers rewritten (oakrender:: -> oak_render::, oakcore_rs:: ->
oak_core::, ...) across all 226 referencing files.

The GUI application moves from the workspace root into
crates/oak-app/: src/, build.rs (paths fixed for the new location) and
tests/ travel with it, the root Cargo.toml becomes workspace-only
([workspace] + workspace.package + profiles), and the app package
inherits the workspace version. The screenshots example becomes a
standalone crate examples/simple_player/ with its own Cargo.toml.

Every crate now inherits the single workspace version
(version.workspace = true), and the workflows' crate paths and the
build docs follow the renames.

Validated with a clean cargo check --workspace.
2026-08-22 16:58:37 +08:00