- process_shader_job: bind all texture params by name, recurse into nested
shader payloads (depth cap 8), fall back to frame size without inputs
- run_effect: take iterative_input so dropshadow previous_iteration_in works
- merge: actually composite inputs; keyer mask, opacity modulation, math
texture ops and mrg generator layers now bind their textures
- transform distort: real fragment-side inverse-matrix sampling
- time offset / time remap: wire NodeBehavior time adjustment hooks
- plugin: fix first-node identity colliding with unbound sentinel
oak-common is gone; its modules (configstore, xmlutils, ocioutils,
oiioutils, colormath, colortransform, videoparams, ffmpegutils, ...)
now live in oak-core alongside the value types. The render value/GPU
types moved too: backend (wgpu context + DisplayRenderer), color
(ColorProcessor over ocio-rs), texture, frame, and the commonutil
config helpers.
Fix-ups to make the merged tree build and pass tests:
- oak-core Cargo.toml: wgpu back to 25 (the moved backend code is
written against that API generation); add the toml/quick-xml/image
deps oak-common carried.
- lib.rs: drop the duplicate 'pub mod error;'.
- error.rs: unified OAKCORE_* codes; restore Error::new() and
From<OcioError> from oak-common's error type.
- backend.rs/color.rs: oak_core::/oak_render:: self-references
rewritten to crate::; the shaderfx-dependent GPU effect test moved
to oak-render's shaderfx tests (shaderfx depends on oak-node and
cannot live in oak-core).
- oak-render's error module re-exports oak_core::error::{Error,
Result}; the OAKRENDER_* codes stay as the public-code contract.
- oak-node jobs.rs: ColorProcessor imported from oak_core::color.
- Integration tests repointed at oak_core::{texture, frame, backend,
color, colormath}.
- the display-ICC regression test treats an empty OAK_DISPLAY_ICC as
unset, matching displayicc::env_override_icc.
- open_hw_accel marks the device unavailable when the decoder OPEN fails
(cuvidCreateDecoder OOM at 4K) too, not just device-context creation:
without it every subsequent decoder session retried CUDA and flooded
the log per open.
- Decoder gains hardware_decoding(); the oak-render decode-session LRU
evicts hardware sessions first (each pins a GPU surface pool — ~100 MB
at 4K), so a full cache cannot exhaust video memory before the next
open.
- Worker pool count now factors GPU vram: per-worker budget = 1 GiB
(1080p peak) scaled by pixel ratio + 256 MiB idle floor, 10% reserve
of free vram; applied when hardware decoding is on (nvidia-smi query,
None otherwise falls back to the RAM/CPU policy).
- Dynamic pool resize: ProcessDispatcher::set_target_workers grows or
retires workers; retiring ones stop claiming, drain their in-flight
batch (future playback frames included), then exit naturally on the
shutdown signal — no mid-work kill (30 s deadline only as a hung-
decoder last resort). A retiring worker that dies re-queues its frames
to surviving workers. Resizes are throttled to 2 s (a resolution burst
merges; only the latest target applies) so 1080p<->4K flaps cannot
thrash process spawns.
- RenderManager::set_workspace_size announces the sequence resolution;
RealEngine calls it from refresh_sequence_info.
- Integration test: shrink 3->1 mid-wave (all frames complete, retired
workers exit naturally) then regrow 1->3 and render a fresh wave.
On multi-GPU boxes VAAPI's default render node can point at a device
with no VA driver (NVIDIA without libva-nvidia-driver), failing before
the working AMD/Intel node is ever tried; trying CUDA (NVDEC) first is
both the discrete-GPU path and sidesteps that misdirection. VAAPI stays
as the fallback for AMD/Intel-only machines.
Adds a hwcheck example printing which hw device a stream opens with and
the raw av_hwdevice_ctx_create result codes per device type.
threads=auto spawns one decode thread per logical core per decoder (16
observed on a 24-core box); with cores-2 render workers decoding
concurrently the machine is oversubscribed several times over and the
playback-audio render thread gets starved (late chunks are dropped).
4 threads decode 1080p h264 far past real-time; preview throughput comes
from the worker pool, not per-decoder threading. Applies to both the
software open path and the hwaccel open path.
retrieve_audio_to decodes whole codec frames but copies only the part
inside the requested chunk; the tail of the frame crossing the chunk end
(up to 1023 samples for AAC) was consumed by the decoder and lost, so
the next chunk started with a hole. On the playback grid (1920 samples
at 25fps/48kHz) the hole cycles 128..896 samples and hits 7 of 8 chunk
starts -- the heavy stutter/noise heard during playback.
Keep the overflow (decode and resampler-flush tails) in a per-session
carry buffer and serve it at the start of the next contiguous chunk;
clear it on seek, format change and chunk failure. Also make seek()
actually drop the cached resampler as its comment claimed.
Verified sample-exact: chunked decode of a 440Hz tone now matches a
one-shot decode bit for bit, and chunked renders of real media line up
with the ffmpeg CLI reference at correlation 1.0 / drift 0.
Adds a regression test (playback_sized_chunks_match_oneshot_sample_exact)
with a new tone fixture -- demo.mp4's audio is -90dB digital silence and
cannot expose the holes -- plus a render_audio_wav example used to
diagnose chunk-boundary artifacts offline.
Preview now follows the project output colorspace end to end: the
display chain derives its content space from the project's OutputColorSpec
instead of a hardcoded sRGB name, self-managed ICC transforms go through
an XYZ D65 interchange stage (OCIO cie_xyz_d65_interchange) for non-sRGB
targets, and the platform layer declares the content colorspace (gpui
submodule bump). macOS defaults to OS-managed (fixes wide-gamut UI
oversaturation); Windows ACM warns once on non-sRGB targets.
Multi-monitor: the display ICC is looked up per the window's current
screen (macOS display id, Windows per-monitor DC, X11 RandR output
profile) with a throttled poll that invalidates frame caches on moves.
Pipeline precision: 10-bit+ sources fall back to YUV444P16LE + a Rust
matrix conversion when swscale lacks F32 output (no more 8-bit
truncation); BT.709/2020 SDR decodes with BT.1886 gamma 2.4 instead of
the sRGB EOTF; working-space compositing no longer clamps RGB to [0,1]
(alpha still clamped); the output node clamps to the target gamut;
frames without colorimetry metadata convert with BT.709 defaults
(warned once) instead of passing through; scopes read the
output-colorspace signal on both F32 paths.
Also: only emit rerun-if-changed for .env when it exists (a missing file
made every build fully dirty).
mov/mp4 muxers rewrite the stream time base at write_header (mov timescale
>= 10000) and FFmpeg 9 no longer rescales packets for us, so mpeg2video
clips were muxed with pts in encoder ticks -- a 10-frame/10fps clip became
1ms long and every seek past t=0 decoded to the EOF frame.
- read back the real stream time base after write_header and rescale
packets (including the flush path) before write_interleaved
- warn once when retrieve_frame's EOF fallback returns a frame far from
the requested timestamp
- testmedia round-trip now asserts container duration and per-frame
decode instead of only t=0
Match the timeline UI (V_max drawn topmost): composite tracks from
V1 up to V_max so the highest-numbered track is composited last, in
both the montage path and direct graph evaluation.
Root causes found for the 4K stalls and the second-footage memory
blowup (audit + code review):
- ticket bookkeeping leaked unbounded: the procpool ticket table and
the arena slot map only ever grew (50-100 tickets/sec during
playback, each pinning montage params and shm region views).
Completed/cancelled/superseded/crashed entries are now removed, and
the arena reaps fire-and-forget tickets once finished; the sync poll
path reaps via a terminal result() read. InFlight duplicate submits
now answer State immediately instead of sitting in the map forever.
- decode ran a full-resolution swscale to F32 RGBA (~132 MB at 4K)
plus a second full-res copy before downscaling to the 480px proxy:
RetrieveVideoParams.target_size lets swscale convert AND resize in
one pass (bilinear, matching the old Rust resampler), so a 4K
preview frame costs ~1 MB instead of ~260 MB of churn. This applies
to proxy AND full-res requests alike.
- per-process decoder cache was unbounded (each session pins an FFmpeg
context + 2 native decoded frames): LRU-capped at 16, eviction drops
the map entry (in-flight renders keep their Arc; Drop releases
FFmpeg).
- playback window completions were not generation-gated: a stale
render from before an edit landed in the rebuilt window (wrong frame
displayed, fresh request blocked). Stale completions now return
their shm slot credit instead.
- async audio prefetch used the polling ticket submit without ever
polling: switched to the fire-and-forget submit so entries reap.
- docs/zh/plans: finished plans move to completed/ (the RIIR series, the
event-bridge and dependency plans, the v04 manual test plan).
- New design docs: the external (functional) plugin system
(process-isolated, JSON-RPC/shm) and its protocol.
- ai-agent-design refreshed; README pointers follow the moves.
ci (Windows): the runner's msys2 shell starts as the base MSYS
environment (MSYSTEM=MSYS), which install-deps.sh rejects. Set the
job-level MSYSTEM=UCRT64 env and prepend /ucrt64/bin to PATH in every
Windows step, so pacman installs and the toolchain resolve against the
mingw-w64-ucrt-x86_64 packages.
fixtures: the restructured workspace moved the media fixtures into
crates/oak-app/tests; the remaining references pointed at the old
repo-root tests/ — oak-codec realmedia_tests + hwdecode, oak-node
serializer golden, oak-worker procpool integration. All point at
../oak-app/tests now (oak-app's own tests/demo.mp4 references were
already correct after the move).
All crates take the oak-* kebab-case naming (oak-audio, oak-codec,
oak-common, oak-core, oak-ffmpeg-link, oak-node, oak-otio, oak-plugin,
oak-render, oak-storage, oak-task, oak-timeline, oak-undo), with the
lib identifiers rewritten (oakrender:: -> oak_render::, oakcore_rs:: ->
oak_core::, ...) across all 226 referencing files.
The GUI application moves from the workspace root into
crates/oak-app/: src/, build.rs (paths fixed for the new location) and
tests/ travel with it, the root Cargo.toml becomes workspace-only
([workspace] + workspace.package + profiles), and the app package
inherits the workspace version. The screenshots example becomes a
standalone crate examples/simple_player/ with its own Cargo.toml.
Every crate now inherits the single workspace version
(version.workspace = true), and the workflows' crate paths and the
build docs follow the renames.
Validated with a clean cargo check --workspace.