- Autocache range jobs now post at Background priority
(submit_video_background): they used to go through the Seek path and,
after the M4 seek over-admission, jumped ahead of playback and past the
render-queue bound. The interactive single-frame preview keeps Seek.
- Job.cancelled: the arena installs the slot's cancel atom, and
execute_job finishes a cancelled job with Error::State before running
the producer — a cancel no longer burns a full render/GPU pass only to
discard the result. Exactly-once delivery is unchanged.
- DECODE_LRU_CAP 8 -> 2: the decode service's LRU is a hand-off buffer,
not the cache of record (the eval-side decoded_frames LRU is); the
double-cache footprint at 1080p F32 drops by ~6 frames. A hand-off miss
is served from the eval cache without a new decode.
- Tests: sequence-aware preview cancel, over-admitted seek ordering,
deterministic prefetch LRU reuse, cancelled-job skip, autocache
priority. docs §3.4 backfilled with the A/B/C audit outcomes.
Adjustment layers (docs/zh/plans/adjustment-layers-and-transitions.md):
a new timeline block type whose effect chain grades the composite of
every video track below it, over its own range (spanning clips or a
slice of one). The graph path flushes the lower tracks at the block's
track boundary and sweeps the composite through the chain via a
transient texture-source node; the montage path mirrors it with
AdjustmentSpan tickets (wire-compatible), so worker previews and
exports agree. An empty-area context menu creates one; the block
trims/moves/deletes like a clip, with undo everywhere.
Transitions: seam blocks come alive - cross dissolve/fade/wipe/slide
evaluate both neighbors through the graph path with progress from the
transition's own range (never the whole clip). Ctrl+Shift+D or the clip
menu inserts a default transition; the gpui wedges render and drag to
resize offsets undoably, and TransitionRemoveCommand now restores
offsets and edges on undo. The transitionfx node form runs the same
shaders on an adjustment layer with progress_in auto-filled from the
layer's span (explicit value wins).
Also: every built-in effect name and parameter name is now
translatable (360 node.* keys per locale, zh-CN fully translated, two
coverage tests guard future gaps); the new nodes register in
nodes/mod.rs with the factory smoke table updated; textfootage and
adjustment-layer i18n keys included.
- open_hw_accel marks the device unavailable when the decoder OPEN fails
(cuvidCreateDecoder OOM at 4K) too, not just device-context creation:
without it every subsequent decoder session retried CUDA and flooded
the log per open.
- Decoder gains hardware_decoding(); the oak-render decode-session LRU
evicts hardware sessions first (each pins a GPU surface pool — ~100 MB
at 4K), so a full cache cannot exhaust video memory before the next
open.
- Worker pool count now factors GPU vram: per-worker budget = 1 GiB
(1080p peak) scaled by pixel ratio + 256 MiB idle floor, 10% reserve
of free vram; applied when hardware decoding is on (nvidia-smi query,
None otherwise falls back to the RAM/CPU policy).
- Dynamic pool resize: ProcessDispatcher::set_target_workers grows or
retires workers; retiring ones stop claiming, drain their in-flight
batch (future playback frames included), then exit naturally on the
shutdown signal — no mid-work kill (30 s deadline only as a hung-
decoder last resort). A retiring worker that dies re-queues its frames
to surviving workers. Resizes are throttled to 2 s (a resolution burst
merges; only the latest target applies) so 1080p<->4K flaps cannot
thrash process spawns.
- RenderManager::set_workspace_size announces the sequence resolution;
RealEngine calls it from refresh_sequence_info.
- Integration test: shrink 3->1 mid-wave (all frames complete, retired
workers exit naturally) then regrow 1->3 and render a fresh wave.
Preview now follows the project output colorspace end to end: the
display chain derives its content space from the project's OutputColorSpec
instead of a hardcoded sRGB name, self-managed ICC transforms go through
an XYZ D65 interchange stage (OCIO cie_xyz_d65_interchange) for non-sRGB
targets, and the platform layer declares the content colorspace (gpui
submodule bump). macOS defaults to OS-managed (fixes wide-gamut UI
oversaturation); Windows ACM warns once on non-sRGB targets.
Multi-monitor: the display ICC is looked up per the window's current
screen (macOS display id, Windows per-monitor DC, X11 RandR output
profile) with a throttled poll that invalidates frame caches on moves.
Pipeline precision: 10-bit+ sources fall back to YUV444P16LE + a Rust
matrix conversion when swscale lacks F32 output (no more 8-bit
truncation); BT.709/2020 SDR decodes with BT.1886 gamma 2.4 instead of
the sRGB EOTF; working-space compositing no longer clamps RGB to [0,1]
(alpha still clamped); the output node clamps to the target gamut;
frames without colorimetry metadata convert with BT.709 defaults
(warned once) instead of passing through; scopes read the
output-colorspace signal on both F32 paths.
Also: only emit rerun-if-changed for .env when it exists (a missing file
made every build fully dirty).
Adding an effect to a clip did nothing: the sequence render is
flattened into a montage (decode + composite), and MontageClip carried
no effect data at all.
- MontageClip gains an ordered effect stack (type id / enabled /
effect input / parameter values); protocol v2 carries it as an
additive wire field (older peers default to an empty stack).
- renderops::video_montage fills the stack from the effect chain
(the footage source node — the chain end without an effect input —
is dropped; the montage decodes the footage itself). Export
(oak-task) and the multicam single-track montage fill it too.
- The worker applies the stack between decode and composite: built-in
Opacity gets a CPU evaluator (C++ opacity.frag parity — whole vec4,
alpha included, unity pass-through); everything else dispatches as an
OFX plugin job through a new instance-factory slot (oak-plugin
lazily creates + caches one instance per identifier per render
process) with the montage's parameters injected. Disabled effects
bypass (the C++ traverser pushes the effect input through). Unknown
types warn once per type id and pass through — no silent no-ops.
Not covered (explicitly): Transform/Crop and the other ~30 built-in
effects have no CPU evaluator in oak-render (they pass through with a
warning), keyframed parameter animation, audio effect chains, and the
CLI's simplified montage.
Acceptance: a real 50% Opacity on real media quarters the rendered
pixels both in-process (renderops test) and through a real worker
process over IPC + shared memory (procpool_integration test);
disabling restores the plain render byte-for-byte.
convert_bgra8 applied packed u8 pixels to the default (F32-finalized)
OCIO CPU processor, which rejects them with a bit-depth mismatch; the
callers swallow the error, so the display-ICC transform was silently
inert on the shm preview path. ocio-rs exposes no Uint8-finalized CPU
processor, so the conversion now detours through F32. Verified against
the machine's actual display profile with the new
display_icc_bgra8_never_outputs_black test (OAK_DISPLAY_ICC).
Also:
- procpool_integration: audio tickets are Seek priority and claimable
by any worker now, so the shard-spread assertion goes (rendering on a
live worker is what matters).
- OAK_DEBUG_VIEWER=1: the program viewer logs frame pushes and dumps
the displayed frame to /tmp/oak_viewer_frame.ppm (the black-screen
investigation tooling).
ci (Windows): the runner's msys2 shell starts as the base MSYS
environment (MSYSTEM=MSYS), which install-deps.sh rejects. Set the
job-level MSYSTEM=UCRT64 env and prepend /ucrt64/bin to PATH in every
Windows step, so pacman installs and the toolchain resolve against the
mingw-w64-ucrt-x86_64 packages.
fixtures: the restructured workspace moved the media fixtures into
crates/oak-app/tests; the remaining references pointed at the old
repo-root tests/ — oak-codec realmedia_tests + hwdecode, oak-node
serializer golden, oak-worker procpool integration. All point at
../oak-app/tests now (oak-app's own tests/demo.mp4 references were
already correct after the move).
All crates take the oak-* kebab-case naming (oak-audio, oak-codec,
oak-common, oak-core, oak-ffmpeg-link, oak-node, oak-otio, oak-plugin,
oak-render, oak-storage, oak-task, oak-timeline, oak-undo), with the
lib identifiers rewritten (oakrender:: -> oak_render::, oakcore_rs:: ->
oak_core::, ...) across all 226 referencing files.
The GUI application moves from the workspace root into
crates/oak-app/: src/, build.rs (paths fixed for the new location) and
tests/ travel with it, the root Cargo.toml becomes workspace-only
([workspace] + workspace.package + profiles), and the app package
inherits the workspace version. The screenshots example becomes a
standalone crate examples/simple_player/ with its own Cargo.toml.
Every crate now inherits the single workspace version
(version.workspace = true), and the workflows' crate paths and the
build docs follow the renames.
Validated with a clean cargo check --workspace.
A claim mixing audio and video tickets is delivered as the video
message first and the audio message second, and the worker pops one
free-ring slot per ticket in that message order, checking each pop
against the assignment. The dispatcher however assigned slots in the
scheduler's interleaved frame order, so every audio ticket inside a
mixed batch mismatched, and each mismatch consumed a worker slot
without recycling it — cascading into the 'slot assignment mismatch'
flood and failed frames during playback.
Slot assignment now partitions the claim: video tickets first, then
audio. The mixed_audio_video integration test forces mixed claims
(queue depth > slot count with immediate releases) and fails with the
exact production signature when the fix is reverted.
- Audio tickets join the process backend: render_audio_batch wire
message, workers mix straight into shm slots (SLOT_FORMAT_AUDIO_F32),
ShmAudio payload with release semantics, crash isolation covers audio
renders; playback audio uses an async 4-chunk prefetch drained on the
UI tick (also fixes the sub-60fps chunk truncation bug); oversized
ranges and dispatcher outages fall back to in-process inline.
- Per-ticket slot formats: force_format is honored (exports request
F32 slots, dropping the BGRA8 round-trip and its 8-bit quantization);
segments grow on demand via worker-idle rebuild with generation
handoff; the scheduler filters over-capacity tickets.
- Adaptive defaults: 128-256MB/worker segment budgets drive slots per
worker, batch size follows workers/slots; bench_process example
measures throughput and adjacent-frame completion deltas
(e.g. 4 workers: 841 fps, 4.6ms mean delta).
- WorkerPool thread pool deleted; RenderManager defaults to the
Processes backend (oak-worker children), Threads kept as a test-only
inline dispatcher; audio tickets stay in-process until S3.
- Onscreen path reads worker shm slots directly: BGRA8 slot format,
RenderedFrame::Shm wrapped into the display buffer (single disclosed
GPU-staging memcpy), scopes analyze BGRA8; the long-lived full-res /
thumbnail paths take the counted slot_to_vec copy and release.
- Playback pre-render window: forward 120 frames (configurable) fed to
the PreviewScheduler at Playback priority, interleaved across
workers, cached in shm slots until the playhead consumes them;
generation-based invalidation cancels and releases on edits.
- oaktask export and oak-cli run on private ProcessDispatchers (fixed
a pump-while-locked self-deadlock in the export loop); facade
get_frame handles ShmFrame payloads.
- Acceptance: preview path main_heap_frame_copies == 0 with spawned
workers, CLI transcode/render verified end to end.