- Mark the raw-pointer interop entry points unsafe with # Safety docs
(oak-core upload/download/frame-from-pixels, oak-audio convert) and
satisfy the existing callers (tests).
- mut_from_ref: allow with the ABI contract documented (the handle
get_mut helpers in oak-timeline/oak-render/oak-task take the shared
reference the C ABI passes; exclusivity is the caller's unsafe
contract).
- Fix the eq_op in the white-balance normalization (green / green).
- Apply cargo clippy --fix across the workspace (redundant closures and
field names, field reassignment, items after test modules, ...).
- Revert the replace_box fix in image_effect's clip_define: a
redefinition must allocate a new box, otherwise the old clip handle
stays valid and the HS-map replace contract (clip != clip2) breaks.
- 283 warnings remain; they are all non-machine-applicable
(chunks_exact -> as_chunks needs a manual iter_mut, too_many_arguments,
complex types, missing Safety docs, ...) and are tracked as the
follow-up.
- Autocache range jobs now post at Background priority
(submit_video_background): they used to go through the Seek path and,
after the M4 seek over-admission, jumped ahead of playback and past the
render-queue bound. The interactive single-frame preview keeps Seek.
- Job.cancelled: the arena installs the slot's cancel atom, and
execute_job finishes a cancelled job with Error::State before running
the producer — a cancel no longer burns a full render/GPU pass only to
discard the result. Exactly-once delivery is unchanged.
- DECODE_LRU_CAP 8 -> 2: the decode service's LRU is a hand-off buffer,
not the cache of record (the eval-side decoded_frames LRU is); the
double-cache footprint at 1080p F32 drops by ~6 frames. A hand-off miss
is served from the eval cache without a new decode.
- Tests: sequence-aware preview cancel, over-admitted seek ordering,
deterministic prefetch LRU reuse, cancelled-job skip, autocache
priority. docs §3.4 backfilled with the A/B/C audit outcomes.
docs/zh/plans/render-pipeline-threads.md M4: the thread pipeline now
keeps its decode thread ahead of the render thread and the app's
playback window consumes in-process frames.
- Render queue: priority-ordered by JobSchedule.priority (Seek >
Playback > Background, FIFO within a class), so interactive frames
jump playback exports/autocache. Seek posts may over-admit the bound:
priority only reorders queued jobs, so a full queue of background work
must not park the UI thread until an export frame finishes.
- Decode queue: rendezvous Requests are served ahead of queued
Prefetches (a frame the renderer needs never waits behind speculative
decodes); Sync barriers stay FIFO. The queue is a bounded
Mutex+Condvar structure, preserving the request backpressure and the
wait_idle contract.
- Playback read-ahead: a Playback job's footage decode requests are
derived from its montage/footage spec on post (same media time, size
and force_format.unwrap_or(F32) as the eval) and queued immediately,
so frame N+1 decodes while frame N runs its GPU passes.
- App window: PreviewWindow slots are generalized to
PreviewSlot::{Shm, Video}; the pipeline's in-process TicketPayload is
cached and consumed by cpu_frame exactly like a worker slot.
PipelineBackend::preview_window_capacity reports the render-queue
headroom, so playback posts are capped to what the queue can take;
cancel_preview_frame drops queued frames the playhead has passed,
matched on the full (sequence, frame, version) key so one monitor's
window never drops the other sequence's same-numbered frame.
- Tests: decode-queue preemption/FIFO, render-queue ordering, request
derivation, and deterministic end-to-end M4 tests: a prefetch that
must be reused by the render request (LRU hit, single decode — the
read-ahead claim is falsifiable), a parked-render-thread priority test
where a full queue of background work still lets a Seek over-admit and
run first, and a sequence-aware cancel test. The playback prefetch
smoke asserts prefetches == distinct decodes == frames; it does not
claim zero heap copies (Frame.data is deep-copied at the eval-cache
and service-LRU boundaries today).
- bench_playback gains a pipeline mode with CPU (self+children) and
first-frame latency; both backends now produce F32 frames so the
comparison is like-for-like. The §3.4 backfill records the numbers:
at the proxy size the pipeline is faster with a lower first frame; at
1080p peak throughput is below the multi-worker pool, but that is an
artifact of the decode still being CPU software (M5), not a case for
pooling decode threads — GPU decode is a single device/queue and the
zero-copy import shares one GPU memory pool, so the single decode
thread stays the target shape.
docs/zh/plans/render-pipeline-threads.md M2: the graph's textures stay
on the GPU from evaluation through presentation, and presentation runs
on the UI's own wgpu device.
- wgpu 25 -> 29 (naga 29) across the engine, unifying it with
gpui_wgpu so engine textures are directly sampleable by the presenter
(a single wgpu remains in the lockfile).
- GpuContext::adopt/install_shared: the app registers the window's
device at startup and the render thread renders on it;
texture_handle hands the raw Arc<wgpu::Texture> to
SurfaceSource::Texture - zero-copy present on Linux/FreeBSD. The
shared slot replaces an engine context that has not touched the GPU
yet (startup-order guard) and refuses once it has.
- Texture::Gpu shares a GpuLease so clones release the registry token
exactly once; the compositor, transitions and adjustment sweeps keep
GPU textures end to end (no per-clip readbacks; GPU clears for
black/generated frames).
- Color management stays on the GPU: the output node + display ICC
chain is baked into a 65^3 3D LUT with the exact CPU reference and
applied by the present WGSL pass (manual trilinear);
ColorTransformJob bakes its OCIO processor the same way. Neither
path skips color management.
- The explicit readback boundaries accept GPU textures: export
encoder, CLI, worker shm, disk cache; CPU OpenFX already read back.
- M5 dependency: the YUV->RGB GPU pass (BT.601/709/2020 x
limited/full) matches colormath::yuv444p16_to_rgb_f32.
- Acceptance: gpu_transfer_counters; single-clip and layered
(multi-track + transition + adjustment) playback tests assert zero
GPU->CPU readbacks, and the app test asserts adopted-device present
is zero-copy. GPU tests hard-fail when OAK_REQUIRE_GPU is set (CI
lavapipe) instead of skipping silently.
docs/zh/plans/render-pipeline-threads.md M1: an in-process
alternative to the worker-process pool, behind OAK_PIPELINE=threads
(processes stays the default and is fully retained).
- pipeline.rs: PipelineBackend implements JobDispatch over a single
render thread draining a bounded FIFO (cap 8; blocking post with
condvar backpressure and a one-ahead exception for re-posts from
the render thread itself; shutdown drains with Error::State like
the inline dispatcher). The DecodeService is a single decode
thread behind a bounded command queue with a real LRU (tick-based
eviction), rendezvous requests (None on shutdown -> the caller
decodes inline), prefetch gated on render-queue room, and a Sync
barrier; it installs into a process-wide slot that eval's footage
path consults per frame (no service -> the synchronous decode it
always was).
- The manager gains RenderBackendChoice::Pipeline; init() reads
OAK_PIPELINE (threads -> pipeline, anything else -> the process
pool), audio stays deliberately inline.
- Present mapping: the UI thread consumes through the ticket
completion, unchanged — no fourth thread is invented.
- Tests: decode-service unit tests (rendezvous, LRU hit/eviction,
error propagation, backpressure gate) plus a six-case integration
suite matrixed over inline vs pipeline — consecutive-frame and
out-of-order seek pixel equality asserted byte for byte, with
decode counters proving the service (not the caller) did the
codec work.
254 warnings (320 counting replayed-cache re-emitters) cleaned:
unused mut/imports/variables, irrefutable if-lets and unreachable
patterns, dead code removed or annotated #[allow(dead_code)] with
the reason (C++ parity value sets, cfg(test) helpers, public API
reservations), drop(&ref) no-ops removed, fn-pointer identity via
std::ptr::fn_addr_eq, the test-stubs feature declared in
oak-node's manifest, missing docs filled. Every unused-Result site
was judged individually: meaningful errors propagate, intentional
ignores are let _ = with a note.
Two pre-existing latent bugs are documented in place, behavior
preserved: app.rs's timeline-tool observer and dialogs.rs's format
subscription both drop the returned Subscription immediately, so
they never fire.
Generator effects (bars, checkerboard) can be dragged from the library
onto the timeline, where they land as a standalone five-second clip
built from the node factory; the inspector shows the generator's
parameters as the clip's own chain.
Transitions are no longer junction-only. The render planner accepts a
transition with at least one wired neighbor and blends the missing
side against transparent black, so head transitions fade in from black
and tail transitions fade out to black. add_transition_at_edge creates
those single-sided blocks (wired to just the IN or OUT block), the
default-transition command covers both ends of a lone clip, and an
effect drag dropped near a clip edge routes to the nearest seam or
edge within a one-second window.
Translating, rotating or warping content past the frame edge used to
smear the clamped edge row/column across the vacated region. The
transform, position, swirl, ripple and wave shaders now multiply the
sample by an in-bounds mask so off-frame pixels come out transparent
(and composite as black when nothing sits below). Tile deliberately
keeps its wrapping lookup.
Adjustment layers (docs/zh/plans/adjustment-layers-and-transitions.md):
a new timeline block type whose effect chain grades the composite of
every video track below it, over its own range (spanning clips or a
slice of one). The graph path flushes the lower tracks at the block's
track boundary and sweeps the composite through the chain via a
transient texture-source node; the montage path mirrors it with
AdjustmentSpan tickets (wire-compatible), so worker previews and
exports agree. An empty-area context menu creates one; the block
trims/moves/deletes like a clip, with undo everywhere.
Transitions: seam blocks come alive - cross dissolve/fade/wipe/slide
evaluate both neighbors through the graph path with progress from the
transition's own range (never the whole clip). Ctrl+Shift+D or the clip
menu inserts a default transition; the gpui wedges render and drag to
resize offsets undoably, and TransitionRemoveCommand now restores
offsets and edges on undo. The transitionfx node form runs the same
shaders on an adjustment layer with progress_in auto-filled from the
layer's span (explicit value wins).
Also: every built-in effect name and parameter name is now
translatable (360 node.* keys per locale, zh-CN fully translated, two
coverage tests guard future gaps); the new nodes register in
nodes/mod.rs with the factory smoke table updated; textfootage and
adjustment-layer i18n keys included.
Text is no longer a hand-written HTML effect (docs in
docs/zh/plans/text-footage-redesign.md):
- textv3 gains structured inputs - plain text, font family/size, font
color, outline (enable/color/width), glow (enable/color/radius); the
legacy text_in HTML is hidden and auto-migrated to plain text on load.
- Outline and glow render as GPU post-process chains (dilate/blur +
colorize under the text); the font color tints the raster
premultiplied. Plain text now rasterizes even with both passes off
(previously a null deferred job), fixing a use-after-free where the
handle was lifted out of an owning Option<NodeValue> before addref.
- A cosmic-text backend (the lockfile's 0.19) installs at engine
startup through the textbackend hooks and feeds the font-family combo.
- The project panel gains 添加文本素材 next to 新建序列: a text entry
in the bin that drops onto the timeline as a clip (one undo row), its
parameters shown as structured fields in the inspector (multiline
text area, no HTML anywhere). text3 is hidden from the effect add
menus; legacy text3 chains keep evaluating.
21 built-in effects reimplemented as native GPU nodes from the
OpenFX-Misc algorithm references (cleanroom, docs in
docs/zh/plans/ofx-misc-gpu-cleanroom.md):
- Color: Color Correct, Gamma, Saturation, Invert, Clamp, Grade
- Matrix/morphology: Color Matrix, Edge Detect, Dilate, Erode
- Blur: Directional Blur, Sharpen (unsharp mask)
- Merge: Dissolve, Key Mix, Premultiply, Unpremultiply
- Geometry/generators: Position, Mirror, Checkerboard, Color Bars, Ramp
Every node carries unit tests plus GPU pixel tests (28 cases over five
ofxmisc_* suites). The effect library groups built-ins by category
(color/filter/distort/keying/generator/math/general) with collapsible
group headers persisted to the config; the inspector's add menu groups
the same way. Registration wiring and the factory smoke table land with
the adjustment/transition wave sharing the same files.
Preview renders at proxy size while a paused frame renders full-res, so
anchoring resolution_in to the render target made every sequence-pixel
effect (shape size/pos, transform offsets, corner pin points, drop
shadow distance) change apparent size whenever the transport stopped.
Pre-fill resolution_in from the sequence's video params (C++ inserts
the NodeGlobals square resolution at job-build time), covering the
nested generator job inside a merge as well; a node that inserted its
own resolution_in keeps it.
oak-common is gone; its modules (configstore, xmlutils, ocioutils,
oiioutils, colormath, colortransform, videoparams, ffmpegutils, ...)
now live in oak-core alongside the value types. The render value/GPU
types moved too: backend (wgpu context + DisplayRenderer), color
(ColorProcessor over ocio-rs), texture, frame, and the commonutil
config helpers.
Fix-ups to make the merged tree build and pass tests:
- oak-core Cargo.toml: wgpu back to 25 (the moved backend code is
written against that API generation); add the toml/quick-xml/image
deps oak-common carried.
- lib.rs: drop the duplicate 'pub mod error;'.
- error.rs: unified OAKCORE_* codes; restore Error::new() and
From<OcioError> from oak-common's error type.
- backend.rs/color.rs: oak_core::/oak_render:: self-references
rewritten to crate::; the shaderfx-dependent GPU effect test moved
to oak-render's shaderfx tests (shaderfx depends on oak-node and
cannot live in oak-core).
- oak-render's error module re-exports oak_core::error::{Error,
Result}; the OAKRENDER_* codes stay as the public-code contract.
- oak-node jobs.rs: ColorProcessor imported from oak_core::color.
- Integration tests repointed at oak_core::{texture, frame, backend,
color, colormath}.
- the display-ICC regression test treats an empty OAK_DISPLAY_ICC as
unset, matching displayicc::env_override_icc.
- Project explorer header gains a 新建序列 button (opens the existing
new-sequence dialog, seeded like the menu action).
- Sequence presets gain 4K UHD (3840x2160@25) and 4K DCI (4096x2160@24);
the sequence-properties dialog re-selects them on reopen. Format fields
are also seedable from a probed footage format (the drop flow).
- Chroma Key (and Color Difference Key) value() now box a ShaderJobPayload
like Despill: the old OCIO-processor gate pushed nothing (the processor
is never populated without the render bridge), so the traverser handed
the clip NodeValue::None and the rendered frame lost the clip. The
renderer resolves the OCIO stub at compile time from OCIO_SHADER_STUBS.
End-to-end graph test: green key on a green frame keys out, red key
keeps it.
- The OFX color picker's SV palette / hue bar / preview / swatch canvases
get size_full(): the bare canvases collapsed to zero height in the
block layout, so the palette painted nothing (the reported 色板没显示).
Regression test clicks the palette center and expects mid s/v.
Preview now follows the project output colorspace end to end: the
display chain derives its content space from the project's OutputColorSpec
instead of a hardcoded sRGB name, self-managed ICC transforms go through
an XYZ D65 interchange stage (OCIO cie_xyz_d65_interchange) for non-sRGB
targets, and the platform layer declares the content colorspace (gpui
submodule bump). macOS defaults to OS-managed (fixes wide-gamut UI
oversaturation); Windows ACM warns once on non-sRGB targets.
Multi-monitor: the display ICC is looked up per the window's current
screen (macOS display id, Windows per-monitor DC, X11 RandR output
profile) with a throttled poll that invalidates frame caches on moves.
Pipeline precision: 10-bit+ sources fall back to YUV444P16LE + a Rust
matrix conversion when swscale lacks F32 output (no more 8-bit
truncation); BT.709/2020 SDR decodes with BT.1886 gamma 2.4 instead of
the sRGB EOTF; working-space compositing no longer clamps RGB to [0,1]
(alpha still clamped); the output node clamps to the target gamut;
frames without colorimetry metadata convert with BT.709 defaults
(warned once) instead of passing through; scopes read the
output-colorspace signal on both F32 paths.
Also: only emit rerun-if-changed for .env when it exists (a missing file
made every build fully dirty).
Match the timeline UI (V_max drawn topmost): composite tracks from
V1 up to V_max so the highest-numbered track is composited last, in
both the montage path and direct graph evaluation.
Root causes found for the 4K stalls and the second-footage memory
blowup (audit + code review):
- ticket bookkeeping leaked unbounded: the procpool ticket table and
the arena slot map only ever grew (50-100 tickets/sec during
playback, each pinning montage params and shm region views).
Completed/cancelled/superseded/crashed entries are now removed, and
the arena reaps fire-and-forget tickets once finished; the sync poll
path reaps via a terminal result() read. InFlight duplicate submits
now answer State immediately instead of sitting in the map forever.
- decode ran a full-resolution swscale to F32 RGBA (~132 MB at 4K)
plus a second full-res copy before downscaling to the 480px proxy:
RetrieveVideoParams.target_size lets swscale convert AND resize in
one pass (bilinear, matching the old Rust resampler), so a 4K
preview frame costs ~1 MB instead of ~260 MB of churn. This applies
to proxy AND full-res requests alike.
- per-process decoder cache was unbounded (each session pins an FFmpeg
context + 2 native decoded frames): LRU-capped at 16, eviction drops
the map entry (in-flight renders keep their Arc; Drop releases
FFmpeg).
- playback window completions were not generation-gated: a stale
render from before an edit landed in the rebuilt window (wrong frame
displayed, fresh request blocked). Stale completions now return
their shm slot credit instead.
- async audio prefetch used the polling ticket submit without ever
polling: switched to the fire-and-forget submit so entries reap.
All crates take the oak-* kebab-case naming (oak-audio, oak-codec,
oak-common, oak-core, oak-ffmpeg-link, oak-node, oak-otio, oak-plugin,
oak-render, oak-storage, oak-task, oak-timeline, oak-undo), with the
lib identifiers rewritten (oakrender:: -> oak_render::, oakcore_rs:: ->
oak_core::, ...) across all 226 referencing files.
The GUI application moves from the workspace root into
crates/oak-app/: src/, build.rs (paths fixed for the new location) and
tests/ travel with it, the root Cargo.toml becomes workspace-only
([workspace] + workspace.package + profiles), and the app package
inherits the workspace version. The screenshots example becomes a
standalone crate examples/simple_player/ with its own Cargo.toml.
Every crate now inherits the single workspace version
(version.workspace = true), and the workflows' crate paths and the
build docs follow the renames.
Validated with a clean cargo check --workspace.