Commit Graph
55 Commits
Author SHA1 Message Date
Mike-Solar 9b3a031d58 test(oak-render): probe true in /usr/bin for the dispatcher tests
macOS ships true in /usr/bin (there is no /bin/true), so the immediate
exit worker spawn failed with ENOENT and
start_failure_restart_budget_and_accessors never reached a restart.
Resolve the binary from the usual locations, falling back to PATH.
2026-09-24 16:22:39 +08:00
Mike-Solar 61cb4b71ed fix(oak-render): fold over-long POSIX shm names for macOS
macOS caps shm_open names at PSHMNAMLEN (31 bytes including the leading
slash) and answers ENAMETOOLONG beyond that, while Linux allows 255. The
dispatcher tests name segments oak-procpool-ut-<pid>-<label>, so every
procpool test failed on macOS as soon as the earlier crash stopped
masking them. Fold longer keys into a deterministic FNV-1a short name in
one place; owner, worker and unlink all derive the name through it, so
they keep meeting on the same segment.
2026-09-24 16:22:39 +08:00
Mike-Solar d0a8fcd2a5 test(oak-render): make frame-path and worker-bin tests platform-neutral
Windows CI exposed two assertions that assumed POSIX paths:
frame_filename was checked with ends_with("/450") (separator is '\\'
there) and the OAK_WORKER_BIN override used /bin/sh, which never exists
on Windows so the sibling probe legitimately won. Compare the last path
component and point the override at the test executable instead.
2026-09-24 16:11:13 +08:00
Mike-Solar c50d489dc0 fix(oak-render): composite the CPU track stack bottom-up
The GPU path composites frames bottom (last) to top (first); the CPU
fallback iterated top-first, so on machines without a working adapter
every multi-layer frame had its layering order inverted (the
composite_tracks test caught it as 0.8125 vs the documented 0.625).
Factor the CPU half into composite_tracks_cpu, iterate it in reverse
and pin the math in the test by calling the CPU path directly (the old
assertion silently exercised the GPU path whenever another test had
installed a shared context).
2026-09-24 16:11:10 +08:00
Mike-Solar b13ef477a8 fix(tests): make the platform-specific suites portable and deterministic
Four independent CI failures the first real cross-platform run surfaced:

- macOS SIGSEGV: the real-GL unit tests in gl_bridge ran wherever CGL is
  available, including the headless CI runner. Gate them with the same
  OAK_GPU_TESTS switch the integration GL tests already use (skip on CI,
  opt in on a real Mac).
- Windows build: examples compile under `cargo test`, and bench_playback
  used libc::getrusage unconditionally. Keep the Unix CPU accounting
  behind #[cfg(unix)] and report zero CPU seconds elsewhere.
- openKylin arm64: engine_without_a_project_hits_the_guard_paths assumed
  the library backend was unconfigured while a parallel config test
  transiently set Storage/Backend=sqlite. Take the shared config lock and
  pin the key off for the test's duration.
- openKylin x64: the prefetch smoke test asserted an exact decode count,
  but the hand-off LRU holds only DECODE_LRU_CAP (2) frames, so a request
  can miss the prefetched copy under scheduling pressure and re-run the
  producer (the eval cache still serves the pixels). Bound the count
  instead of pinning it; the deterministic sibling test pins read-ahead
  usage.
2026-09-24 14:45:12 +08:00
Mike-Solar 666ac9b4d4 test(oak-render, oak-worker): eval, procpool and worker coverage
Render evaluation fallbacks, the process pool (dispatch, cancel,
restart, teardown), half-float display packing, and the worker's
shared-memory job paths; includes the M5 footage import acceptance
tests and the software-decode byte-exactness guard.
2026-09-22 20:54:04 +08:00
Mike-Solar f188bc79e7 feat(gpu-decode): zero-copy hardware imports and the planar pipeline
Adds VAAPI DMA-BUF, D3D11VA shared-handle and VideoToolbox IOSurface
imports behind a tri-state outcome (imported / unsupported / failed),
planar textures with bounded residency and a CPU staging fallback, the
staged montage decode path, reference-counted decoder frames, VAAPI-first
device selection on Linux, and the host-GPU context plumbing used by the
app and worker. See docs/zh/plans/render-pipeline-threads.md (M5).
2026-09-22 20:54:03 +08:00
Mike-Solar ee7ea18d94 clippy: clear the workspace errors and apply the machine fixes
- Mark the raw-pointer interop entry points unsafe with # Safety docs
  (oak-core upload/download/frame-from-pixels, oak-audio convert) and
  satisfy the existing callers (tests).
- mut_from_ref: allow with the ABI contract documented (the handle
  get_mut helpers in oak-timeline/oak-render/oak-task take the shared
  reference the C ABI passes; exclusivity is the caller's unsafe
  contract).
- Fix the eq_op in the white-balance normalization (green / green).
- Apply cargo clippy --fix across the workspace (redundant closures and
  field names, field reassignment, items after test modules, ...).
- Revert the replace_box fix in image_effect's clip_define: a
  redefinition must allocate a new box, otherwise the old clip handle
  stays valid and the HS-map replace contract (clip != clip2) breaks.
- 283 warnings remain; they are all non-machine-applicable
  (chunks_exact -> as_chunks needs a manual iter_mut, too_many_arguments,
  complex types, missing Safety docs, ...) and are tracked as the
  follow-up.
2026-09-15 19:29:32 +08:00
Mike-Solar 7e87ec135e render: the M4 audit follow-ups — autocache priority, cancel-in-flight, decode LRU
- Autocache range jobs now post at Background priority
  (submit_video_background): they used to go through the Seek path and,
  after the M4 seek over-admission, jumped ahead of playback and past the
  render-queue bound. The interactive single-frame preview keeps Seek.
- Job.cancelled: the arena installs the slot's cancel atom, and
  execute_job finishes a cancelled job with Error::State before running
  the producer — a cancel no longer burns a full render/GPU pass only to
  discard the result. Exactly-once delivery is unchanged.
- DECODE_LRU_CAP 8 -> 2: the decode service's LRU is a hand-off buffer,
  not the cache of record (the eval-side decoded_frames LRU is); the
  double-cache footprint at 1080p F32 drops by ~6 frames. A hand-off miss
  is served from the eval cache without a new decode.
- Tests: sequence-aware preview cancel, over-admitted seek ordering,
  deterministic prefetch LRU reuse, cancelled-job skip, autocache
  priority. docs §3.4 backfilled with the A/B/C audit outcomes.
2026-09-15 19:29:17 +08:00
Mike-Solar ba1143e7a3 render: the M4 playback prefetch — dependency window, priorities and backpressure
docs/zh/plans/render-pipeline-threads.md M4: the thread pipeline now
keeps its decode thread ahead of the render thread and the app's
playback window consumes in-process frames.

- Render queue: priority-ordered by JobSchedule.priority (Seek >
  Playback > Background, FIFO within a class), so interactive frames
  jump playback exports/autocache. Seek posts may over-admit the bound:
  priority only reorders queued jobs, so a full queue of background work
  must not park the UI thread until an export frame finishes.
- Decode queue: rendezvous Requests are served ahead of queued
  Prefetches (a frame the renderer needs never waits behind speculative
  decodes); Sync barriers stay FIFO. The queue is a bounded
  Mutex+Condvar structure, preserving the request backpressure and the
  wait_idle contract.
- Playback read-ahead: a Playback job's footage decode requests are
  derived from its montage/footage spec on post (same media time, size
  and force_format.unwrap_or(F32) as the eval) and queued immediately,
  so frame N+1 decodes while frame N runs its GPU passes.
- App window: PreviewWindow slots are generalized to
  PreviewSlot::{Shm, Video}; the pipeline's in-process TicketPayload is
  cached and consumed by cpu_frame exactly like a worker slot.
  PipelineBackend::preview_window_capacity reports the render-queue
  headroom, so playback posts are capped to what the queue can take;
  cancel_preview_frame drops queued frames the playhead has passed,
  matched on the full (sequence, frame, version) key so one monitor's
  window never drops the other sequence's same-numbered frame.
- Tests: decode-queue preemption/FIFO, render-queue ordering, request
  derivation, and deterministic end-to-end M4 tests: a prefetch that
  must be reused by the render request (LRU hit, single decode — the
  read-ahead claim is falsifiable), a parked-render-thread priority test
  where a full queue of background work still lets a Seek over-admit and
  run first, and a sequence-aware cancel test. The playback prefetch
  smoke asserts prefetches == distinct decodes == frames; it does not
  claim zero heap copies (Frame.data is deep-copied at the eval-cache
  and service-LRU boundaries today).
- bench_playback gains a pipeline mode with CPU (self+children) and
  first-frame latency; both backends now produce F32 frames so the
  comparison is like-for-like. The §3.4 backfill records the numbers:
  at the proxy size the pipeline is faster with a lower first frame; at
  1080p peak throughput is below the multi-worker pool, but that is an
  artifact of the decode still being CPU software (M5), not a case for
  pooling decode threads — GPU decode is a single device/queue and the
  zero-copy import shares one GPU memory pool, so the single decode
  thread stays the target shape.
2026-09-15 17:25:01 +08:00
Mike-Solar fec6e9dba7 render: the M3 OFX host — one oak-worker --ofx-host process for every plugin job
docs/zh/plans/render-pipeline-threads.md M3 (design 3.2): OpenFX crash
isolation moves from "every worker hosts plugins" to a single dedicated
host process, served over NDJSON + shared memory.

- oak-worker --ofx-host mode (src/ofx_host.rs): loads every plugin once,
  resolves jobs by the cross-process-stable OFX identifier, and renders
  through the same in-process executor the workers used to install.
- oak-render/ofxhost.rs: the single-host client. The render manager
  creates and installs it for the Pipeline backend (lazy spawn on the
  first plugin job); eval::process_plugin_job prefers it and falls back
  to the in-process executor otherwise, so the process backend keeps its
  current behavior until M4.
- Data plane: input/output FrameSlotPool pairs (the handshake's input_*
  fields are used for the first time). Named clips and the source frame
  are written to input slots after the explicit CPU readback; the plugin
  output returns through an output slot. Pool size/capacity grow by a
  host restart when a job needs more (safe: submissions are serialized
  and one job is in flight).
- Crash loop: reader EOF fails the in-flight submit, which respawns the
  host and re-posts the same job (frames are read back once); after three
  consecutive crashes the client is permanently dead and the evaluator
  falls back to a purple frame. The dead child is reaped immediately, and
  a submit mutex enforces the one-job-in-flight contract.
- Progress/cancel: the host flushes plugin_progress immediately (live
  progress), and reads stdin on its own thread so plugin_cancel takes
  effect mid-render at the plugin's next progressUpdate; the sticky flag
  resets at progressStart and request_plugin_cancel_all broadcasts to
  both the worker pool and the host.
- JobSpec::Plugin / PluginJobPayload carry the plugin type_id (stable
  across processes); `--ofx-crash-once` / `--ofx-crash-always` are the
  deterministic crash hooks, matching the worker's env hooks.
- Tests: wire round-trips; host unit tests (crash budget, cancel-flag
  reset through the factory, source mapping); oak-worker integration
  tests against the real host + bundled test plugin (render + progress,
  crash respawn and re-post, three-crash give-up, mid-render cancel on
  the new slow variant, concurrent submits); eval's purple fallback.
2026-09-12 23:10:43 +08:00
Mike-Solar 48e99e56b7 render: the M2 GPU zero-copy pipeline — wgpu 29, shared gpui device, GPU color LUTs
docs/zh/plans/render-pipeline-threads.md M2: the graph's textures stay
on the GPU from evaluation through presentation, and presentation runs
on the UI's own wgpu device.

- wgpu 25 -> 29 (naga 29) across the engine, unifying it with
  gpui_wgpu so engine textures are directly sampleable by the presenter
  (a single wgpu remains in the lockfile).
- GpuContext::adopt/install_shared: the app registers the window's
  device at startup and the render thread renders on it;
  texture_handle hands the raw Arc<wgpu::Texture> to
  SurfaceSource::Texture - zero-copy present on Linux/FreeBSD. The
  shared slot replaces an engine context that has not touched the GPU
  yet (startup-order guard) and refuses once it has.
- Texture::Gpu shares a GpuLease so clones release the registry token
  exactly once; the compositor, transitions and adjustment sweeps keep
  GPU textures end to end (no per-clip readbacks; GPU clears for
  black/generated frames).
- Color management stays on the GPU: the output node + display ICC
  chain is baked into a 65^3 3D LUT with the exact CPU reference and
  applied by the present WGSL pass (manual trilinear);
  ColorTransformJob bakes its OCIO processor the same way. Neither
  path skips color management.
- The explicit readback boundaries accept GPU textures: export
  encoder, CLI, worker shm, disk cache; CPU OpenFX already read back.
- M5 dependency: the YUV->RGB GPU pass (BT.601/709/2020 x
  limited/full) matches colormath::yuv444p16_to_rgb_f32.
- Acceptance: gpu_transfer_counters; single-clip and layered
  (multi-track + transition + adjustment) playback tests assert zero
  GPU->CPU readbacks, and the app test asserts adopted-device present
  is zero-copy. GPU tests hard-fail when OAK_REQUIRE_GPU is set (CI
  lavapipe) instead of skipping silently.
2026-09-12 20:52:17 +08:00
Mike-Solar a5b0b2a1b1 render: the M1 thread pipeline — one render thread, one decode thread
docs/zh/plans/render-pipeline-threads.md M1: an in-process
alternative to the worker-process pool, behind OAK_PIPELINE=threads
(processes stays the default and is fully retained).

- pipeline.rs: PipelineBackend implements JobDispatch over a single
  render thread draining a bounded FIFO (cap 8; blocking post with
  condvar backpressure and a one-ahead exception for re-posts from
  the render thread itself; shutdown drains with Error::State like
  the inline dispatcher). The DecodeService is a single decode
  thread behind a bounded command queue with a real LRU (tick-based
  eviction), rendezvous requests (None on shutdown -> the caller
  decodes inline), prefetch gated on render-queue room, and a Sync
  barrier; it installs into a process-wide slot that eval's footage
  path consults per frame (no service -> the synchronous decode it
  always was).
- The manager gains RenderBackendChoice::Pipeline; init() reads
  OAK_PIPELINE (threads -> pipeline, anything else -> the process
  pool), audio stays deliberately inline.
- Present mapping: the UI thread consumes through the ticket
  completion, unchanged — no fourth thread is invented.
- Tests: decode-service unit tests (rendezvous, LRU hit/eviction,
  error propagation, backpressure gate) plus a six-case integration
  suite matrixed over inline vs pipeline — consecutive-frame and
  out-of-order seek pixel equality asserted byte for byte, with
  decode counters proving the service (not the caller) did the
  codec work.
2026-09-11 19:37:27 +08:00
Mike-Solar 4f0f5cbba6 workspace: zero compiler warnings across all targets
254 warnings (320 counting replayed-cache re-emitters) cleaned:
unused mut/imports/variables, irrefutable if-lets and unreachable
patterns, dead code removed or annotated #[allow(dead_code)] with
the reason (C++ parity value sets, cfg(test) helpers, public API
reservations), drop(&ref) no-ops removed, fn-pointer identity via
std::ptr::fn_addr_eq, the test-stubs feature declared in
oak-node's manifest, missing docs filled. Every unused-Result site
was judged individually: meaningful errors propagate, intentional
ignores are let _ = with a note.

Two pre-existing latent bugs are documented in place, behavior
preserved: app.rs's timeline-tool observer and dialogs.rs's format
subscription both drop the returned Subscription immediately, so
they never fire.
2026-09-11 16:38:44 +08:00
Mike-Solar 80e6b8bb6e render: skip posix_fallocate off Linux (macOS CI)
macOS has no posix_fallocate (and the libc crate rightly does not
expose it there), so the shm segment setup failed to compile. The
eager reservation is a tmpfs concern; off Linux the call is skipped
and the existing touch-every-page fallback runs instead.
2026-09-11 15:49:36 +08:00
Mike-Solar 29204d1f63 node: virtual graph endpoints and the Kahn-order BFS sweep (M0b core)
Per docs/zh/plans/render-pipeline-threads.md §3.8:

- oak-node/nodes/graphendpoints.rs: the GraphInput/GraphOutput
  virtual node pair — factory-registered but hidden from every create
  menu, duplicate refused, real value() semantics (the input forwards
  its feed_in row, the output publishes its tex_in as the frame).
  The input endpoint also declares a connectable feed_in port
  (documented deviation: footage/generator sources have no connectable
  inputs, so the walk needs a feeder anchor).
- graph.rs: ensure_endpoints/endpoints/is_endpoint — idempotent,
  identified by type id, default input->output edge only while the
  output's tex_in is free; remove_node refuses endpoints.
- project.rs + serializer.rs: every project graph carries the pair;
  a legacy file without endpoints migrates on load (roundtrip and
  legacy-migration tests, re-save is idempotent).
- traverser.rs: eval_graph_bfs — the endpoint-to-endpoint Kahn
  sweep. Live set = (input's forward cone U its feeder cone) INTERSECT
  (output's backward cone); multi-input nodes dequeue at zero
  in-degree over the live subgraph; deterministic ascending-id ready
  order (Graph::edges is a BTreeSet, so insertion order is
  unrecoverable — documented); time-shifted upstreams pull through
  the shared DFS memo (walk_dfs, factored out of evaluate);
  un-orderable remainder reports a named cycle; missing endpoints /
  unreachable output are errors. Eight BFS tests cover the plan's
  acceptance bullets.
- oak-render: bfs_endpoint_sweep_renders_footage_through_position —
  real clip through a real Position node via the sweep, shifted
  pixels asserted against a reference decode.
- Endpoint names localized in all eight i18n packs; storage/structure
  tests updated for the two extra nodes.
2026-09-11 15:13:53 +08:00
Mike-Solar 3a48dd4991 render: Job enum in the tables, single-loop match resolve, real CacheJob
M0a of the render-pipeline plan (docs/zh/plans/render-pipeline-threads.md):

- oak-node: every payload push site (58 across footage.rs, plugin.rs
  and the nodes/* effects) now boxes the Job enum instead of the raw
  payload. The enum gains CacheJob with a CacheJobPayload (path +
  time + fallback value, the C++ cachejob.h shape), plus safe as_*
  accessors and unsafe probe helpers beside job_ref.
- oak-render: RenderEvalHooks::resolve is one loop over the table —
  a single get_checked::<Job> probe per texture value, a match
  dispatch to process_footage/shader/plugin/color_transform/cache,
  and recursive resolution of the job boxes embedded in a payload's
  inputs (depth-capped, cycle-guarded) — replacing the four
  sequential full-table scans (resolve_*_jobs, deleted).
- The disk frame cache is real: frameio.rs implements a minimal
  self-describing F32 container (magic/version/dims/format/timestamp
  + payload, tmp-write + atomic rename, full header validation on
  load) because the OIIO bridge is a stub and EXR is unavailable in
  this build; process_cache_job genuinely reads the file before
  falling back to the job's (already resolved) fallback value.
- Tests: CacheJob roundtrip (save -> resolve -> pixel equality),
  missing-file fallback, nested cache-job-through-shader resolution,
  plus four frameio container tests. 2330 passed, 0 failed across
  the workspace.
2026-09-11 10:38:12 +08:00
Mike-Solar 2817ac286c timeline: drag generators onto the timeline, drop transitions at clip edges
Generator effects (bars, checkerboard) can be dragged from the library
onto the timeline, where they land as a standalone five-second clip
built from the node factory; the inspector shows the generator's
parameters as the clip's own chain.

Transitions are no longer junction-only. The render planner accepts a
transition with at least one wired neighbor and blends the missing
side against transparent black, so head transitions fade in from black
and tail transitions fade out to black. add_transition_at_edge creates
those single-sided blocks (wired to just the IN or OUT block), the
default-transition command covers both ends of a lone clip, and an
effect drag dropped near a clip edge routes to the nearest seam or
edge within a one-second window.
2026-09-11 08:49:08 +08:00
Mike-Solar 16364d414a nodes: mask off-frame samples to transparent in the distort shaders
Translating, rotating or warping content past the frame edge used to
smear the clamped edge row/column across the vacated region. The
transform, position, swirl, ripple and wave shaders now multiply the
sample by an in-bounds mask so off-frame pixels come out transparent
(and composite as black when nothing sits below). Tile deliberately
keeps its wrapping lookup.
2026-09-11 08:48:48 +08:00
Mike-Solar 19b7d3ac78 render: move the text engine into oak-render and install it in the worker
The render worker process never installed a text backend, so text clips
rendered as empty frames in playback and export. The cosmic-text engine
now lives in oak-render (the crate both the app and the worker link),
and the worker installs it during runtime initialization.
2026-09-11 08:47:33 +08:00
Mike-Solar 4a2614b3fc timeline: adjustment layers and first-class transitions
Adjustment layers (docs/zh/plans/adjustment-layers-and-transitions.md):
a new timeline block type whose effect chain grades the composite of
every video track below it, over its own range (spanning clips or a
slice of one). The graph path flushes the lower tracks at the block's
track boundary and sweeps the composite through the chain via a
transient texture-source node; the montage path mirrors it with
AdjustmentSpan tickets (wire-compatible), so worker previews and
exports agree. An empty-area context menu creates one; the block
trims/moves/deletes like a clip, with undo everywhere.

Transitions: seam blocks come alive - cross dissolve/fade/wipe/slide
evaluate both neighbors through the graph path with progress from the
transition's own range (never the whole clip). Ctrl+Shift+D or the clip
menu inserts a default transition; the gpui wedges render and drag to
resize offsets undoably, and TransitionRemoveCommand now restores
offsets and edges on undo. The transitionfx node form runs the same
shaders on an adjustment layer with progress_in auto-filled from the
layer's span (explicit value wins).

Also: every built-in effect name and parameter name is now
translatable (360 node.* keys per locale, zh-CN fully translated, two
coverage tests guard future gaps); the new nodes register in
nodes/mod.rs with the factory smoke table updated; textfootage and
adjustment-layer i18n keys included.
2026-09-10 22:03:15 +08:00
Mike-Solar 1e0d48578e nodes: text becomes structured footage with a real text engine
Text is no longer a hand-written HTML effect (docs in
docs/zh/plans/text-footage-redesign.md):
- textv3 gains structured inputs - plain text, font family/size, font
  color, outline (enable/color/width), glow (enable/color/radius); the
  legacy text_in HTML is hidden and auto-migrated to plain text on load.
- Outline and glow render as GPU post-process chains (dilate/blur +
  colorize under the text); the font color tints the raster
  premultiplied. Plain text now rasterizes even with both passes off
  (previously a null deferred job), fixing a use-after-free where the
  handle was lifted out of an owning Option<NodeValue> before addref.
- A cosmic-text backend (the lockfile's 0.19) installs at engine
  startup through the textbackend hooks and feeds the font-family combo.
- The project panel gains 添加文本素材 next to 新建序列: a text entry
  in the bin that drops onto the timeline as a clip (one undo row), its
  parameters shown as structured fields in the inspector (multiline
  text area, no HTML anywhere). text3 is hidden from the effect add
  menus; legacy text3 chains keep evaluating.
2026-09-10 22:02:57 +08:00
Mike-Solar a7916aa93d nodes: OpenFX-Misc cleanroom GPU ports, grouped and collapsible in the library
21 built-in effects reimplemented as native GPU nodes from the
OpenFX-Misc algorithm references (cleanroom, docs in
docs/zh/plans/ofx-misc-gpu-cleanroom.md):
- Color: Color Correct, Gamma, Saturation, Invert, Clamp, Grade
- Matrix/morphology: Color Matrix, Edge Detect, Dilate, Erode
- Blur: Directional Blur, Sharpen (unsharp mask)
- Merge: Dissolve, Key Mix, Premultiply, Unpremultiply
- Geometry/generators: Position, Mirror, Checkerboard, Color Bars, Ramp

Every node carries unit tests plus GPU pixel tests (28 cases over five
ofxmisc_* suites). The effect library groups built-ins by category
(color/filter/distort/keying/generator/math/general) with collapsible
group headers persisted to the config; the inspector's add menu groups
the same way. Registration wiring and the factory smoke table land with
the adjustment/transition wave sharing the same files.
2026-09-10 22:02:38 +08:00
Mike-Solar d028a45ffa nodes: pivot transform rotation/scale around the frame center
The transform shader sampled in a top-left-origin pixel space while
Olive's transform semantics (and every other node) are center-origin:
rotation swung the image around the top-left corner, pushing it partly
off-frame - reading exactly like an unwanted zoom. Match the C++
transform.vert projection: position (0,0) is the frame center and
rotation/scale pivot around the anchor, so rotation and scale stay
independent user controls. GPU tests pin the 90-degree landing spot
(no smearing) and the 2x scale centroid (stays centered).
2026-09-10 18:13:05 +08:00
Mike-Solar fc90604245 render: anchor resolution_in to the sequence resolution
Preview renders at proxy size while a paused frame renders full-res, so
anchoring resolution_in to the render target made every sequence-pixel
effect (shape size/pos, transform offsets, corner pin points, drop
shadow distance) change apparent size whenever the transport stopped.
Pre-fill resolution_in from the sequence's video params (C++ inserts
the NodeGlobals square resolution at job-build time), covering the
nested generator job inside a merge as well; a node that inserted its
own resolution_in keeps it.
2026-09-10 16:53:55 +08:00
Mike-Solar 37df1d3d34 nodes: spell shape/despill shader dispatch as if/else chains
CI / Build & test (Linux) (push) Successful in 24m7s
CI / Build & test (Windows) (push) Successful in 30m31s
naga's WGSL emitter rejects fall-through-capable GLSL switch blocks, so
every shape and despill job failed to compile and silently fell back to
the effect input - both effects were no-ops. Rewrite the type/method
dispatch as if/else chains (same semantics as the C++ shaders) and
cover all three shape types plus green-screen despill with GPU pixel
tests. Also drop the now-stale nested-payload/merge-binding TODO notes.
2026-09-09 16:51:13 +08:00
Mike-Solar 4f404f8cc6 tests: replace placebo assertions with real behavior checks 2026-09-09 16:33:35 +08:00
Mike-Solar 5ab12b937f render: real texture binding, generator layers and iteration feedback in shader passes
- process_shader_job: bind all texture params by name, recurse into nested
  shader payloads (depth cap 8), fall back to frame size without inputs
- run_effect: take iterative_input so dropshadow previous_iteration_in works
- merge: actually composite inputs; keyer mask, opacity modulation, math
  texture ops and mrg generator layers now bind their textures
- transform distort: real fragment-side inverse-matrix sampling
- time offset / time remap: wire NodeBehavior time adjustment hooks
- plugin: fix first-node identity colliding with unbound sentinel
2026-09-09 16:33:29 +08:00
Mike-Solar cde4dddf32 structure: move jobs out of nodes/ 2026-09-03 17:44:46 +08:00
Mike-Solar 4babbf5de8 core: merge oak-common into oak-core
CI / Build & test (Linux) (push) Successful in 24m6s
CI / Build & test (Windows) (push) Successful in 31m14s
oak-common is gone; its modules (configstore, xmlutils, ocioutils,
oiioutils, colormath, colortransform, videoparams, ffmpegutils, ...)
now live in oak-core alongside the value types. The render value/GPU
types moved too: backend (wgpu context + DisplayRenderer), color
(ColorProcessor over ocio-rs), texture, frame, and the commonutil
config helpers.

Fix-ups to make the merged tree build and pass tests:

- oak-core Cargo.toml: wgpu back to 25 (the moved backend code is
  written against that API generation); add the toml/quick-xml/image
  deps oak-common carried.
- lib.rs: drop the duplicate 'pub mod error;'.
- error.rs: unified OAKCORE_* codes; restore Error::new() and
  From<OcioError> from oak-common's error type.
- backend.rs/color.rs: oak_core::/oak_render:: self-references
  rewritten to crate::; the shaderfx-dependent GPU effect test moved
  to oak-render's shaderfx tests (shaderfx depends on oak-node and
  cannot live in oak-core).
- oak-render's error module re-exports oak_core::error::{Error,
  Result}; the OAKRENDER_* codes stay as the public-code contract.
- oak-node jobs.rs: ColorProcessor imported from oak_core::color.
- Integration tests repointed at oak_core::{texture, frame, backend,
  color, colormath}.
- the display-ICC regression test treats an empty OAK_DISPLAY_ICC as
  unset, matching displayicc::env_override_icc.
2026-09-03 17:42:20 +08:00
Mike-Solar 49fed365a4 nodes: real OCIO color grading (linear + log) on the GPU
The OCIO grading nodes previously pushed null texture handles; they now
push real ShaderJobPayloads whose GLSL is the OCIO-generated dynamic
grading-primary GPU shader — the exact code the C++ path applies, so no
approximation:

- color: grading_primary_function_shader(style) builds a dynamic
  GradingPrimaryTransform (LIN/LOG) on the default config, extracts the
  GLSL via GpuShaderDesc (function 'ove_grading_primary', resource
  prefix ocio_, no LUT textures) and caches it per style + config id.

- eval: OCIO_GRADING_STUBS maps the two node type ids to the grading
  style; process_shader_job resolves the stub and splices it into the
  node's %1 marker (same wiring as the chromakey OCIO stub); the
  pipeline cache key folds the stub text so a config change recompiles.

- nodes: value() pushes a ShaderJobPayload with the C++ value()
  rewrite applied to the row — vec4 (RGBM x=master) grading inputs to
  the vec3 GPU uniform form (lin: contrast RGB=c*m, offset RGB=c+m,
  exposure RGB=2^(c+m); log: lift RGB=c+m, gain c*m, gamma c*m), plus
  pivot/saturation floats, the log pivotBlack/pivotWhite normalization
  range (0/1), clamp sentinels (NoClampBlack -1 / NoClampWhite 2),
  white>black enforcement per frame, and localBypass=false. Generated
  uniform names bind by name (the log node's OCIO_NAMESPACE_ id text
  normalizes to the ocio_ resource prefix).

- Tests: grading stub generation (analytic GLSL, cache) in color,
  end-to-end GPU exposure doubling for lin (+1 stop on 0.2 gray -> 0.4)
  and lift for log, node payload rewrite assertions, and the
  all-shaders sweep now retries grading stubs. oak-render 181,
  oak-node 441, oak-app 272 lib tests pass.
2026-09-02 20:04:55 +08:00
Mike-Solar b06c4fbbb6 nodes: real polygon and mask implementations on the GPU
Polygon and mask previously pushed null texture handles ('fake'
implementations). They now generate real ShaderJobPayloads and render
through the existing GPU shader pipeline:

- shaderfx: std140 uniform array support (Vec4Array(N)) — the parser
  accepts 'uniform <type> <name>[N];' declarations, translate()
  re-emits them as vec4[N] block members with per-element std140
  offsets, and pack_uniforms writes array items from the new
  NodeValue::Vec4Array value, padding short arrays to the declared N.

- polygon: value() collects the inherited points array (row element
  keys 'points_in[i]', else the node's own per-element values — an
  unconnected array resolves to the default pentagon via
  GetValueAtTime parity) and pushes a ShaderJobPayload; the 'rgb'
  fragment shader rasterizes the closed point loop with an odd-even
  fill in screen space (center-translated, y-flipped to match the C++
  point convention) and outputs color_in inside / transparent outside.
  The CPU QPainterPath generate_frame stays a documented no-op.

- mask: value() pushes a single ShaderJobPayload whose new 'mask'
  fragment shader folds the whole C++ chain into one GPU pass — base
  texture multiplied by the polygon matte, optional invert, and the
  optional feather gaussian softens the matte during sampling (the
  separable blur.frag h/v iterations as a one-pass product,
  density-normalized, radius capped at 16 px).

- Tests: translate/pack array coverage in shaderfx, GPU end-to-end
  rasterization of the pentagon (center white, corner transparent),
  mask multiply/invert/feather on real frames, and updated oak-node
  payload assertions. oak-render 178, oak-node 440, oak-app 272 lib
  tests pass.
2026-09-02 16:59:49 +08:00
Mike-Solar 4ef5b3e31d render: snapshot store temp dir is unique per store (test parallel flake)
CI / Build & test (Linux) (push) Successful in 23m2s
CI / Build & test (Windows) (push) Successful in 30m53s
Tests run in parallel inside one process; each GraphSnapshotStore::new()
used the same 'oakrender-snapshots-<pid>' root, so one test's cleanup()
deleted another test's live snapshot — the
acquire_rewrite_forces_file_rewrite_on_same_key flake on the 16-core
Windows runner (file written, then exists() == false). Append a per-store
sequence number to the directory.
2026-09-02 16:19:01 +08:00
Mike-Solar 6153a2ac33 tests: fix three stale assertions
CI / Build & test (Windows) (push) Failing after 13m15s
CI / Build & test (Linux) (push) Successful in 24m22s
- multicamnode: type_id() assertions called the method on dyn
  NodeBehavior, which method resolution routed to std::any::Any's
  TypeId::of — qualify via NodeBehavior::type_id so the trait method
  (the &str node type id) is compared
- procpool: the per-worker GPU budget grew a security headroom (×2)
  for the CUDA-OOM flood; the two budget tests now assert against the
  real formula (2 GiB + 256 MiB at 1080p24, 2 workers at 4K/24 GiB)
- cli info fixture: the fixture runs 29.97 fps; the assertion expected
  30/1 (stale from the older fixture)
2026-09-02 12:34:39 +08:00
Mike-Solar be42620e22 multicam: throttle angle refresh, shrink pre-render window
- MulticamPanel: playback angle refresh runs one cycle per 3 ticks
  instead of re-requesting every source every tick — each angle decode
  is a keyframe-scanning FFmpeg seek that stole worker capacity from
  the main viewer
- PreRender frames default 120 -> 12: a window larger than what the
  pool can render in real time queues far ahead of the playhead, so
  the painted frame lags seconds behind (playback frozen); a smaller
  window keeps the backlog bounded
- render_graph_frame: OAK_PERF clip-level timing
2026-09-01 21:52:56 +08:00
Mike-Solar 7f41570596 render: stop evicting the only hardware decoder session on every open
The per-process 'hardware session == 1, evict before every new open'
guard was turning every frame's open into a decoder re-open: after
inserting the fresh session, the NEXT frame's pre-open eviction dropped
it again, so no request ever hit the cache ([open] (request) on every
single frame, zero CACHED hits). Every decode then cost a full FFmpeg
session open (~0.5 s) + a keyframe-scanning seek — playback could never
keep up (the 'main viewer barely moves' report; [perf] showed 1.6-2.2 s
per frame).

Hardware VRAM is bounded by the LRU cap (hardware-first eviction at
MAX_CACHED_DECODERS) and the gpu-vram worker-count policy; the eager
pre-open eviction was the regression.
2026-09-01 21:46:01 +08:00
Mike-Solar 1b0a15192e multicam: wizard + AFV + node-graph preview + performance
CI / Build & test (Linux) (push) Failing after 24s
CI / Build & test (Windows) (push) Failing after 3m38s
- wizard: angle multi-select, sync modes, auto-align, create sequence;
  keeps the host sequence current (host clip = the multicam clip)
- build_multicam_sequence: per-angle clip -> SOURCES_INPUT[element],
  array slot growth, audio angle tracks for AFV
- AFV: host linked audio clip follows the switched source (one undo),
  muted host audio track disables it
- node-graph preview: inline producer renders through the traverser
  (viewer/project-matched graph frames); sequence viewer uses the graph
- multicam node value(): element-tagged row keys (sources_in[i]),
  reads the current source; build_row keys array inputs by element
- angle grid: viewer=0 (single-track montage, not whole-graph)
- project explorer: rename (dialog) + delete (undoable) real items
- performance: decoded-frame LRU, per-process NVDEC quota (1 session,
  evict before open), GPU composite fail-once fallback, snapshot
  upload debounced on the engine tick, worker vram budget headroom
- timeline clip: multicam overlay via ClipDecorator
- wizard menu item moved to Sequence menu
2026-09-01 20:27:59 +08:00
Mike-Solar 8de9af705e render: query GPU vram on AMD/Intel Linux too, document the UMA/Windows fallbacks
gpu_vram_bytes() now chains per vendor/platform:
- NVIDIA everywhere: nvidia-smi (shipped by the NVIDIA driver on every
  OS) — the NVDEC path's primary device.
- AMD/Intel on Linux: the DRM mem_info_vram_total/used sysfs attributes
  (amdgpu, i915, Xe). Free = total - used; the first non-zero card wins
  (an iGPU without dedicated vram reports 0 and is skipped). The walk is
  now testable via an injectable sysfs root; a card missing the attrs is
  skipped, never aborts the walk.
- Apple Silicon: unified memory — no separate vram exists; the RAM/4
  budget IS the correct bound for decode surfaces and render targets, so
  no query (an explicit vram budget would double-count the same pool).
- Windows AMD/Intel: no portable CLI; DXGI QueryVideoMemoryInfo is the
  real API but wgpu 25 does not expose it. Falling back to the RAM
  policy is safe-side (under-sized pool loses throughput, never OOMs the
  device).

Tests: the sysfs walk (fixture with a missing-attr card, a 0-total
iGPU and a discrete winner) and the per-worker budget scaling (1080p
baseline, 4K ~4x, 60 fps over-provision).
2026-08-30 22:03:05 +08:00
Mike-Solar 9352fca4a9 codec/render: silence hw-decode failures, vram-aware dynamic worker pool
- open_hw_accel marks the device unavailable when the decoder OPEN fails
  (cuvidCreateDecoder OOM at 4K) too, not just device-context creation:
  without it every subsequent decoder session retried CUDA and flooded
  the log per open.
- Decoder gains hardware_decoding(); the oak-render decode-session LRU
  evicts hardware sessions first (each pins a GPU surface pool — ~100 MB
  at 4K), so a full cache cannot exhaust video memory before the next
  open.
- Worker pool count now factors GPU vram: per-worker budget = 1 GiB
  (1080p peak) scaled by pixel ratio + 256 MiB idle floor, 10% reserve
  of free vram; applied when hardware decoding is on (nvidia-smi query,
  None otherwise falls back to the RAM/CPU policy).
- Dynamic pool resize: ProcessDispatcher::set_target_workers grows or
  retires workers; retiring ones stop claiming, drain their in-flight
  batch (future playback frames included), then exit naturally on the
  shutdown signal — no mid-work kill (30 s deadline only as a hung-
  decoder last resort). A retiring worker that dies re-queues its frames
  to surviving workers. Resizes are throttled to 2 s (a resolution burst
  merges; only the latest target applies) so 1080p<->4K flaps cannot
  thrash process spawns.
- RenderManager::set_workspace_size announces the sequence resolution;
  RealEngine calls it from refresh_sequence_info.
- Integration test: shrink 3->1 mid-wave (all frames complete, retired
  workers exit naturally) then regrow 1->3 and render a fresh wave.
2026-08-30 21:59:06 +08:00
Mike-Solar 5937e557a7 app: project-explorer new-sequence button, 4K presets, keying node fix, color-picker canvases
- Project explorer header gains a 新建序列 button (opens the existing
  new-sequence dialog, seeded like the menu action).
- Sequence presets gain 4K UHD (3840x2160@25) and 4K DCI (4096x2160@24);
  the sequence-properties dialog re-selects them on reopen. Format fields
  are also seedable from a probed footage format (the drop flow).
- Chroma Key (and Color Difference Key) value() now box a ShaderJobPayload
  like Despill: the old OCIO-processor gate pushed nothing (the processor
  is never populated without the render bridge), so the traverser handed
  the clip NodeValue::None and the rendered frame lost the clip. The
  renderer resolves the OCIO stub at compile time from OCIO_SHADER_STUBS.
  End-to-end graph test: green key on a green frame keys out, red key
  keeps it.
- The OFX color picker's SV palette / hue bar / preview / swatch canvases
  get size_full(): the bare canvases collapsed to zero height in the
  block layout, so the palette painted nothing (the reported 色板没显示).
  Regression test clicks the palette center and expects mid s/v.
2026-08-30 21:58:26 +08:00
Mike-Solar 9dab9efc38 codec: carry codec-frame overflow across audio chunk boundaries
retrieve_audio_to decodes whole codec frames but copies only the part
inside the requested chunk; the tail of the frame crossing the chunk end
(up to 1023 samples for AAC) was consumed by the decoder and lost, so
the next chunk started with a hole. On the playback grid (1920 samples
at 25fps/48kHz) the hole cycles 128..896 samples and hits 7 of 8 chunk
starts -- the heavy stutter/noise heard during playback.

Keep the overflow (decode and resampler-flush tails) in a per-session
carry buffer and serve it at the start of the next contiguous chunk;
clear it on seek, format change and chunk failure. Also make seek()
actually drop the cached resampler as its comment claimed.

Verified sample-exact: chunked decode of a 440Hz tone now matches a
one-shot decode bit for bit, and chunked renders of real media line up
with the ffmpeg CLI reference at correlation 1.0 / drift 0.

Adds a regression test (playback_sized_chunks_match_oneshot_sample_exact)
with a new tone fixture -- demo.mp4's audio is -90dB digital silence and
cannot expose the holes -- plus a render_audio_wav example used to
diagnose chunk-boundary artifacts offline.
2026-08-29 19:51:59 +08:00
Mike-Solar e128a6aca1 render: clamp mixed audio to [-1,1] before it reaches the device
Overlapping clips sum linearly in mix_audio_montage and can exceed full
scale (two hot clips reach +/-2; gain > 1 would too); the cpal sink
forwarded samples unclamped, so overlaps clipped at the DAC. Clamp the
accumulator after the montage mix (both the heap and shm-slot paths
share mix_audio_montage) and document it on render_audio_samples.
2026-08-29 19:51:25 +08:00
Mike-Solar cbd4ba2442 optimize: video and audio render
CI / Build & test (Linux) (push) Canceled after 0s
CI / Build & test (Windows) (push) Canceled after 0s
2026-08-29 05:01:55 +08:00
Mike-Solar fdb5caabd5 color: non-sRGB preview, per-monitor display ICC, pipeline hardening
Preview now follows the project output colorspace end to end: the
display chain derives its content space from the project's OutputColorSpec
instead of a hardcoded sRGB name, self-managed ICC transforms go through
an XYZ D65 interchange stage (OCIO cie_xyz_d65_interchange) for non-sRGB
targets, and the platform layer declares the content colorspace (gpui
submodule bump). macOS defaults to OS-managed (fixes wide-gamut UI
oversaturation); Windows ACM warns once on non-sRGB targets.

Multi-monitor: the display ICC is looked up per the window's current
screen (macOS display id, Windows per-monitor DC, X11 RandR output
profile) with a throttled poll that invalidates frame caches on moves.

Pipeline precision: 10-bit+ sources fall back to YUV444P16LE + a Rust
matrix conversion when swscale lacks F32 output (no more 8-bit
truncation); BT.709/2020 SDR decodes with BT.1886 gamma 2.4 instead of
the sRGB EOTF; working-space compositing no longer clamps RGB to [0,1]
(alpha still clamped); the output node clamps to the target gamut;
frames without colorimetry metadata convert with BT.709 defaults
(warned once) instead of passing through; scopes read the
output-colorspace signal on both F32 paths.

Also: only emit rerun-if-changed for .env when it exists (a missing file
made every build fully dirty).
2026-08-29 00:24:15 +08:00
Mike-Solar 9c0269bcbb render: cache worker frames and decouple audio from the video workers
- oak-worker gains an LRU frame cache (default 64 MiB) keyed by the
  render-deterministic spec subset; repeat frames (paused frames,
  scrubs over rendered ranges, re-renders after an effect change) are
  memcpy-cheap, which is what made adding an OFX plugin -- and the
  in-flight batches after removing one -- stall the UI
- audio dispatch on the Processes backend now mixes inline on the UI
  tick instead of queueing behind video batches in the worker pool, so
  video stalls no longer starve the ~100ms cpal output buffer into
  silence
2026-08-27 20:35:50 +08:00
Mike-Solar 5d21f83e1e render/app: display bit depth option (10-bit default, 8-bit optional)
Preferences gains a display-bit-depth combo stored in the config store;
at startup the app forwards it to gpui_wgpu via OAK_DISPLAY_BIT_DEPTH,
which prefers Rgb10a2Unorm for 10-bit presentation. Takes effect after
restart (noted in the dialog); i18n in all eight packs.
2026-08-27 19:08:55 +08:00
Mike-Solar c03f1ec604 render: stack higher-numbered video tracks on top
Match the timeline UI (V_max drawn topmost): composite tracks from
V1 up to V_max so the highest-numbered track is composited last, in
both the montage path and direct graph evaluation.
2026-08-27 15:20:18 +08:00
Mike-Solar f7352ae19d render: fall back to CPU when the adapter cannot render Rgba32Float
CI / Build & test (Linux) (push) Successful in 19m8s
CI / Build & test (Windows) (push) Successful in 32m35s
2026-08-27 13:52:03 +08:00
Mike-Solar c51a349070 render: translate node GLSL shaders to WGSL and run them as wgpu passes
CI / Build & test (Linux) (push) Failing after 16m57s
CI / Build & test (Windows) (push) Successful in 31m46s
2026-08-27 07:07:09 +08:00
Mike-Solar fa3951344b render: fix 4K playback memory growth and decode-to-target-size
Root causes found for the 4K stalls and the second-footage memory
blowup (audit + code review):

- ticket bookkeeping leaked unbounded: the procpool ticket table and
  the arena slot map only ever grew (50-100 tickets/sec during
  playback, each pinning montage params and shm region views).
  Completed/cancelled/superseded/crashed entries are now removed, and
  the arena reaps fire-and-forget tickets once finished; the sync poll
  path reaps via a terminal result() read. InFlight duplicate submits
  now answer State immediately instead of sitting in the map forever.
- decode ran a full-resolution swscale to F32 RGBA (~132 MB at 4K)
  plus a second full-res copy before downscaling to the 480px proxy:
  RetrieveVideoParams.target_size lets swscale convert AND resize in
  one pass (bilinear, matching the old Rust resampler), so a 4K
  preview frame costs ~1 MB instead of ~260 MB of churn. This applies
  to proxy AND full-res requests alike.
- per-process decoder cache was unbounded (each session pins an FFmpeg
  context + 2 native decoded frames): LRU-capped at 16, eviction drops
  the map entry (in-flight renders keep their Arc; Drop releases
  FFmpeg).
- playback window completions were not generation-gated: a stale
  render from before an edit landed in the rebuilt window (wrong frame
  displayed, fresh request blocked). Stale completions now return
  their shm slot credit instead.
- async audio prefetch used the polling ticket submit without ever
  polling: switched to the fire-and-forget submit so entries reap.
2026-08-26 01:31:10 +08:00