Commit Graph
23 Commits
Author SHA1 Message Date
Mike-Solar 3a48dd4991 render: Job enum in the tables, single-loop match resolve, real CacheJob
M0a of the render-pipeline plan (docs/zh/plans/render-pipeline-threads.md):

- oak-node: every payload push site (58 across footage.rs, plugin.rs
  and the nodes/* effects) now boxes the Job enum instead of the raw
  payload. The enum gains CacheJob with a CacheJobPayload (path +
  time + fallback value, the C++ cachejob.h shape), plus safe as_*
  accessors and unsafe probe helpers beside job_ref.
- oak-render: RenderEvalHooks::resolve is one loop over the table —
  a single get_checked::<Job> probe per texture value, a match
  dispatch to process_footage/shader/plugin/color_transform/cache,
  and recursive resolution of the job boxes embedded in a payload's
  inputs (depth-capped, cycle-guarded) — replacing the four
  sequential full-table scans (resolve_*_jobs, deleted).
- The disk frame cache is real: frameio.rs implements a minimal
  self-describing F32 container (magic/version/dims/format/timestamp
  + payload, tmp-write + atomic rename, full header validation on
  load) because the OIIO bridge is a stub and EXR is unavailable in
  this build; process_cache_job genuinely reads the file before
  falling back to the job's (already resolved) fallback value.
- Tests: CacheJob roundtrip (save -> resolve -> pixel equality),
  missing-file fallback, nested cache-job-through-shader resolution,
  plus four frameio container tests. 2330 passed, 0 failed across
  the workspace.
2026-09-11 10:38:12 +08:00
Mike-Solar 2817ac286c timeline: drag generators onto the timeline, drop transitions at clip edges
Generator effects (bars, checkerboard) can be dragged from the library
onto the timeline, where they land as a standalone five-second clip
built from the node factory; the inspector shows the generator's
parameters as the clip's own chain.

Transitions are no longer junction-only. The render planner accepts a
transition with at least one wired neighbor and blends the missing
side against transparent black, so head transitions fade in from black
and tail transitions fade out to black. add_transition_at_edge creates
those single-sided blocks (wired to just the IN or OUT block), the
default-transition command covers both ends of a lone clip, and an
effect drag dropped near a clip edge routes to the nearest seam or
edge within a one-second window.
2026-09-11 08:49:08 +08:00
Mike-Solar 16364d414a nodes: mask off-frame samples to transparent in the distort shaders
Translating, rotating or warping content past the frame edge used to
smear the clamped edge row/column across the vacated region. The
transform, position, swirl, ripple and wave shaders now multiply the
sample by an in-bounds mask so off-frame pixels come out transparent
(and composite as black when nothing sits below). Tile deliberately
keeps its wrapping lookup.
2026-09-11 08:48:48 +08:00
Mike-Solar 4a2614b3fc timeline: adjustment layers and first-class transitions
Adjustment layers (docs/zh/plans/adjustment-layers-and-transitions.md):
a new timeline block type whose effect chain grades the composite of
every video track below it, over its own range (spanning clips or a
slice of one). The graph path flushes the lower tracks at the block's
track boundary and sweeps the composite through the chain via a
transient texture-source node; the montage path mirrors it with
AdjustmentSpan tickets (wire-compatible), so worker previews and
exports agree. An empty-area context menu creates one; the block
trims/moves/deletes like a clip, with undo everywhere.

Transitions: seam blocks come alive - cross dissolve/fade/wipe/slide
evaluate both neighbors through the graph path with progress from the
transition's own range (never the whole clip). Ctrl+Shift+D or the clip
menu inserts a default transition; the gpui wedges render and drag to
resize offsets undoably, and TransitionRemoveCommand now restores
offsets and edges on undo. The transitionfx node form runs the same
shaders on an adjustment layer with progress_in auto-filled from the
layer's span (explicit value wins).

Also: every built-in effect name and parameter name is now
translatable (360 node.* keys per locale, zh-CN fully translated, two
coverage tests guard future gaps); the new nodes register in
nodes/mod.rs with the factory smoke table updated; textfootage and
adjustment-layer i18n keys included.
2026-09-10 22:03:15 +08:00
Mike-Solar d028a45ffa nodes: pivot transform rotation/scale around the frame center
The transform shader sampled in a top-left-origin pixel space while
Olive's transform semantics (and every other node) are center-origin:
rotation swung the image around the top-left corner, pushing it partly
off-frame - reading exactly like an unwanted zoom. Match the C++
transform.vert projection: position (0,0) is the frame center and
rotation/scale pivot around the anchor, so rotation and scale stay
independent user controls. GPU tests pin the 90-degree landing spot
(no smearing) and the 2x scale centroid (stays centered).
2026-09-10 18:13:05 +08:00
Mike-Solar fc90604245 render: anchor resolution_in to the sequence resolution
Preview renders at proxy size while a paused frame renders full-res, so
anchoring resolution_in to the render target made every sequence-pixel
effect (shape size/pos, transform offsets, corner pin points, drop
shadow distance) change apparent size whenever the transport stopped.
Pre-fill resolution_in from the sequence's video params (C++ inserts
the NodeGlobals square resolution at job-build time), covering the
nested generator job inside a merge as well; a node that inserted its
own resolution_in keeps it.
2026-09-10 16:53:55 +08:00
Mike-Solar 37df1d3d34 nodes: spell shape/despill shader dispatch as if/else chains
CI / Build & test (Linux) (push) Successful in 24m7s
CI / Build & test (Windows) (push) Successful in 30m31s
naga's WGSL emitter rejects fall-through-capable GLSL switch blocks, so
every shape and despill job failed to compile and silently fell back to
the effect input - both effects were no-ops. Rewrite the type/method
dispatch as if/else chains (same semantics as the C++ shaders) and
cover all three shape types plus green-screen despill with GPU pixel
tests. Also drop the now-stale nested-payload/merge-binding TODO notes.
2026-09-09 16:51:13 +08:00
Mike-Solar 5ab12b937f render: real texture binding, generator layers and iteration feedback in shader passes
- process_shader_job: bind all texture params by name, recurse into nested
  shader payloads (depth cap 8), fall back to frame size without inputs
- run_effect: take iterative_input so dropshadow previous_iteration_in works
- merge: actually composite inputs; keyer mask, opacity modulation, math
  texture ops and mrg generator layers now bind their textures
- transform distort: real fragment-side inverse-matrix sampling
- time offset / time remap: wire NodeBehavior time adjustment hooks
- plugin: fix first-node identity colliding with unbound sentinel
2026-09-09 16:33:29 +08:00
Mike-Solar cde4dddf32 structure: move jobs out of nodes/ 2026-09-03 17:44:46 +08:00
Mike-Solar 4babbf5de8 core: merge oak-common into oak-core
CI / Build & test (Linux) (push) Successful in 24m6s
CI / Build & test (Windows) (push) Successful in 31m14s
oak-common is gone; its modules (configstore, xmlutils, ocioutils,
oiioutils, colormath, colortransform, videoparams, ffmpegutils, ...)
now live in oak-core alongside the value types. The render value/GPU
types moved too: backend (wgpu context + DisplayRenderer), color
(ColorProcessor over ocio-rs), texture, frame, and the commonutil
config helpers.

Fix-ups to make the merged tree build and pass tests:

- oak-core Cargo.toml: wgpu back to 25 (the moved backend code is
  written against that API generation); add the toml/quick-xml/image
  deps oak-common carried.
- lib.rs: drop the duplicate 'pub mod error;'.
- error.rs: unified OAKCORE_* codes; restore Error::new() and
  From<OcioError> from oak-common's error type.
- backend.rs/color.rs: oak_core::/oak_render:: self-references
  rewritten to crate::; the shaderfx-dependent GPU effect test moved
  to oak-render's shaderfx tests (shaderfx depends on oak-node and
  cannot live in oak-core).
- oak-render's error module re-exports oak_core::error::{Error,
  Result}; the OAKRENDER_* codes stay as the public-code contract.
- oak-node jobs.rs: ColorProcessor imported from oak_core::color.
- Integration tests repointed at oak_core::{texture, frame, backend,
  color, colormath}.
- the display-ICC regression test treats an empty OAK_DISPLAY_ICC as
  unset, matching displayicc::env_override_icc.
2026-09-03 17:42:20 +08:00
Mike-Solar 49fed365a4 nodes: real OCIO color grading (linear + log) on the GPU
The OCIO grading nodes previously pushed null texture handles; they now
push real ShaderJobPayloads whose GLSL is the OCIO-generated dynamic
grading-primary GPU shader — the exact code the C++ path applies, so no
approximation:

- color: grading_primary_function_shader(style) builds a dynamic
  GradingPrimaryTransform (LIN/LOG) on the default config, extracts the
  GLSL via GpuShaderDesc (function 'ove_grading_primary', resource
  prefix ocio_, no LUT textures) and caches it per style + config id.

- eval: OCIO_GRADING_STUBS maps the two node type ids to the grading
  style; process_shader_job resolves the stub and splices it into the
  node's %1 marker (same wiring as the chromakey OCIO stub); the
  pipeline cache key folds the stub text so a config change recompiles.

- nodes: value() pushes a ShaderJobPayload with the C++ value()
  rewrite applied to the row — vec4 (RGBM x=master) grading inputs to
  the vec3 GPU uniform form (lin: contrast RGB=c*m, offset RGB=c+m,
  exposure RGB=2^(c+m); log: lift RGB=c+m, gain c*m, gamma c*m), plus
  pivot/saturation floats, the log pivotBlack/pivotWhite normalization
  range (0/1), clamp sentinels (NoClampBlack -1 / NoClampWhite 2),
  white>black enforcement per frame, and localBypass=false. Generated
  uniform names bind by name (the log node's OCIO_NAMESPACE_ id text
  normalizes to the ocio_ resource prefix).

- Tests: grading stub generation (analytic GLSL, cache) in color,
  end-to-end GPU exposure doubling for lin (+1 stop on 0.2 gray -> 0.4)
  and lift for log, node payload rewrite assertions, and the
  all-shaders sweep now retries grading stubs. oak-render 181,
  oak-node 441, oak-app 272 lib tests pass.
2026-09-02 20:04:55 +08:00
Mike-Solar be42620e22 multicam: throttle angle refresh, shrink pre-render window
- MulticamPanel: playback angle refresh runs one cycle per 3 ticks
  instead of re-requesting every source every tick — each angle decode
  is a keyframe-scanning FFmpeg seek that stole worker capacity from
  the main viewer
- PreRender frames default 120 -> 12: a window larger than what the
  pool can render in real time queues far ahead of the playhead, so
  the painted frame lags seconds behind (playback frozen); a smaller
  window keeps the backlog bounded
- render_graph_frame: OAK_PERF clip-level timing
2026-09-01 21:52:56 +08:00
Mike-Solar 7f41570596 render: stop evicting the only hardware decoder session on every open
The per-process 'hardware session == 1, evict before every new open'
guard was turning every frame's open into a decoder re-open: after
inserting the fresh session, the NEXT frame's pre-open eviction dropped
it again, so no request ever hit the cache ([open] (request) on every
single frame, zero CACHED hits). Every decode then cost a full FFmpeg
session open (~0.5 s) + a keyframe-scanning seek — playback could never
keep up (the 'main viewer barely moves' report; [perf] showed 1.6-2.2 s
per frame).

Hardware VRAM is bounded by the LRU cap (hardware-first eviction at
MAX_CACHED_DECODERS) and the gpu-vram worker-count policy; the eager
pre-open eviction was the regression.
2026-09-01 21:46:01 +08:00
Mike-Solar 1b0a15192e multicam: wizard + AFV + node-graph preview + performance
CI / Build & test (Linux) (push) Failing after 24s
CI / Build & test (Windows) (push) Failing after 3m38s
- wizard: angle multi-select, sync modes, auto-align, create sequence;
  keeps the host sequence current (host clip = the multicam clip)
- build_multicam_sequence: per-angle clip -> SOURCES_INPUT[element],
  array slot growth, audio angle tracks for AFV
- AFV: host linked audio clip follows the switched source (one undo),
  muted host audio track disables it
- node-graph preview: inline producer renders through the traverser
  (viewer/project-matched graph frames); sequence viewer uses the graph
- multicam node value(): element-tagged row keys (sources_in[i]),
  reads the current source; build_row keys array inputs by element
- angle grid: viewer=0 (single-track montage, not whole-graph)
- project explorer: rename (dialog) + delete (undoable) real items
- performance: decoded-frame LRU, per-process NVDEC quota (1 session,
  evict before open), GPU composite fail-once fallback, snapshot
  upload debounced on the engine tick, worker vram budget headroom
- timeline clip: multicam overlay via ClipDecorator
- wizard menu item moved to Sequence menu
2026-09-01 20:27:59 +08:00
Mike-Solar 9352fca4a9 codec/render: silence hw-decode failures, vram-aware dynamic worker pool
- open_hw_accel marks the device unavailable when the decoder OPEN fails
  (cuvidCreateDecoder OOM at 4K) too, not just device-context creation:
  without it every subsequent decoder session retried CUDA and flooded
  the log per open.
- Decoder gains hardware_decoding(); the oak-render decode-session LRU
  evicts hardware sessions first (each pins a GPU surface pool — ~100 MB
  at 4K), so a full cache cannot exhaust video memory before the next
  open.
- Worker pool count now factors GPU vram: per-worker budget = 1 GiB
  (1080p peak) scaled by pixel ratio + 256 MiB idle floor, 10% reserve
  of free vram; applied when hardware decoding is on (nvidia-smi query,
  None otherwise falls back to the RAM/CPU policy).
- Dynamic pool resize: ProcessDispatcher::set_target_workers grows or
  retires workers; retiring ones stop claiming, drain their in-flight
  batch (future playback frames included), then exit naturally on the
  shutdown signal — no mid-work kill (30 s deadline only as a hung-
  decoder last resort). A retiring worker that dies re-queues its frames
  to surviving workers. Resizes are throttled to 2 s (a resolution burst
  merges; only the latest target applies) so 1080p<->4K flaps cannot
  thrash process spawns.
- RenderManager::set_workspace_size announces the sequence resolution;
  RealEngine calls it from refresh_sequence_info.
- Integration test: shrink 3->1 mid-wave (all frames complete, retired
  workers exit naturally) then regrow 1->3 and render a fresh wave.
2026-08-30 21:59:06 +08:00
Mike-Solar e128a6aca1 render: clamp mixed audio to [-1,1] before it reaches the device
Overlapping clips sum linearly in mix_audio_montage and can exceed full
scale (two hot clips reach +/-2; gain > 1 would too); the cpal sink
forwarded samples unclamped, so overlaps clipped at the DAC. Clamp the
accumulator after the montage mix (both the heap and shm-slot paths
share mix_audio_montage) and document it on render_audio_samples.
2026-08-29 19:51:25 +08:00
Mike-Solar cbd4ba2442 optimize: video and audio render
CI / Build & test (Linux) (push) Canceled after 0s
CI / Build & test (Windows) (push) Canceled after 0s
2026-08-29 05:01:55 +08:00
Mike-Solar fdb5caabd5 color: non-sRGB preview, per-monitor display ICC, pipeline hardening
Preview now follows the project output colorspace end to end: the
display chain derives its content space from the project's OutputColorSpec
instead of a hardcoded sRGB name, self-managed ICC transforms go through
an XYZ D65 interchange stage (OCIO cie_xyz_d65_interchange) for non-sRGB
targets, and the platform layer declares the content colorspace (gpui
submodule bump). macOS defaults to OS-managed (fixes wide-gamut UI
oversaturation); Windows ACM warns once on non-sRGB targets.

Multi-monitor: the display ICC is looked up per the window's current
screen (macOS display id, Windows per-monitor DC, X11 RandR output
profile) with a throttled poll that invalidates frame caches on moves.

Pipeline precision: 10-bit+ sources fall back to YUV444P16LE + a Rust
matrix conversion when swscale lacks F32 output (no more 8-bit
truncation); BT.709/2020 SDR decodes with BT.1886 gamma 2.4 instead of
the sRGB EOTF; working-space compositing no longer clamps RGB to [0,1]
(alpha still clamped); the output node clamps to the target gamut;
frames without colorimetry metadata convert with BT.709 defaults
(warned once) instead of passing through; scopes read the
output-colorspace signal on both F32 paths.

Also: only emit rerun-if-changed for .env when it exists (a missing file
made every build fully dirty).
2026-08-29 00:24:15 +08:00
Mike-Solar c03f1ec604 render: stack higher-numbered video tracks on top
Match the timeline UI (V_max drawn topmost): composite tracks from
V1 up to V_max so the highest-numbered track is composited last, in
both the montage path and direct graph evaluation.
2026-08-27 15:20:18 +08:00
Mike-Solar c51a349070 render: translate node GLSL shaders to WGSL and run them as wgpu passes
CI / Build & test (Linux) (push) Failing after 16m57s
CI / Build & test (Windows) (push) Successful in 31m46s
2026-08-27 07:07:09 +08:00
Mike-Solar fa3951344b render: fix 4K playback memory growth and decode-to-target-size
Root causes found for the 4K stalls and the second-footage memory
blowup (audit + code review):

- ticket bookkeeping leaked unbounded: the procpool ticket table and
  the arena slot map only ever grew (50-100 tickets/sec during
  playback, each pinning montage params and shm region views).
  Completed/cancelled/superseded/crashed entries are now removed, and
  the arena reaps fire-and-forget tickets once finished; the sync poll
  path reaps via a terminal result() read. InFlight duplicate submits
  now answer State immediately instead of sitting in the map forever.
- decode ran a full-resolution swscale to F32 RGBA (~132 MB at 4K)
  plus a second full-res copy before downscaling to the 480px proxy:
  RetrieveVideoParams.target_size lets swscale convert AND resize in
  one pass (bilinear, matching the old Rust resampler), so a 4K
  preview frame costs ~1 MB instead of ~260 MB of churn. This applies
  to proxy AND full-res requests alike.
- per-process decoder cache was unbounded (each session pins an FFmpeg
  context + 2 native decoded frames): LRU-capped at 16, eviction drops
  the map entry (in-flight renders keep their Arc; Drop releases
  FFmpeg).
- playback window completions were not generation-gated: a stale
  render from before an edit landed in the rebuilt window (wrong frame
  displayed, fresh request blocked). Stale completions now return
  their shm slot credit instead.
- async audio prefetch used the polling ticket submit without ever
  polling: switched to the fire-and-forget submit so entries reap.
2026-08-26 01:31:10 +08:00
Mike-Solar f2aab8ce15 render: apply clip effect stacks in the montage path
CI / Build & test (Windows) (push) Failing after 16m26s
CI / Build & test (Linux) (push) Successful in 19m12s
Adding an effect to a clip did nothing: the sequence render is
flattened into a montage (decode + composite), and MontageClip carried
no effect data at all.

- MontageClip gains an ordered effect stack (type id / enabled /
  effect input / parameter values); protocol v2 carries it as an
  additive wire field (older peers default to an empty stack).
- renderops::video_montage fills the stack from the effect chain
  (the footage source node — the chain end without an effect input —
  is dropped; the montage decodes the footage itself). Export
  (oak-task) and the multicam single-track montage fill it too.
- The worker applies the stack between decode and composite: built-in
  Opacity gets a CPU evaluator (C++ opacity.frag parity — whole vec4,
  alpha included, unity pass-through); everything else dispatches as an
  OFX plugin job through a new instance-factory slot (oak-plugin
  lazily creates + caches one instance per identifier per render
  process) with the montage's parameters injected. Disabled effects
  bypass (the C++ traverser pushes the effect input through). Unknown
  types warn once per type id and pass through — no silent no-ops.

Not covered (explicitly): Transform/Crop and the other ~30 built-in
effects have no CPU evaluator in oak-render (they pass through with a
warning), keyframed parameter animation, audio effect chains, and the
CLI's simplified montage.

Acceptance: a real 50% Opacity on real media quarters the rendered
pixels both in-process (renderops test) and through a real worker
process over IPC + shared memory (procpool_integration test);
disabling restores the plain render byte-for-byte.
2026-08-25 04:49:00 +08:00
Mike-Solar 244d5e860f workspace: kebab-case crates, app under crates/oak-app, shared versions
CI / Build & test (Windows) (push) Failing after 7s
All crates take the oak-* kebab-case naming (oak-audio, oak-codec,
oak-common, oak-core, oak-ffmpeg-link, oak-node, oak-otio, oak-plugin,
oak-render, oak-storage, oak-task, oak-timeline, oak-undo), with the
lib identifiers rewritten (oakrender:: -> oak_render::, oakcore_rs:: ->
oak_core::, ...) across all 226 referencing files.

The GUI application moves from the workspace root into
crates/oak-app/: src/, build.rs (paths fixed for the new location) and
tests/ travel with it, the root Cargo.toml becomes workspace-only
([workspace] + workspace.package + profiles), and the app package
inherits the workspace version. The screenshots example becomes a
standalone crate examples/simple_player/ with its own Cargo.toml.

Every crate now inherits the single workspace version
(version.workspace = true), and the workflows' crate paths and the
build docs follow the renames.

Validated with a clean cargo check --workspace.
2026-08-22 16:58:37 +08:00