NDSplat
Under review

FactorSplat: Appearance-Controllable Gaussian Proxies for Medical Volume Rendering

United Imaging Intelligence, Boston MA, USA
zhongpai.gao@uii-ai.com *Corresponding author
+1.52 dB
PSNR over VEG · interpolation
524 FPS
At 1600² (B200)
1.17 ms
Per TF switch
FactorSplat, the path-traced reference, and VEG under two vascular transfer-function edits

Change appearance without retraining. Two vascular transfer-function edits: highlight selected tissue (top) and change tissue colors (bottom). FactorSplat (left) tracks the path-traced reference (middle); VEG (right) misses the edit. Full-frame PSNR.

Interactive Demo

A live session in the FactorSplat viewer. Stepping through transfer-function presets on the intestine, lower-body, and heart scans reuses one trained checkpoint per scan, with no retraining between presets.

Abstract

Transfer functions (TFs) control color and visibility in medical volume rendering, but image-trained Gaussian proxies typically bake one transfer function into their appearance. We present FactorSplat, a per-scene N-dimensional Gaussian splatting (N-DGS) proxy that accepts region-specific intensity-to-RGBA curves at inference.

A local lookup applies the authored color and opacity change, while a shared functional encoder and low-rank per-Gaussian factors learn the residual appearance response. Geometry and directional appearance remain shared across presets, with visibility control and TF-aware pruning preserving the ability to hide and reveal structures.

On seven CT and MR scans, FactorSplat improves mean PSNR and changed-region error over region-aware VEG across validation, interpolation, unseen composition, and out-of-distribution (OOD) edits. Across these four splits, seven-scan mean PSNR gains over VEG range from 1.10 to 1.52 dB. One checkpoint per scan supports unseen edits without retraining. At 1600², the cached fast renderer averages 524 FPS with 1.17 ms TF switches.

How It Works

The TF is a function, not a preset ID, and its effect differs across the scan. FactorSplat reads each Gaussian's local region and intensity samples, then updates only its view-independent color and opacity. Geometry and the N-DGS view response stay shared across every preset.

FactorSplat method overview: inputs, local conditioning, one model for many transfer functions
Direct · physical lookup

Authored TF response

Weighted local samples apply the authored color and opacity change directly, centred on a base TF so the base preset is reproduced exactly.

Learned · functional residual

Appearance correction

A shared pointwise encoder and rank-8 per-Gaussian factors learn how the cinematic renderer's image departs from the lookup.

Coverage · visibility gate

Hide and reveal structures

An analytic visibility gate and TF-aware pruning keep primitives that any training preset needs, so hidden anatomy is still there to reveal.

Quantitative Results

Seven scans: five CT and two MR. Each has 41 TF conditions: 24 training presets, then held-out validation, interpolation, unseen composition, and five OOD edits. Means weight scans equally within each modality.

Model ΔTF↓ (×10−2) PSNR↑ FPS↑Train (min)↓
valinterpcompOOD valinterpcompOOD
CT5 scansN-DGS (identity floor) 3.606.498.9617.7431.1127.8825.5021.118516.9
N-DGS (specialist) —30.7030.8231.0931.9486692.8
VEG 1.402.553.059.1531.8030.9030.6425.3713444.3
VEG (native-quantized) 1.412.563.069.1731.6530.7830.5425.3213644.3
FactorSplat (Ours) 1.362.172.888.0933.0832.2931.3326.5850810.5
MR2 scansN-DGS (identity floor) 2.884.255.4811.0936.5835.5431.1727.869256.0
N-DGS (specialist) —37.0337.0437.0738.0093076.4
VEG 1.191.891.595.1638.8238.5237.8231.2416942.6
VEG (native-quantized) 1.191.901.595.1538.6738.3937.7031.2017142.6
FactorSplat (Ours) 1.091.621.284.6540.6840.3639.9533.135628.3

ΔTF = changed-region edit error between predicted and reference image differences from the base TF. The identity floor has no TF input; specialists train one model per target TF, so they are preset-specific references rather than upper bounds. FPS on B200 at 1600². Bold = best per column.

Ablation: every branch earns its place

Removing the physical lookup costs the most. The learned residual helps on all eight columns, the functional code matters for edits far from training presets, and the gate helps when material is removed.

Variant (PSNR↑) heart vascular robustness
valint.cmpOOD valint.cmpOOD rank↓gap (dB)↓
FactorSplat (Full) 30.7929.3328.6823.77 34.2033.6832.5930.88 1.690.48
− exact gate 30.7929.3328.6823.45 34.2033.6832.5930.36 1.940.52
− learned residual 30.1628.6927.6123.22 33.7032.5531.9229.28 3.621.60
− physical lookup 29.7727.4826.1120.30 33.3132.1329.9429.18 4.883.47
− functional code 30.7628.3828.4222.75 34.3033.6833.0728.04 2.882.84

Rank 8. Rank and worst gap are over the eight scene–split comparisons.

Qualitative Comparison

Held-out cameras and TFs against the N-DGS identity floor, per-TF N-DGS specialists, and VEG. Full-frame PSNR under each image; bold marks each row's best.

Held-out CT and MR examples: heart, vascular, nose, hand

CT and MR: heart (rows 1–3), vascular (4–5), nose (6), hand (7).

Additional held-out CT examples: vascular, heart, intestine, lower body

CT: vascular (rows 1–3), heart (4–5), intestine (6), lower body (7).