NeurIPS 2026

ODDR

One-Step Deshadow Diffusion via Reward Guidance

Junseong Shin, Kijun Kim, Minseong Kim, Dongjin Kim, Tae Hyun Kim
Hanyang University, VILAB
Paper coming soon OpenReview coming soon arXiv →
1
denoising step
(vs. 20–25 for prior diffusion models)
0
real-world paired images
used for supervision
0
human annotations
to train the reward model
5.3×
faster inference
than SSR (20-step diffusion)
ODDR outperforms diffusion-based shadow removal baselines on unseen datasets while using a single denoising step
One step is all it takes. Compared with diffusion-based shadow removal models (DeS3, SSR), ODDR achieves the best PSNR/RMSE on unseen LRSS and UIUC datasets with a single denoising step (NFE = 1), running at 127 ms and 158 GFLOPs — over 5× faster and 14× lighter than SSR.

Recent advances in deep learning for shadow removal have significantly enhanced image quality and realism. However, most approaches rely on real-world paired datasets, which are costly to collect and often limited in scene diversity, leading to limited generalization. To address these limitations, we propose One-step Deshadow Diffusion via Reward guidance (ODDR), a new framework that achieves efficient and high-fidelity shadow removal without relying on real-world paired supervision. Our method begins with One-step Deshadow Diffusion (ODD), a baseline model trained on synthetic shadow data for efficient one-step shadow-free reconstruction. We further adapt ODD into ODDR using ShadowReward. In contrast to traditional, annotation-heavy approaches, ShadowReward is the first reward model for shadow removal trained entirely without human annotation. It learns to mimic human perceptual judgments by ranking synthetically generated images with controlled degradations, such as texture distortion and boundary artifacts. This reward-guided fine-tuning enables ODDR to close the synthetic-to-real domain gap. Extensive experiments show that ODD achieves strong performance without relying on real-world paired supervision, and ODDR further improves the results, narrowing the gap to fully supervised methods trained on real-world paired data while maintaining higher computational efficiency as a single-step model.

Train on synthetic shadows, refine on real ones with a reward.

Two-stage training pipeline

ODDR training pipeline: ODD is trained on synthetic shadows, then fine-tuned into ODDR on real shadow images guided by ShadowReward
(a) The baseline ODD is trained on synthetic shadow images (xsyn), optimizing ODD-LoRA and ODD-Conv while the pretrained Stable Diffusion weights stay frozen. (b) ODDR is then fine-tuned from ODD on real shadow images (xreal) — with no ground truth — guided by the frozen ShadowReward model rθ, optimizing only the ODDR-LoRA module.

Human-annotation-free preference data

Example degradation list with increasing severity used to train ShadowReward without human labels
We synthesize a perceptually ordered list {x(j)} for each clean image by applying controlled degradations (texture distortion, noise, chromatic shift, luminance change, boundary artifacts) of increasing severity to the shadow region. The ordering itself provides the ranking supervision — no human annotation is needed.

ShadowReward

ShadowReward architecture: frozen DINOv2 encoder, lightweight convolutional head, Top-K pooling, trained with a ListMLE ranking loss
ShadowReward encodes the shadow input and a candidate restoration with a frozen DINOv2 backbone, scores patches with a lightweight convolutional head, aggregates them via Top-K pooling, and is trained with a ListMLE ranking loss on the ordered degradation lists — the first reward model for shadow removal trained entirely without human annotation.
Why does this work? The alignment loss anchors ODDR to the frozen ODD pseudo-target, while the reward term is the only gradient that moves the model beyond ODD. Reward-guided fine-tuning on unpaired real shadow images therefore directly targets the synthetic-to-real gap — improving shadow regions while preserving non-shadow content.

Clean shadow regions, untouched everything else.

Visual comparison on AISTD

Qualitative comparison on the AISTD benchmark
ODDR removes shadows cleanly while preserving non-shadow regions, whereas competing methods leave residual shadows or alter regions that should stay unchanged.

Generalization to unseen datasets (SRD / LRSS / UIUC)

Qualitative comparison on unseen SRD, LRSS and UIUC datasets
Without ever training on these datasets, ODDR consistently removes both hard and soft shadows while preserving natural color and structure — demonstrating strong robustness to domain shifts.

Effect of reward-guided fine-tuning (ODD → ODDR)

Comparison between ODD and reward-fine-tuned ODDR
ODDR corrects the residual artifacts and color shifts of ODD, yielding cleaner shadow regions and more consistent illumination across the entire image.
@inproceedings{shin2026oddr,
    title={ODDR: One-Step Deshadow Diffusion via Reward Guidance},
    author={Junseong Shin and Kijun Kim and Minseong Kim and Dongjin Kim and Tae Hyun Kim},
    booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
    year={2026}
}