2501.00740
Existing image object-removal methods suffer from incomplete removal, incorrect content synthesis, and blurry regions, giving low success rates. The authors argue the root cause is a lack of high-qua…
A method for training a reliable image object remover by manufacturing high-quality paired training data rather than relying on the self-supervised in-painting paradigm that confuses object synthesis with background restoration. Starting from an initial remover trained on 60K open-source pairs, human feedback selects high-quality removal results; those selections train a discriminator that automates further pair filtering, and iterating this human-in-the-loop, semi-supervised loop grows a dataset beyond 200K pairs. Fine-tuning a pre-trained Stable Diffusion model on it yields RORem, achieving state-of-the-art object-removal reliability and image quality and improving success rate over prior methods by more than 18%.
Existing image object-removal methods suffer from incomplete removal, incorrect content synthesis, and blurry regions, giving low success rates. The authors argue the root cause is a lack of high-qua…