Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL
Paper • 2605.24001 • Published
How to use Junyi-W/DIDR with Diffusers:
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("Junyi-W/DIDR", dtype=torch.bfloat16, device_map="cuda")
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL (NeurIPS 2026), Junyi Wu, Weijian Luo, Haoyang Zheng, Ruizhe Zhang, Guang Lin
Feel free to contact us if you have any questions about the paper!
Junyi Wu [email protected]
We can use the standard diffusers pipeline:
import torch
from diffusers import DiffusionPipeline, UNet2DConditionModel, LCMScheduler
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
base_model_id = "stabilityai/stable-diffusion-xl-base-1.0"
repo_name = "Junyi-W/DIDR"
ckpt_name = "didr_sdxl_1step_unet_fp16.safetensors"
# Load model.
unet = UNet2DConditionModel.from_config(UNet2DConditionModel.load_config(base_model_id, subfolder="unet")).to("cuda", torch.float16)
unet.load_state_dict(load_file(hf_hub_download(repo_name, ckpt_name), device="cuda"))
pipe = DiffusionPipeline.from_pretrained(base_model_id, unet=unet, torch_dtype=torch.float16, variant="fp16").to("cuda")
pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config)
prompt="a photo of a cat"
image=pipe(prompt=prompt, num_inference_steps=1, guidance_scale=0, timesteps=[399]).images[0]
import torch
from diffusers import DiffusionPipeline, UNet2DConditionModel, LCMScheduler
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
base_model_id = "stabilityai/stable-diffusion-xl-base-1.0"
repo_name = "Junyi-W/DIDR"
ckpt_name = "didr_longer_sdxl_1step_unet_fp16.safetensors"
# Load model.
unet = UNet2DConditionModel.from_config(UNet2DConditionModel.load_config(base_model_id, subfolder="unet")).to("cuda", torch.float16)
unet.load_state_dict(load_file(hf_hub_download(repo_name, ckpt_name), device="cuda"))
pipe = DiffusionPipeline.from_pretrained(base_model_id, unet=unet, torch_dtype=torch.float16, variant="fp16").to("cuda")
pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config)
prompt="a photo of a cat"
image=pipe(prompt=prompt, num_inference_steps=1, guidance_scale=0, timesteps=[399]).images[0]
import torch
from diffusers import ZImagePipeline, ZImageTransformer2DModel
base_model_id = "Tongyi-MAI/Z-Image-Turbo"
repo_name = "Junyi-W/DIDR"
# Load model.
transformer = ZImageTransformer2DModel.from_pretrained(repo_name, subfolder="zimage_transformer", torch_dtype=torch.bfloat16)
pipe = ZImagePipeline.from_pretrained(base_model_id, transformer=transformer, torch_dtype=torch.bfloat16).to("cuda")
prompt="A photorealistic portrait of a young woman in traditional Chinese red Hanfu, intricate golden embroidery, dramatic lighting, ultra high definition"
image=pipe(prompt=prompt, height=1024, width=1024, num_inference_steps=1, guidance_scale=0.0).images[0]
For more information, please refer to the code repository
DIDR is released under Creative Commons Attribution-NonCommercial 4.0 International License.
If you find DIDR useful or relevant to your research, please kindly cite our paper:
@inproceedings{wu2026didr,
title={Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL},
author={Wu, Junyi and Luo, Weijian and Zheng, Haoyang and Zhang, Ruizhe and Lin, Guang},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2026}
}
Our SDXL generators are initialized from DMD2 and our Z-Image generators from Z-Image-Turbo. We thank the authors for releasing their models.
Base model
Tongyi-MAI/Z-Image-Turbo