랭킹으로 돌아가기

thu-ml/prolificdreamer

Pythonml.cs.tsinghua.edu.cn/prolificdreamer/

ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation (NeurIPS 2023 Spotlight)

diffusion-modeldreamfusionnerftext-to-3dprolificdreamerstablediffusion
스타 성장
스타
1.6k
포크
43
주간 성장
이슈
20
1k1.5k
2023년 5월2024년 5월2025년 6월2026년 7월
아티팩트PyPIpip install prolificdreamer
README

ProlificDreamer

Official implementation of ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation, published in NeurIPS 2023 (Spotlight).

Installation

The codebase is built on stable-dreamfusion. For installation,

pip install -r requirements.txt

Training

ProlificDreamer includes 3 stages for high-fidelity text-to-3d generation.

# --------- Stage 1 (NeRF, VSD guidance) --------- #
# This costs approximately 27GB GPU memory, with rendering resolution of 512x512
CUDA_VISIBLE_DEVICES=0 python main.py --text "A pineapple." --iters 25000 --lambda_entropy 10 --scale 7.5 --n_particles 1 --h 512  --w 512 --workspace exp-nerf-stage1/
# If you find the result is foggy, you can increase the --lambda_entropy. For example
CUDA_VISIBLE_DEVICES=0 python main.py --text "A pineapple." --iters 25000 --lambda_entropy 100 --scale 7.5 --n_particles 1 --h 512  --w 512 --workspace exp-nerf-stage1/
# Generate with multiple particles. Notice that generating with multiple particles is only supported in Stage 1.
CUDA_VISIBLE_DEVICES=0 python main.py --text "A pineapple." --iters 100000 --lambda_entropy 10 --scale 7.5 --n_particles 4 --h 512  --w 512 --t5_iters 20000 --workspace exp-nerf-stage1/

# --------- Stage 2 (Geometry Refinement) --------- #
# This costs <20GB GPU memory
CUDA_VISIBLE_DEVICES=0 python main.py --text "A pineapple." --iters 15000 --scale 100 --dmtet --mesh_idx 0  --init_ckpt /path/to/stage1/ckpt --normal True --sds True --density_thresh 0.1 --lambda_normal 5000 --workspace exp-dmtet-stage2/
# If the results are with maney floaters, you can increase --density_thresh. Notice that the value of --density_thresh must be consistent in stage2 and stage3.
CUDA_VISIBLE_DEVICES=0 python main.py --text "A pineapple." --iters 15000 --scale 100 --dmtet --mesh_idx 0  --init_ckpt /path/to/stage1/ckpt --normal True --sds True --density_thresh 0.4 --lambda_normal 5000 --workspace exp-dmtet-stage2/

# --------- Stage 3 (Texturing, VSD guidance) --------- #
# texturing with 512x512 rasterization
CUDA_VISIBLE_DEVICES=0 python main.py --text "A pineapple." --iters 30000 --scale 7.5 --dmtet --mesh_idx 0  --init_ckpt /path/to/stage2/ckpt --density_thresh 0.1 --finetune True --workspace exp-dmtet-stage3/

We also provide a script that can automatically run these 3 stages.

bash run.sh gpu_id text_prompt

For example,

bash run.sh 0 "A pineapple."

Limitations: (1) Our work ultilizes the original Stable Diffusion without any 3D data, thus the multi-face Janus problem is prevalent in the results. Ultilizing text-to-image diffusion which has been finetuned on multi-view images will alleviate this problem. (2) If the results are not satisfactory, try different seeds. This is helpful if the results have a good quality but suffer from the multi-face Janus problem.

TODO List

  • Release our code.
  • Combine MVDream with VSD to alleviate the multi-face problem.

BibTeX

If you find our work useful for your project, please consider citing the following paper.

@inproceedings{wang2023prolificdreamer,
  title={ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation},
  author={Zhengyi Wang and Cheng Lu and Yikai Wang and Fan Bao and Chongxuan Li and Hang Su and Jun Zhu},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  year={2023}
}
관련 저장소
MoonInTheRiver/DiffSinger

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code

PythonPyPIMIT Licensetext-to-speechdiffusion-speedup
4.8k825
Alpha-VLLM/Lumina-T2X

Lumina-T2X is a unified framework for Text to Any Modality Generation

PythonPyPIMIT Licenseaigctransformer
2.2k95
wangkai930418/awesome-diffusion-categorized

collection of diffusion model papers categorized by their subareas

diffusionstable-diffusion
2.2k102
PKU-YuanGroup/Helios

Helios: Real Real-Time Long Video Generation Model

PythonPyPIApache License 2.0accelerationdiffusion
pku-yuangroup.github.io/Helios-Page
2k158
Janspiry/Palette-Image-to-Image-Diffusion-Models

Unofficial implementation of Palette: Image-to-Image Diffusion Models by Pytorch

PythonPyPIMIT Licenseddpmimage-restoration
1.8k239
Stability-AI/stable-virtual-camera

Stable Virtual Camera: Generative View Synthesis with Diffusion Models

PythonPyPIOtherdiffusion-modelimage-to-video
stable-virtual-camera.github.io
1.6k122
opendilab/awesome-diffusion-model-in-rl

A curated list of Diffusion Model in RL resources (continually updated)

Apache License 2.0deep-reinforcement-learningdiffusion-model
1.6k78
wenhaochai/StableVideo

[ICCV 2023] StableVideo: Text-driven Consistency-aware Diffusion Video Editing

PythonPyPIApache License 2.0aigccomputer-vision
rese1f.github.io/StableVideo/
1.4k88
rese1f/StableVideo

[ICCV 2023] StableVideo: Text-driven Consistency-aware Diffusion Video Editing

PythonPyPIApache License 2.0aigccomputer-vision
rese1f.github.io/StableVideo/
1.4k89
shivammehta25/Matcha-TTS

[ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching

Jupyter NotebookMIT Licensedeep-learningflow-matching
shivammehta25.github.io/Matcha-TTS/
1.3k212
thu-ml/Motus

Official code of Motus: A Unified Latent Action World Model

PythonPyPIApache License 2.0roboticsworld-model
motus-robotics.github.io/motus
1.2k70