JiaKui Hu\, Yuxiao Yang\, Jialun Liu, Jinbo Wu, Chen Zhao, Yanye Lu
Model source
Source excerpt
JiaKui Hu, Yuxiao Yang, Jialun Liu, Jinbo Wu, Chen Zhao, Yanye Lu
Sources
1 sourceVerified Aug 1
Model artifacts
3 artifactsSource excerpts
2 excerptspt · 630 MB · SHA-256 2e0aa79ccd82…2af9 · Hugging Face
--- license: cc-by-nc-4.0 pipeline_tag: text-to-3d tags: - multi-view generation - auto-regressive --- # Auto-Regressively Generating Multi-View Consistent Images [JiaKui Hu](https://jkhu29.github.io/)\*, [Yuxiao Yang](https://yuxiaoyang23.github.io/)\*, [Jialun Liu](https://scholar.google.com/citations?user=OkMMP2AAAAAJ), [Jinbo Wu](https://scholar.google.com/citations?user=9OecN2sAAAAJ), [Chen Zhao](), [Yanye Lu](https://scholar.google.com/citations?user=WSFToOMAAAAJ) <br>PKU, BaiduVis, THU<br> ## Introduction  Diffusion-based multi-view image generation methods use a specific reference view for predicting subsequent views, which becomes problematic when overlap between the reference view and the predicted view is minimal, affecting image quality and multi-view consistency. Our MV-AR addresses this by using the preceding view with significant overlap for conditioning. ## Results ### Text to Multiview images  ### Image to Multiview images  ### Text + Geometric to Multiview images  ## Quick Start ### Requirements > Please follow the instructions in [code](https://github.com/MILab-PKU/MVAR). ### Reproduce 1. Please download [flan-t5-xl](https://huggingface.co/google/flan-t5-xl) in `./pretrained_models`; 2. Please download [Cap3D_automated_Objaverse_full.csv](https://huggingface.co/datasets/tiange/Cap3D/blob/main/Cap3D_automated_Objaverse_full.csv) in `dataset/captions`; 3. Please download models here, put them in `./pretrained_models`; 4. Run: ```shell # For t2mv on objaverse sh sample_tcam2i.sh # For t2mv on GSO sh sample_icam2i_gso.sh # For i2mv on GSO sh sample_icam2i_gso.sh ``` The generated images will be saved to `samples_objaverse_nv_ray/`. ## Acknowledgement This repository is heavily based on [LlamaGen](https://github.com/FoundationVision/LlamaGen). We would like to thank the authors of these work for publicly releasing their code. For help or issues using this git, please feel free to submit a GitHub issue. For other communications relate...
Source context: 0 downloads · 1 likes · Pipeline text-to-3d · Repo Jiakui/MV-AR