JiaKui Hu\, Yuxiao Yang\, Jialun Liu, Jinbo Wu, Chen Zhao, Yanye Lu
Fonte do modelo
Trecho da fonte
JiaKui Hu, Yuxiao Yang, Jialun Liu, Jinbo Wu, Chen Zhao, Yanye Lu
Fontes
1 fonteVerificado 1 de ago.
Artefatos de modelo
3 artefatosTrechos de fonte
2 trechospt · 630 MB · SHA-256 2e0aa79ccd82…2af9 · Hugging Face
--- license: cc-by-nc-4.0 pipeline_tag: text-to-3d tags: - multi-view generation - auto-regressive --- # Auto-Regressively Generating Multi-View Consistent Images [JiaKui Hu](https://jkhu29.github.io/)\*, [Yuxiao Yang](https://yuxiaoyang23.github.io/)\*, [Jialun Liu](https://scholar.google.com/citations?user=OkMMP2AAAAAJ), [Jinbo Wu](https://scholar.google.com/citations?user=9OecN2sAAAAJ), [Chen Zhao](), [Yanye Lu](https://scholar.google.com/citations?user=WSFToOMAAAAJ) <br>PKU, BaiduVis, THU<br> ## Introduction  Diffusion-based multi-view image generation methods use a specific reference view for predicting subsequent views, which becomes problematic when overlap between the reference view and the predicted view is minimal, affecting image quality and multi-view consistency. Our MV-AR addresses this by using the preceding view with significant overlap for conditioning. ## Results ### Text to Multiview images  ### Image to Multiview images  ### Text + Geometric to Multiview images  ## Quick Start ### Requirements > Please follow the instructions in [code](https://github.com/MILab-PKU/MVAR). ### Reproduce 1. Please download [flan-t5-xl](https://huggingface.co/google/flan-t5-xl) in `./pretrained_models`; 2. Please download [Cap3D_automated_Objaverse_full.csv](https://huggingface.co/datasets/tiange/Cap3D/blob/main/Cap3D_automated_Objaverse_full.csv) in `dataset/captions`; 3. Please download models here, put them in `./pretrained_models`; 4. Run: ```shell # For t2mv on objaverse sh sample_tcam2i.sh # For t2mv on GSO sh sample_icam2i_gso.sh # For i2mv on GSO sh sample_icam2i_gso.sh ``` The generated images will be saved to `samples_objaverse_nv_ray/`. ## Acknowledgement This repository is heavily based on [LlamaGen](https://github.com/FoundationVision/LlamaGen). We would like to thank the authors of these work for publicly releasing their code. For help or issues using this git, please feel free to submit a GitHub issue. For other communications relate...
Source context: 0 downloads · 1 likes · Pipeline text-to-3d · Repo Jiakui/MV-AR