Asset Details
MbrlCatalogueTitleDetail
Do you wish to reserve the book?
Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
by
Baldridge, Jason M
, Qian, Rui
, Rosston, Sarah
, Guo, Mandy
, Pont-Tuset, Jordi
, Yan, Jimmy
, Luo, Shixin
, Xu, Keyang
, Kajic, Ivana
, Fleet, David J
, Waters, Austin
, Zhou, Wenlei
, Parekh, Zarana
, Swersky, Kevin
, Bunner, Andrew
, Wang, Oliver
, Onoe, Yasumasa
, Nandwani, Henna
, Garg, Roopal
, Wang, Su
, Li, Yeqing
, Rashwan, Abdullah
, Walker, Trevor
, Vasconcelos, Cristina N
, Hongliang Fei
in
Alignment
/ Datasets
/ Greedy algorithms
/ High resolution
/ Image resolution
/ Pixels
/ Regularization
2024
Hey, we have placed the reservation for you!
By the way, why not check out events that you can attend while you pick your title.
You are currently in the queue to collect this book. You will be notified once it is your turn to collect the book.
Oops! Something went wrong.
Looks like we were not able to place the reservation. Kindly try again later.
Are you sure you want to remove the book from the shelf?
Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
by
Baldridge, Jason M
, Qian, Rui
, Rosston, Sarah
, Guo, Mandy
, Pont-Tuset, Jordi
, Yan, Jimmy
, Luo, Shixin
, Xu, Keyang
, Kajic, Ivana
, Fleet, David J
, Waters, Austin
, Zhou, Wenlei
, Parekh, Zarana
, Swersky, Kevin
, Bunner, Andrew
, Wang, Oliver
, Onoe, Yasumasa
, Nandwani, Henna
, Garg, Roopal
, Wang, Su
, Li, Yeqing
, Rashwan, Abdullah
, Walker, Trevor
, Vasconcelos, Cristina N
, Hongliang Fei
in
Alignment
/ Datasets
/ Greedy algorithms
/ High resolution
/ Image resolution
/ Pixels
/ Regularization
2024
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
Do you wish to request the book?
Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
by
Baldridge, Jason M
, Qian, Rui
, Rosston, Sarah
, Guo, Mandy
, Pont-Tuset, Jordi
, Yan, Jimmy
, Luo, Shixin
, Xu, Keyang
, Kajic, Ivana
, Fleet, David J
, Waters, Austin
, Zhou, Wenlei
, Parekh, Zarana
, Swersky, Kevin
, Bunner, Andrew
, Wang, Oliver
, Onoe, Yasumasa
, Nandwani, Henna
, Garg, Roopal
, Wang, Su
, Li, Yeqing
, Rashwan, Abdullah
, Walker, Trevor
, Vasconcelos, Cristina N
, Hongliang Fei
in
Alignment
/ Datasets
/ Greedy algorithms
/ High resolution
/ Image resolution
/ Pixels
/ Regularization
2024
Please be aware that the book you have requested cannot be checked out. If you would like to checkout this book, you can reserve another copy
We have requested the book for you!
Your request is successful and it will be processed during the Library working hours. Please check the status of your request in My Requests.
Oops! Something went wrong.
Looks like we were not able to place your request. Kindly try again later.
Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
Paper
Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
2024
Request Book From Autostore
and Choose the Collection Method
Overview
We address the long-standing problem of how to learn effective pixel-based image diffusion models at scale, introducing a remarkably simple greedy growing method for stable training of large-scale, high-resolution models. without the needs for cascaded super-resolution components. The key insight stems from careful pre-training of core components, namely, those responsible for text-to-image alignment ıt vs. high-resolution rendering. We first demonstrate the benefits of scaling a ıt Shallow UNet, with no down(up)-sampling enc(dec)oder. Scaling its deep core layers is shown to improve alignment, object structure, and composition. Building on this core model, we propose a greedy algorithm that grows the architecture into high-resolution end-to-end models, while preserving the integrity of the pre-trained representation, stabilizing training, and reducing the need for large high-resolution datasets. This enables a single stage model capable of generating high-resolution images without the need of a super-resolution cascade. Our key results rely on public datasets and show that we are able to train non-cascaded models up to 8B parameters with no further regularization schemes. Vermeer, our full pipeline model trained with internal datasets to produce 1024x1024 images, without cascades, is preferred by 44.0% vs. 21.4% human evaluators over SDXL.
Publisher
Cornell University Library, arXiv.org
Subject
MBRLCatalogueRelatedBooks
Related Items
Related Items
This website uses cookies to ensure you get the best experience on our website.