Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
123
result(s) for
"Zhou, Wenlei"
Sort by:
Experimental and Numerical Simulation of the Resistance Characteristics and Desulfurization Efficiency of Rod-Shaped Turbulators in WFGD for Green Power Systems
by
Li, Xiangpeng
,
Yang, Likun
,
Liu, Xunliang
in
Coal-fired power plants
,
Design
,
desulfurization efficiency
2025
The Wet Flue Gas Desulfurization (WFGD) process has always been an important part in the low-carbon/green power realization process of traditional power plants. Adding a turbulator to the spray scrubber can improve the desulfurization efficiency, whereas it also increases the flow resistance. In this study, a small experiment device based on a spray scrubber with a turbulator in a power plant was built on a 1:10 scale to address the problem. The influence of the flue gas flow rate and the liquid–gas ratio on the flow resistance was investigated. A mathematical model was established for the two-phase flow and the SO2 liquid phase absorption reaction, and numerical simulations were achieved by the Fluent code. The resistance characteristics and the liquid droplet residence time were studied in detail. By fitting the experimental data, the relationship between the resistance coefficient, the Reynolds number, and the liquid–gas ratio in the tower was determined as the following: f = 0.0288 Re0.359(L/G)0.754. The desulfurization efficiency was calculated by adopting a user-defined function (UDF) code in a computational fluid dynamics (CFD) model, and the effects of the flue gas flow rate, temperature, and the liquid–gas ratio were analyzed. The results show that the effect of the rod-shaped turbulator on the flow resistance is much less than the effect of the liquid spray. The residence time of droplets around the turbulator is doubled. The pressure loss in the scrubber increases with the liquid–gas ratio (associated with the number of spray layer) and the flue gas flow rate. The turbulator can improve the uniformity of the flue gas velocity to some extent and increase the utilization rate of the spraying liquid, thereby increasing the desulfurization efficiency by 2.49%. Considering the operation cost, the reasonable value range of the liquid–gas ratio is 20~30. This work presents a good demonstration of combining the experiment and numerical simulations on a laboratory scale for large systems and associated components research, which is helpful for the engineering design and optimization of modern green power systems.
Journal Article
DDPM Simulation for Fluidization Behavior and Reduction of Iron Ore Fines with Hydrogen in the Fluidized Bed
by
Huo, Hailong
,
Su, Fuyong
,
Zhang, Sizong
in
Direct reduced iron
,
Ferric oxide
,
Fluidized bed reduction
2024
In this paper, the hydrogen direct reduction of iron ore fines is numerically studied by using the Dense Discrete Phase Model (DDPM) in the fluidized bed. The fluidization behavior at different inlet gas velocities (Ug) as well as the influence of Ug and hydrogen concentration on reduction degree (RD) are comprehensively investigated. The result indicates the increase of time-averaged solids volume fraction for the same cross-sectional heights with increasing Ug when the bed height (H) exceeds 0.06 m. Furthermore, the reduction rate of mineral powder increases with higher Ug value, and the RD reaches almost 100 pct after 4000 seconds of reduction time with Ug ranging from 0.35 to 0.65 m/s. The reduction rate increases noticeably with the increase of hydrogen concentration in the range of 10 to 100 pct, and Fe2O3 can be completely converted to Fe under condition of 65 pct H2 concentration after 4000 seconds. Moreover, higher H2 concentration leads to faster rate of Fe2O3 consumption and Fe production. The mass fraction peak values of Fe3O4 and FeO are in the range of 0.29 to 0.34 and 0.21 to 0.24 under different H2 concentrations, respectively.
Journal Article
Image Generators are Generalist Vision Learners
2026
Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language understanding and reasoning from generative pretraining. While it has long been conjectured that the ability to create visual content implies an ability to understand it, there has been limited evidence that generative vision models have developed strong understanding capabilities. In this work, we demonstrate that image generation training serves a role similar to LLM pretraining, and lets models learn powerful and general visual representations that enable SOTA performance on various vision tasks. We introduce Vision Banana, a generalist model built by instruction-tuning Nano Banana Pro (NBP) on a mixture of its original training data alongside a small amount of vision task data. By parameterizing the output space of vision tasks as RGB images, we seamlessly reframe perception as image generation. Our generalist model, Vision Banana, achieves SOTA results on a variety of vision tasks involving both 2D and 3D understanding, beating or rivaling zero-shot domain-specialists, including Segment Anything Model 3 on segmentation tasks, and the Depth Anything series on metric depth estimation. We show that these results can be achieved with lightweight instruction-tuning without sacrificing the base model's image generation capabilities. The superior results suggest that image generation pretraining is a generalist vision learner. It also shows that image generation serves as a unified and universal interface for vision tasks, similar to text generation's role in language understanding and reasoning. We could be witnessing a major paradigm shift for computer vision, where generative vision pretraining takes a central role in building Foundational Vision Models for both generation and understanding.
Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use
2024
From content moderation to wildlife conservation, the number of applications that require models to recognize nuanced or subjective visual concepts is growing. Traditionally, developing classifiers for such concepts requires substantial manual effort measured in hours, days, or even months to identify and annotate data needed for training. Even with recently proposed Agile Modeling techniques, which enable rapid bootstrapping of image classifiers, users are still required to spend 30 minutes or more of monotonous, repetitive data labeling just to train a single classifier. Drawing on Fiske's Cognitive Miser theory, we propose a new framework that alleviates manual effort by replacing human labeling with natural language interactions, reducing the total effort required to define a concept by an order of magnitude: from labeling 2,000 images to only 100 plus some natural language interactions. Our framework leverages recent advances in foundation models, both large language models and vision-language models, to carve out the concept space through conversation and by automatically labeling training data points. Most importantly, our framework eliminates the need for crowd-sourced annotations. Moreover, our framework ultimately produces lightweight classification models that are deployable in cost-sensitive scenarios. Across 15 subjective concepts and across 2 public image classification datasets, our trained models outperform traditional Agile Modeling as well as state-of-the-art zero-shot classification models like ALIGN, CLIP, CuPL, and large visual question-answering models like PaLI-X.
Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
2026
We introduce Gemini Embedding 2, a native multimodal embedding model that allows embedding video, audio, image, and text modalities in a unified representation space. We leverage the multimodal capabilities of Gemini to produce embeddings for arbitrary combinations of interleaved inputs across all these modalities that generalize well across a wide variety of tasks. Applying large-scale contrastive learning in a multi-task multi-stage training setup, we achieve state-of-the-art performance on key embedding benchmarks including unimodal, cross-modal, and multimodal retrieval spanning a diverse set of tasks. We show that our embedding model demonstrates strong performance (with a score of 62.9 R@1 on MSCOCO, 68.8 NDCG@10 on Vatex, 69.9 on MTEB multilingual and 84.0 on MTEB Code) across a variety of tasks surpassing the performance of specialized models. These unified capabilities make Gemini Embedding 2 a promising candidate for downstream use cases such as RAG, recommendation and search. Furthermore, its robust zero-shot performance across distinct fields - from astronomy and bioscience to fine arts and the culinary arts - establishes it as a highly reliable, out-of-the-box representation even for specialized domains.
Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
2024
We address the long-standing problem of how to learn effective pixel-based image diffusion models at scale, introducing a remarkably simple greedy growing method for stable training of large-scale, high-resolution models. without the needs for cascaded super-resolution components. The key insight stems from careful pre-training of core components, namely, those responsible for text-to-image alignment ıt vs. high-resolution rendering. We first demonstrate the benefits of scaling a ıt Shallow UNet, with no down(up)-sampling enc(dec)oder. Scaling its deep core layers is shown to improve alignment, object structure, and composition. Building on this core model, we propose a greedy algorithm that grows the architecture into high-resolution end-to-end models, while preserving the integrity of the pre-trained representation, stabilizing training, and reducing the need for large high-resolution datasets. This enables a single stage model capable of generating high-resolution images without the need of a super-resolution cascade. Our key results rely on public datasets and show that we are able to train non-cascaded models up to 8B parameters with no further regularization schemes. Vermeer, our full pipeline model trained with internal datasets to produce 1024x1024 images, without cascades, is preferred by 44.0% vs. 21.4% human evaluators over SDXL.
EmbeddingGemma: Powerful and Lightweight Text Representations
2025
We introduce EmbeddingGemma, a new lightweight, open text embedding model based on the Gemma 3 language model family. Our innovative training recipe strategically captures knowledge from larger models via encoder-decoder initialization and geometric embedding distillation. We improve model robustness and expressiveness with a spread-out regularizer, and ensure generalizability by merging checkpoints from varied, optimized mixtures. Evaluated on the Massive Text Embedding Benchmark (MTEB) across multilingual, English, and code domains, EmbeddingGemma (300M) achieves state-of-the-art results. Notably, it outperforms prior top models, both proprietary and open, with fewer than 500M parameters, and provides performance comparable to models double its size, offering an exceptional performance-to-cost ratio. Remarkably, this lead persists when quantizing model weights or truncating embedding outputs. This makes EmbeddingGemma particularly well-suited for low-latency and high-throughput use cases such as on-device applications. We provide ablation studies exploring our key design choices. We release EmbeddingGemma to the community to promote further research.
Agile Modeling: From Concept to Classifier in Minutes
2023
The application of computer vision to nuanced subjective use cases is growing. While crowdsourcing has served the vision community well for most objective tasks (such as labeling a \"zebra\"), it now falters on tasks where there is substantial subjectivity in the concept (such as identifying \"gourmet tuna\"). However, empowering any user to develop a classifier for their concept is technically difficult: users are neither machine learning experts, nor have the patience to label thousands of examples. In reaction, we introduce the problem of Agile Modeling: the process of turning any subjective visual concept into a computer vision model through a real-time user-in-the-loop interactions. We instantiate an Agile Modeling prototype for image classification and show through a user study (N=14) that users can create classifiers with minimal effort under 30 minutes. We compare this user driven process with the traditional crowdsourcing paradigm and find that the crowd's notion often differs from that of the user's, especially as the concepts become more subjective. Finally, we scale our experiments with simulations of users training classifiers for ImageNet21k categories to further demonstrate the efficacy.
Gemini Embedding: Generalizable Embeddings from Gemini
2025
In this report, we introduce Gemini Embedding, a state-of-the-art embedding model leveraging the power of Gemini, Google's most capable large language model. Capitalizing on Gemini's inherent multilingual and code understanding capabilities, Gemini Embedding produces highly generalizable embeddings for text spanning numerous languages and textual modalities. The representations generated by Gemini Embedding can be precomputed and applied to a variety of downstream tasks including classification, similarity, clustering, ranking, and retrieval. Evaluated on the Massive Multilingual Text Embedding Benchmark (MMTEB), which includes over one hundred tasks across 250+ languages, Gemini Embedding substantially outperforms prior state-of-the-art models, demonstrating considerable improvements in embedding quality. Achieving state-of-the-art performance across MMTEB's multilingual, English, and code benchmarks, our unified model demonstrates strong capabilities across a broad selection of tasks and surpasses specialized domain-specific models.
Insights on forming N,O-coordinated Cu single-atom catalysts for electrochemical reduction CO2 to methane
2021
Single-atom catalysts (SACs) are promising candidates to catalyze electrochemical CO
2
reduction (ECR) due to maximized atomic utilization. However, products are usually limited to CO instead of hydrocarbons or oxygenates due to unfavorable high energy barrier for further electron transfer on synthesized single atom catalytic sites. Here we report a novel partial-carbonization strategy to modify the electronic structures of center atoms on SACs for lowering the overall endothermic energy of key intermediates. A carbon-dots-based SAC margined with unique CuN
2
O
2
sites was synthesized for the first time. The introduction of oxygen ligands brings remarkably high Faradaic efficiency (78%) and selectivity (99% of ECR products) for electrochemical converting CO
2
to CH
4
with current density of 40 mA·cm
-2
in aqueous electrolytes, surpassing most reported SACs which stop at two-electron reduction. Theoretical calculations further revealed that the high selectivity and activity on CuN
2
O
2
active sites are due to the proper elevated CH
4
and H
2
energy barrier and fine-tuned electronic structure of Cu active sites.
Single-atom catalysts (SACs) are promising candidates to catalyze CO
2
reduction for the formation of high value hydrocarbons but most of the reactions yield CO. Here, the authors show a low-temperature calcining process to fabricate a carbon-dots-based SAC to efficiently convert CO
2
to methane.
Journal Article