Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
167
result(s) for
"Hao, Yongtao"
Sort by:
Gaussian Semantic Segmentation Based on Color and Shape Deformation Fields
2026
Dynamic scene reconstruction has achieved significant milestones with the advent of 3D Gaussian Splatting (3DGS). However, extending this technology from geometric reconstruction to semantic understanding in dynamic environments remains a challenge. Existing methods often rely on external 2D trackers, which lead to temporal inconsistencies and semantic drift, or suffer from the high computational costs of high-dimensional feature fields. In this paper, we propose a novel framework, Gaussian Semantic Segmentation based on Color and Shape Deformation Fields (GSSBC), to address these issues. Building upon our GBC dynamic scene representation, we bind learnable semantic features to deformable Gaussian primitives. We introduce a spatiotemporal contrastive learning strategy guided by the Segment Anything Model (SAM) to enforce semantic consistency without explicit tracking. Furthermore, we employ a density-based clustering algorithm with label propagation to extract discrete object entities efficiently. Experimental results on the HyperNeRF and Neu3D datasets demonstrate that our method achieves superior segmentation accuracy and spatiotemporal stability compared to state-of-the-art approaches, enabling effective semantic understanding in complex dynamic scenes.
Journal Article
Time Series Prediction Based on Multi-Scale Feature Extraction
2024
Time series data are prevalent in the real world, particularly playing a crucial role in key domains such as meteorology, electricity, and finance. Comprising observations at historical time points, these data, when subjected to in-depth analysis and modeling, enable researchers to predict future trends and patterns, providing support for decision making. In current research, especially in the analysis of long time series, effectively extracting and integrating long-term dependencies with short-term features remains a significant challenge. Long-term dependencies refer to the correlation between data points spaced far apart in a time series, while short-term features focus on more recent changes. Understanding and combining these two features correctly are crucial for constructing accurate and reliable predictive models. To efficiently extract and integrate long-term dependencies and short-term features in long time series, this paper proposes a pyramid attention structure model based on multi-scale feature extraction, referred to as the MSFformer model. Initially, a coarser-scale construction module is designed to obtain coarse-grained information. A pyramid data structure is constructed through feature convolution, with the bottom layer representing the original data and each subsequent layer containing feature information extracted across different time step lengths. As a result, nodes higher up in the pyramid integrate information from more time points, such as every Monday or the beginning of each month, while nodes lower down retain their individual information. Additionally, a Skip-PAM is introduced, where a node only calculates attention with its neighboring nodes, parent node, and child nodes, effectively reducing the model’s time complexity to some extent. Notably, the child nodes refer to nodes selected from the next layer by skipping specific time steps. In this study, we not only propose an innovative time series prediction model but also validate the effectiveness of these methods through a series of comprehensive experiments. To comprehensively evaluate the performance of the designed model, we conducted comparative experiments with baseline models, ablation experiments, and hyperparameter studies. The experimental results demonstrate that the MSFformer model improves by 35.87% and 42.6% on the MAE and MSE indicators, respectively, compared to traditional Transformer models. These results highlight the outstanding performance of our proposed deep learning model in handling complex time series data, particularly in capturing long-term dependencies and integrating short-term features.
Journal Article
Gaussian Splatting-Based Color and Shape Deformation Fields for Dynamic Scene Reconstruction
2025
The 3DGS (3D Gaussian Splatting) series of works has achieved significant success in novel view synthesis, but further research is needed for dynamic scene reconstruction tasks. In this paper, we propose a new framework based on 3DGS for handling dynamic scene reconstruction problems involving color changes. Our approach employs a multi-stage training strategy combining motion and color deformation fields to accurately model dynamic geometry and appearance changes. Additionally, we design two modular components: the Dynamic Component for capturing motion variations and the Color Component for managing material and color changes. These components flexibly adapt to different scenes, enhancing our method’s versatility. Experimental results demonstrate that our method achieves real-time rendering at 80 FPS on an RTX 4090 and achieves higher reconstruction accuracy than baseline methods such as HexPlane and Deformable3DGS. Furthermore, it reduces training time by approximately 10%, indicating improved training efficiency. These quantitative results confirm the effectiveness of our approach in delivering high-fidelity 4D reconstruction of complex dynamic environments.
Journal Article
Attention re-alignment in multimodal large language models via intermediate-layer guidance
2026
Multimodal large language models (MLLMs) have achieved impressive performance in understanding and describing visual content, setting new state-of-the-art results on a variety of visual question answering (VQA) benchmarks. However, during decoding, these models often fail to attend to fine-grained visual details in the input image. Our analysis of intermediate attention layers reveals that MLLMs are not inherently incapable of perceiving target objects; rather, attention to visual details becomes diluted in deeper layers due to the dominance of language priors. To address this limitation, we propose a plug-and-play attention re-alignment module (ARA) that enhances suppressed visual grounding. ARA conducts a layer-wise analysis of the relative attention distribution of image-centric attention heads. It incorporates a confidence-aware layer selection mechanism based on attention peak and entropy, enabling the dynamic aggregation of attention maps from the most informative layers. These aggregated maps are subsequently leveraged to guide the generation of semantic masks, enabling the model to emphasize salient visual regions while suppressing irrelevant or noisy content. ARA can be seamlessly integrated into existing MLLMs and demonstrates consistent improvements across multiple VQA benchmarks, validating its effectiveness in enhancing visual detail sensitivity.
Journal Article
MRID: Modeling Radiological Image Differences for Disease Progression Reasoning via Multi-Task Self-Supervision
2026
Automated radiology report generation has become a prominent research topic in medical multimodal learning. However, most existing approaches primarily focus on single-image interpretation and rarely address the task of tracking disease progression across longitudinal chest X-rays. This task presents two major challenges: accurately localizing pathological changes between temporally paired images, and effectively translating visual difference representations into clinically meaningful textual descriptions. To address these challenges, we propose MRID (Modeling Radiological Image Differences for Disease Progression Reasoning), a multi-task self-supervised framework that follows a pretraining–finetuning paradigm. MRID leverages multiple complementary self-supervised objectives to jointly achieve (1) intra-modal spatial alignment of organs and pathological regions across image pairs, and (2) cross-modal semantic alignment between visual difference representations and radiology report embeddings. Furthermore, we introduce a simple yet effective data augmentation strategy to alleviate the imbalance of disease progression categories. Extensive experiments conducted on the Longitudinal-MIMIC and MS-CXR-T datasets demonstrate that MRID effectively captures fine-grained disease progression patterns. In addition, the proposed framework achieves competitive performance on single-image radiology report generation, further highlighting its strong capability in modeling chest X-ray semantics.
Journal Article
A New Pulse Coupled Neural Network (PCNN) for Brain Medical Image Fusion Empowered by Shuffled Frog Leaping Algorithm
by
Che, Wenliang
,
Huang, Chenxi
,
Tian, Ganxun
in
Algorithms
,
computed tomography image
,
Computer science
2019
Recent research has reported the application of image fusion technologies in medical images in a wide range of aspects, such as in the diagnosis of brain diseases, the detection of glioma and the diagnosis of Alzheimer's disease. In our study, a new fusion method based on the combination of the shuffled frog leaping algorithm (SFLA) and the pulse coupled neural network (PCNN) is proposed for the fusion of SPECT and CT images to improve the quality of fused brain images. First, the intensity-hue-saturation (IHS) of a SPECT and CT image are decomposed using a non-subsampled contourlet transform (NSCT) independently, where both low-frequency and high-frequency images, using NSCT, are obtained. We then used the combined SFLA and PCNN to fuse the high-frequency sub-band images and low-frequency images. The SFLA is considered to optimize the PCNN network parameters. Finally, the fused image was produced from the reversed NSCT and reversed IHS transforms. We evaluated our algorithms against standard deviation (SD), mean gradient (Ḡ), spatial frequency (SF) and information entropy (E) using three different sets of brain images. The experimental results demonstrated the superior performance of the proposed fusion method to enhance both precision and spatial resolution significantly.
Journal Article
A New Dynamic Path Planning Approach for Unmanned Aerial Vehicles
by
Pei, Hongbin
,
Liu, Yuchen
,
Huang, Chenxi
in
Aerospace engineering
,
Algorithms
,
Ant colony optimization
2018
Dynamic path planning is one of the key procedures for unmanned aerial vehicles (UAV) to successfully fulfill the diversified missions. In this paper, we propose a new algorithm for path planning based on ant colony optimization (ACO) and artificial potential field. In the proposed algorithm, both dynamic threats and static obstacles are taken into account to generate an artificial field representing the environment for collision free path planning. To enhance the path searching efficiency, a coordinate transformation is applied to move the origin of the map to the starting point of the path and in line with the source-destination direction. Cost functions are established to represent the dynamically changing threats, and the cost value is considered as a scalar value of mobile threats which are vectors actually. In the process of searching for an optimal moving direction for UAV, the cost values of path, mobile threats, and total cost are optimized using ant optimization algorithm. The experimental results demonstrated the performance of the new proposed algorithm, which showed that a smoother planning path with the lowest cost for UAVs can be obtained through our algorithm.
Journal Article
Traditional Chinese Medicine Knowledge Graph Construction Based on Large Language Models
2024
This study explores the use of large language models in constructing a knowledge graph for Traditional Chinese Medicine (TCM) to improve the representation, storage, and application of TCM knowledge. The knowledge graph, based on a graph structure, effectively organizes entities, attributes, and relationships within the TCM domain. By leveraging large language models, we collected and embedded substantial TCM–related data, generating precise representations transformed into a knowledge graph format. Experimental evaluations confirmed the accuracy and effectiveness of the constructed graph, extracting various entities and their relationships, providing a solid foundation for TCM learning, research, and application. The knowledge graph has significant potential in TCM, aiding in teaching, disease diagnosis, treatment decisions, and contributing to TCM modernization. In conclusion, this paper utilizes large language models to construct a knowledge graph for TCM, offering a vital foundation for knowledge representation and application in the field, with potential for future expansion and refinement.
Journal Article
A Stock Prediction Method Based on Multidimensional and Multilevel Feature Dynamic Fusion
2024
Stock price prediction has long been a topic of interest in academia and the financial industry. Numerous factors influence stock prices, such as a company’s performance, industry development, national policies, and other macroeconomic factors. These factors are challenging to quantify, making predicting stock price movements difficult. This paper presents a novel deep neural network framework that leverages the dynamic fusion of multi-dimensional and multi-level features for stock price prediction, which means we utilize fundamental trading data and technical indicators as multi-dimensional data and local and global multi-level information. Firstly, the model dynamically assigns weights to multi-dimensional features of stocks to capture the impact of each feature on stock prices. Next, it applies the Fourier transform to the global features to capture the long-term trends of the global environment and dynamically fuses these with local and global features of the stocks to capture the overall market environment’s impact on individual stocks. Finally, temporal features are captured using an attention layer and an RNN-based model, which incorporates historical price data to forecast future prices. Experiments on stocks from various industries within the Chinese CSI 300 index reveal that the proposed model outperforms traditional methods and other deep learning approaches in terms of stock price prediction. This paper proposes a method that facilitates the dynamic integration of multi-dimensional and multi-level features in an efficient manner and experimental results show that it improves the accuracy of stock price predictions.
Journal Article
Dynamic fusion of multi-source heterogeneous data using MOE mechanism for stock prediction
2025
Stock prices are influenced by numerous factors, including social media, news, and financial reports, serving as indicators of financial market dynamics. However, harnessing diverse information from different sources and structures to predict price trends remains challenging. In this paper, we propose a dual-stage deep learning model based on the Mixture-of-Expert (MoE) mechanism. In stage one, three distinct expert networks encode information about price movements, financial news, and investor sentiments through multi-source interaction attention. In stage two, a gated network dynamically fuses outputs, capturing temporal relationships in windowed data. Experimental results on the Chinese stock market demonstrate our model outperforms existing ones in forecasting tasks.
Journal Article