Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
4 result(s) for "T2T-VIT"
Sort by:
An Improved SAR Ship Classification Method Using Text-to-Image Generation-Based Data Augmentation and Squeeze and Excitation
Synthetic aperture radar (SAR) plays a crucial role in maritime surveillance due to its capability for all-weather, all-day operation. However, SAR ship recognition faces challenges, primarily due to the imbalance and inadequacy of ship samples in publicly available datasets, along with the presence of numerous outliers. To address these issues, this paper proposes a SAR ship classification method based on text-generated images to tackle dataset imbalance. Firstly, an image generation module is introduced to augment SAR ship data. This method generates images from textual descriptions to overcome the problem of insufficient samples and the imbalance between ship categories. Secondly, given the limited information content in the black background of SAR ship images, the Tokens-to-Token Vision Transformer (T2T-ViT) is employed as the backbone network. This approach effectively combines local information on the basis of global modeling, facilitating the extraction of features from SAR images. Finally, a Squeeze-and-Excitation (SE) model is incorporated into the backbone network to enhance the network’s focus on essential features, thereby improving the model’s generalization ability. To assess the model’s effectiveness, extensive experiments were conducted on the OpenSARShip2.0 and FUSAR-Ship datasets. The performance evaluation results indicate that the proposed method achieves higher classification accuracy in the context of imbalanced datasets compared to eight existing methods.
Neonatal jaundice detection using a vision transformer-based deep learning model
Neonatal jaundice is a prevalent and potentially serious condition that can lead to severe complications if undiagnosed or untreated. While traditional diagnostic methods like blood sampling are invasive and time-consuming, and transcutaneous bilirubinometers remain costly, smartphone-based image analysis offers a promising low-cost, non-invasive alternative. However, most existing solutions rely on traditional machine learning techniques with limited accuracy and generalizability. In this study, we introduce a deep learning approach based on the Vision Transformer (T2T-ViT) and compare its performance with three other models, ResNet, Support Vector Machine (SVM), and K-Nearest Neighbors (k-NN), using a clinically annotated dataset of neonatal skin images captured via a smartphone camera. The models were evaluated using multiple performance metrics including accuracy, precision, recall, F1-score, Matthews Correlation Coefficient (MCC), and Area under the Curve (AUC). The T2T-ViT model achieved 99% across all metrics, significantly outperforming both convolutional and traditional machine learning models. These findings demonstrate the feasibility of applying transformer-based deep learning architectures for accessible, scalable, and accurate non-invasive neonatal jaundice screening, potentially enabling early intervention in resource-limited settings. This approach could serve as an accessible, scalable screening tool for neonatal jaundice detection, particularly in low-resource clinical settings.
A Method for the Extraction of Apocynum venetum L. Spatial Distribution in Yuli County, Xinjiang, via an Improved SegFormer Network
Efficient and accurate acquisition of spatial distribution information for Apocynum venetum L. is highly important for the sustainable development of agriculture in Yuli County, Xinjiang. As an important cash crop, Apocynum relies on specific natural conditions for growth, and its survival environment is currently under severe threat. Therefore, accurately quantifying its spatial distribution information is crucial. This research takes Yuli County in Xinjiang as the study area and proposes an enhanced SegFormer model based on deep learning, aiming to realize the effective identification and extraction of Apocynum. The study indicates the following. (1) The improved SegFormer model adds smaller-scale feature layers in the encoder stage, allowing the improved model’s encoder to extract features at five scales: 1/4, 1/8, 1/16, 1/32, and 1/64; meanwhile, integrating the T2T-ViT backbone network into the encoder significantly enhances the precision and efficiency of Apocynum’s spatial distribution extraction. (2) Compared with Unet, TransUNet, and the original SegFormer, the improved SegFormer model outperforms the other models in terms of the mIoU, OA, and mPA metrics, achieving values of 88.22%, 93.98%, and 89.66%, respectively. (3) Ablation experiments show that the T2T_vit_14 model performs best among all the T2T-ViT configurations, with superior extraction effects on fragmented small plots compared with the other models. Therefore, the T2T_vit_14 model is integrated into the SegFormer model. This work improves the extraction accuracy and efficiency of the spatial distribution of Apocynum via an improved SegFormer model, which has strong stability and robustness and offers scientific evidence for resource protection, restoration planting, and germplasm breeding in Yuli County, Xinjiang.
Photovoltaic power prediction based on sky images and tokens-to-token vision transformer
Photovoltaic (PV) power generation has high uncertainties due to the randomness and imbalance nature of solar energy and meteorological parameters. Hence, accurate PV power forecasts are essential in the operation of PV power plants (PVPP) for short-term dispatches and power generation schedules. In this paper, a new deep neural network structure based on vision transformer is proposed to combine sky images and Tokens-To-Token(T2T) for photovoltaic power prediction. The method uses an incremental tokenization module to aggregate neighboring image patches into tokens, which capture the local structural information of the clouds. Then, an efficient T2T-ViT backbone network is used to extract the global attentional relationships of the tokens for power prediction. In order to evaluate the performance of the proposed model, the method was compared with several deep learning architectures such as ResNet and GoogleNet on a dataset collected by the National Renewable Energy Laboratory in Colorado, USA. The results of power prediction were analysed using training loss, prediction error, and linear regression, and they show that the proposed method achieves higher prediction accuracy and lower error compared to the existing methods, especially in short- and ultra-short-term prediction. The paper demonstrates the potential of applying Transformer models to computer vision tasks for renewable energy forecasting. The results show that the proposed method achieves higher prediction accuracy and lower error than several deep learning architectures, such as ResNet and GoogleNet, especially in short- and ultra-short-term prediction.