Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
68
result(s) for
"Jiang Yongyao"
Sort by:
Is ChatGPT a Good Geospatial Data Analyst? Exploring the Integration of Natural Language into Structured Query Language within a Spatial Database
2024
With recent advancements, large language models (LLMs) such as ChatGPT and Bard have shown the potential to disrupt many industries, from customer service to healthcare. Traditionally, humans interact with geospatial data through software (e.g., ArcGIS 10.3) and programming languages (e.g., Python). As a pioneer study, we explore the possibility of using an LLM as an interface to interact with geospatial datasets through natural language. To achieve this, we also propose a framework to (1) train an LLM to understand the datasets, (2) generate geospatial SQL queries based on a natural language question, (3) send the SQL query to the backend database, (4) parse the database response back to human language. As a proof of concept, a case study was conducted on real-world data to evaluate its performance on various queries. The results show that LLMs can be accurate in generating SQL code for most cases, including spatial joins, although there is still room for improvement. As all geospatial data can be stored in a spatial database, we hope that this framework can serve as a proxy to improve the efficiency of spatial data analyses and unlock the possibility of automated geospatial analytics.
Journal Article
Comparative Analysis of BERT and GPT for Classifying Crisis News with Sudan Conflict as an Example
2025
To obtain actionable information for humanitarian and other emergency responses, an accurate classification of news or events is critical. Daily news and social media are hard to classify based on conveyed information, especially when multiple categories of information are embedded. This research used large language models (LLMs) and traditional transformer-based models, such as BERT, to classify news and social media events using the example of the Sudan Conflict. A systematic evaluation framework was introduced to test GPT models using Zero-Shot prompting, Retrieval-Augmented Generation (RAG), and RAG with In-Context Learning (ICL) against standard and hyperparameter-tuned bert-based and bert-large models. BERT outperformed GPT in F1-score and accuracy for multi-label classification (MLC) while GPT outperformed BERT in accuracy for Single-Label classification from Multi-Label Ground Truth (SL-MLG). The results illustrate that a larger model size improves classification accuracy for both BERT and GPT, while BERT benefits from hyperparameter tuning and GPT benefits from its enhanced contextual comprehension capabilities. By addressing challenges such as overlapping semantic categories, task-specific adaptation, and a limited dataset, this study provides a deeper understanding of LLMs’ applicability in constrained, real-world scenarios, particularly in highlighting the potential for integrating NLP with other applications such as GIS in future conflict analyses.
Journal Article
Detection of Salmonella spp. Related Co-Infections Among Children with Diarrheal Diseases in Guangzhou, China
2025
Diarrheal diseases caused by gastrointestinal pathogens contribute to the high morbidity and mortality in children worldwide.
infection is one of the leading causes of diarrhea, especially in children under 5 years of age. This study aimed to assess the prevalence of
infection and its co-infection patterns in relation to clinical symptoms.
A total of 430 stool samples of children with diarrheal diseases were collected from Guangdong Women and Children's Hospital during January 2022 to December 2023 and used for detection. BioFire FilmArray Gastrointestinal (GI) Panel is an efficient and sensitive method used to assess infections caused by enteric pathogens simultaneously based on multiplex polymerase chain reaction (PCR) technologies.
spp. was classified as the predominant pathogen in all stool specimens, with an overall positivity rate of 36.74% (158/430), of which 35.44% (56/158) were identified as
single infection. For
-related bacterial co-infection pattern,
and enteropathogenic
combinations accounted for 12.66% (20/158) of
-related bacterial co-infection patterns, and
plus
was found in 8.86% (14/158). For
-viral co-infection pattern, the most prevalent combination was
and adenovirus (6.33%, 10/158). Notably, the proportion of mucus stools recorded in
plus
infections was statistically higher than that in single
infections (
< 0.05).
This study provides a comprehensive understanding of the nature of
spp. co-infections in diarrheal diseases, and the possibility of clinical symptoms and enhanced treatments.
Journal Article
A Query Understanding Framework for Earth Data Discovery
by
Yang, Chaowei
,
Mcgibbney, Lewis J.
,
Li, Yun
in
Archives & records
,
big spatiotemporal data analytics
,
cybergis
2020
One longstanding complication with Earth data discovery involves understanding a user’s search intent from the input query. Most of the geospatial data portals use keyword-based match to search data. Little attention has focused on the spatial and temporal information from a query or understanding the query with ontology. No research in the geospatial domain has investigated user queries in a systematic way. Here, we propose a query understanding framework and apply it to fill the gap by better interpreting a user’s search intent for Earth data search engines and adopting knowledge that was mined from metadata and user query logs. The proposed query understanding tool contains four components: spatial and temporal parsing; concept recognition; Named Entity Recognition (NER); and, semantic query expansion. Spatial and temporal parsing detects the spatial bounding box and temporal range from a query. Concept recognition isolates clauses from free text and provides the search engine phrases instead of a list of words. Name entity recognition detects entities from the query, which inform the search engine to query the entities detected. The semantic query expansion module expands the original query by adding synonyms and acronyms to phrases in the query that was discovered from Web usage data and metadata. The four modules interact to parse a user’s query from multiple perspectives, with the goal of understanding the consumer’s quest intent for data. As a proof-of-concept, the framework is applied to oceanographic data discovery. It is demonstrated that the proposed framework accurately captures a user’s intent.
Journal Article
Introduction to GIS Programming and Fundamentals with Python and ArcGIS
2017
Combining GIS concepts and fundamental spatial thinking methodology with real programming examples, this book introduces popular Python-based tools and their application to solving real-world problems. It elucidates the programming constructs of Python with its high-level toolkits and demonstrates its integration with ArcGIS Theory. Filled with hands-on computer exercises in a logical learning workflow this book promotes increased interactivity between instructors and students while also benefiting professionals in the field with vital knowledge to sharpen their programming skills. Readers receive expert guidance on modules, package management, and handling shape file formats needed to build their own mini-GIS. Comprehensive and engaging commentary, robust contents, accompanying datasets, and classroom-tested exercises are all housed here to permit users to become competitive in the GIS/IT job market and industry.
A Smart Web-Based Geospatial Data Discovery System with Oceanographic Data as an Example
2018
Discovering and accessing geospatial data presents a significant challenge for the Earth sciences community as massive amounts of data are being produced on a daily basis. In this article, we report a smart web-based geospatial data discovery system that mines and utilizes data relevancy from metadata user behavior. Specifically, (1) the system enables semantic query expansion and suggestion to assist users in finding more relevant data; (2) machine-learned ranking is utilized to provide the optimal search ranking based on a number of identified ranking features that can reflect users’ search preferences; (3) a hybrid recommendation module is designed to allow users to discover related data considering metadata attributes and user behavior; (4) an integrated graphic user interface design is developed to quickly and intuitively guide data consumers to the appropriate data resources. As a proof of concept, we focus on a well-defined domain-oceanography and use oceanographic data discovery as an example. Experiments and a search example show that the proposed system can improve the scientific community’s data search experience by providing query expansion, suggestion, better search ranking, and data recommendation via a user-friendly interface.
Journal Article
A Cloud-Based Framework for Large-Scale Log Mining through Apache Spark and Elasticsearch
2019
The volume, variety, and velocity of different data, e.g., simulation data, observation data, and social media data, are growing ever faster, posing grand challenges for data discovery. An increasing trend in data discovery is to mine hidden relationships among users and metadata from the web usage logs to support the data discovery process. Web usage log mining is the process of reconstructing sessions from raw logs and finding interesting patterns or implicit linkages. The mining results play an important role in improving quality of search-related components, e.g., ranking, query suggestion, and recommendation. While researches were done in the data discovery domain, collecting and analyzing logs efficiently remains a challenge because (1) the volume of web usage logs continues to grow as long as users access the data; (2) the dynamic volume of logs requires on-demand computing resources for mining tasks; (3) the mining process is compute-intensive and time-intensive. To speed up the mining process, we propose a cloud-based log-mining framework using Apache Spark and Elasticsearch. In addition, a data partition paradigm, logPartitioner, is designed to solve the data imbalance problem in data parallelism. As a proof of concept, oceanographic data search and access logs are chosen to validate performance of the proposed parallel log-mining framework.
Journal Article
Reconstructing Sessions from Data Discovery and Access Logs to Build a Semantic Knowledge Base for Improving Data Discovery
by
Yang, Chaowei
,
Li, Yun
,
Armstrong, Edward
in
cluster analysis
,
Construction
,
crawler detection
2016
Big geospatial data are archived and made available through online web discovery and access. However, finding the right data for scientific research and application development is still a challenge. This paper aims to improve the data discovery by mining the user knowledge from log files. Specifically, user web session reconstruction is focused upon in this paper as a critical step for extracting usage patterns. However, reconstructing user sessions from raw web logs has always been difficult, as a session identifier tends to be missing in most data portals. To address this problem, we propose two session identification methods, including time-clustering-based and time-referrer-based methods. We also present the workflow of session reconstruction and discuss the approach of selecting appropriate thresholds for relevant steps in the workflow. The proposed session identification methods and workflow are proven to be able to extract data access patterns for further pattern analyses of user behavior and improvement of data discovery for more relevancy data ranking, suggestion, and navigation.
Journal Article
An Integrated Data Analytics Platform
2019
An Integrated Science Data Analytics Platform is an environment that enables the confluence of resources for scientific investigation. It harmonizes data, tools and computational resources to enable the research community to focus on the investigation rather than spending time on security, data preparation, management, etc. OceanWorks is a NASA technology integration project to establish a cloud-based Integrated Ocean Science Data Analytics Platform for big ocean science at NASA’s Physical Oceanography Distributed Active Archive Center (PO.DAAC) for big ocean science. It focuses on advancement and maturity by bringing together several NASA open-source, big data projects for parallel analytics, anomaly detection, in situ to satellite data matchup, quality-screened data subsetting, search relevancy, and data discovery. Our communities are relying on data available through distributed data centers to conduct their research. In typical investigations, scientists would (1) search for data, (2) evaluate the relevance of that data, (3) download it, and (4) then apply algorithms to identify trends, anomalies, or other attributes of the data. Such a workflow cannot scale if the research involves a massive amount of data or multi-variate measurements. With the upcoming NASA Surface Water and Ocean Topography (SWOT) mission expected to produce over 20PB of observational data during its 3-year nominal mission, the volume of data will challenge all existing Earth Science data archival, distribution and analysis paradigms. This paper discusses how OceanWorks enhances the analysis of physical ocean data where the computation is done on an elastic cloud platform next to the archive to deliver fast, web-accessible services for working with oceanographic measurements.
Journal Article
3D modelling strategy for weather radar data analysis
2018
Weather radar data, which have obvious spatial characteristics, represent an important and essential data source for weather identification and prediction, and the multi-dimensional visualization and analysis of such data in a three-dimensional (3D) environment are important strategies for meteorological assessments of potentially disastrous weather. The previous studies have generally used regular 3D raster grids as a basic structure to represent radar data and reconstruct convective clouds. However, conducting weather radar data analyses based on regular 3D raster grids is time-consuming and inefficient, because such analyses involve considerable amounts of tedious data interpolation, and they cannot be used to address real-time situations or provide rapid-response solutions. Therefore, a new 3D modelling strategy that can be used to efficiently represent and analyse radar data is proposed in this article. According to the mode by which the radar data are obtained, the proposed 3D modelling strategy organizes the radar data using logical objects entitled radar-point, radar-line, radar-sector, and radar-cluster objects. In these logical objects, the radar point is the basic object that carries the real radar data unit detected from the radar scan, and the radar-line, radar-sector, and radar-cluster objects organize the radar-point collection in different spatial levels that are consistent with the intrinsic spatial structure of the radar scan. Radar points can be regarded as spatial points, and their spatial structure can support logical objects; thus, the radar points can be flexibly connected to construct continuous surface data with quads and volume data with hexahedron cells without additional tedious data interpolation. This model can be used to conduct corresponding operations, such as extracting an isosurface with the marching cube method and a radar profile with a designed sectioning algorithm to represent the outer and inner structure of a convective cloud. Finally, a case study is provided to verify that the proposed 3D modelling strategy has a better performance in radar data analysis and can intuitively and effectively represent the 3D structure of convective clouds.
Journal Article