Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
18
result(s) for
"Free-text reports"
Sort by:
Structured panendoscopy reports improve report completeness and documentation time
by
Breuer, Thomas
,
Ernst, Benjamin Philipp
,
Potthast, Georg Long Fei
in
631/67
,
631/67/2322
,
Artificial intelligence
2025
Even today, surgical reports are usually dictated in a free text form (FTR), leading to a wide range in report-quality. This study investigated the use of a fully structured panendoscopy report (SR) compared to FTRs. 64 panendoscopies were performed by three experienced head and neck surgeons. The surgical reports were created as both FTRs and SRs, which were examined regarding time to completion and content using a multilevel regression analysis. User satisfaction was evaluated using a questionnaire. There was no significant difference in time to complete the SRs compared to FTRs. The completeness ratings of SRs were significantly higher than for FTRs (81% vs. 66%,
p
< 0.001), leading to increased report quality. Overall user satisfaction was higher for SRs than for conventional FTRs (VAS 8.1 vs. 3.5,
p
< 0.001). The SRs proved to be fast to complete and more comprehensive with a higher completeness of content. Participating surgeons indicated that they preferred SRs over FTRs because of their advantages in terms of structure, guidance for inexperienced residents and non-native speakers. The data stratification also enables secondary data use to further develop deep learning algorithms in patient care and research.
Journal Article
Structured reporting adds clinical value in primary CT staging of diffuse large B-cell lymphoma
by
Dreyling, Martin
,
Nörenberg, Dominik
,
C Benedikt Westphalen
in
B-cell lymphoma
,
Clinical decision making
,
Completeness
2018
ObjectivesTo evaluate whether template-based structured reports (SRs) add clinical value to primary CT staging in patients with diffuse large B-cell lymphoma (DLBCL) compared to free-text reports (FTRs).MethodsIn this two-centre study SRs and FTRs were acquired for 16 CT examinations. Thirty-two reports were independently scored by four haematologists using a questionnaire addressing completeness of information, structure, guidance for patient management and overall quality. The questionnaire included yes-no, 10-point Likert scale and 5-point scale questions. Altogether 128 completed questionnaires were evaluated. Non-parametric Wilcoxon signed-rank test and McNemar’s test were used for statistical analysis.ResultsSRs contained information on affected organs more often than FTRs (95 % vs. 66 %). More SRs commented on extranodal involvement (91 % vs. 62 %). Sufficient information for Ann-Arbor classification was included in more SRs (89 % vs. 64 %). Information extraction was quicker from SRs (median rating on 10-point Likert scale=9 vs. 6; 7–10 vs. 4–8 interquartile range). SRs had better comprehensibility (9 vs. 7; 8–10 vs. 5–8). Contribution of SRs to clinical decision-making was higher (9 vs. 6; 6–10 vs. 3–8). SRs were of higher quality (p < 0.001). All haematologists preferred SRs over FTRs.ConclusionsStructured reporting of CT examinations for primary staging in patients with DLBCL adds clinical value compared to FTRs by increasing completeness of reports, facilitating information extraction and improving patient management.Key Points• Structured reporting in CT helps clinicians to assess patients with lymphoma.• This two-centre study showed that structured reporting improves information content and extraction.• Patient management may be improved by structured reporting.• Clinicians preferred structured reports over free-text reports.
Journal Article
Structured reporting of B-mode, color Doppler, and CEUS in testicular tumor assessment: a reader study with urologist ratings
by
Waldbillig, Frank
,
Clevert, Dirk-Andre
,
Frölich, Matthias Frank
in
Adult
,
CEUS
,
Classification systems
2026
Purpose
Structured reporting (SR) offers standardized radiological documentation, enhancing clarity and reproducibility. However, its role in contrast-enhanced ultrasound (CEUS) for testicular tumors remains underexplored. This study evaluates urologist-perceived clarity, completeness, and clinical usefulness of SR compared to free-text reporting (FTR).
Methods and materials
In this retrospective, single-center study, 65 male patients with suspected testicular tumors underwent CEUS at LMU University Hospital. Reports were initially documented as FTRs by an experienced radiologist and later converted into SRs using Smart Reporting software. Four board-certified urologists independently assessed both formats using a structured questionnaire. Completeness, readability, trust, and impact on clinical decision-making were evaluated. Statistical analysis included McNemar’s test and the Wilcoxon signed-rank test, with α = 0.05.
Results
SRs significantly improved readability (97.3% vs. 10.0%,
p
< 0.001) and information extraction (98.8% vs. 91.9%,
p
< 0.001). However, completeness (56.9% vs. 60.8%,
p
= 0.427) and clinical decision support (85.7% vs. 84.9%,
p
= 0.152) were comparable. Trust in SRs was lower than in FTRs (4.92 vs. 5.22,
p
< 0.001), likely due to missing diagnostic parameters and retrospective SR generation.
Conclusions
SR was associated with improved reporting clarity and consistency but did not outperform FTR in completeness or clinical decision-making. Interdisciplinary collaboration in template development and the integration of classification systems could improve SR’s diagnostic value. Future prospective, multicenter studies should assess real-time SR implementation and its potential impact on reporting quality, communication, and outcome-based endpoints in prospective settings.
Clinical relevance/application
Structured reporting in multiparametric testicular ultrasound including CEUS improved perceived readability and facilitated information access for referring clinicians. However, SR showed no clear advantage over free-text reporting regarding completeness or clinical decision-making. The lower clinician trust in SR highlights the need for clinically tailored templates developed in interdisciplinary collaboration. The broader clinical value of SR in testicular imaging should be confirmed in prospective real-time studies incorporating outcome-based and workflow-related endpoints.
Journal Article
Fine-tuned lightweight language models for structured extraction of liver cancer imaging free-text report: a comparative analysis with existing large language models
2026
Background
Organizing free-text patient data into a structured format is labor-intensive and time-consuming. This study aims to evaluate the effectiveness of a fine-tuned lightweight language model in structuring liver cancer imaging reports.
Methods
A retrospective dataset of 2,780 liver imaging reports from Sun Yat-sen University Cancer Center (2012–2022), including cases of primary liver cancer and benign liver disease, was collected. Three key entries—Number of Malignant Tumors (NMT), Diameter of the Largest Tumor (DLT), and Vascular Invasion (VI)—were annotated by three radiologists and subsequently reviewed and calibrated by a senior oncologist to ensure data reliability. The annotated dataset was randomly split into training, validation, and test sets at a ratio of 7:1:2. A T5-based lightweight model with 250 M parameters (Liver-T5) was fine-tuned using these data. Performance was evaluated using Accuracy and Macro-F1 metrics. Comparative analysis with LLMs such as ChatGLM4, Qianwen2.0, and Llama3.1 was conducted.
Results
The fine-tuned Liver-T5 model outperformed larger LLMs in Exact Match (EM) rate and key evaluation metrics, achieving an EM of 0.8907 and high accuracy for NMT (0.9355) and VI (0.9910). Specifically, for NMT extraction, Liver-T5 achieved an accuracy of 0.9355, outperforming large models such as Qianwen72B (accuracy 0.9140), LLaMA3 (accuracy 0.8961), and ChatGLM4 (accuracy 0.8226). In the VI extraction, Liver-T5 achieved the highest accuracy of 0.9910, significantly surpassing other models, with Qianwen72B, LLaMA3, and ChatGLM4 achieving accuracies of 0.9606, 0.9462, and 0.7581, respectively. A higher proportion of schema-nonconforming outputs was observed in large general-purpose models (e.g., LLaMA3), while Liver-T5 more consistently generated schema-compliant predictions.
Conclusions
The fine-tuned lightweight language model demonstrates superior accuracy and efficiency in structuring liver cancer imaging reports compared to larger LLMs. This capability addresses critical challenges in clinical workflows by converting unstructured data into structured formats.
Journal Article
Structured Reporting of Lung Cancer Staging: A Consensus Proposal
by
Rengo, Marco
,
Sverzellati, Nicola
,
Granata, Vincenza
in
American Recovery & Reinvestment Act 2009-US
,
computed tomography
,
free text report
2021
Background: Structured reporting (SR) in radiology is becoming necessary and has recently been recognized by major scientific societies. This study aimed to build CT-based structured reports for lung cancer during the staging phase, in order to improve communication between radiologists, members of the multidisciplinary team and patients. Materials and Methods: A panel of expert radiologists, members of the Italian Society of Medical and Interventional Radiology, was established. A modified Delphi exercise was used to build the structural report and to assess the level of agreement for all the report sections. The Cronbach’s alpha (Cα) correlation coefficient was used to assess internal consistency for each section and to perform a quality analysis according to the average inter-item correlation. Results: The final SR version was built by including 16 items in the “Patient Clinical Data” section, 4 items in the “Clinical Evaluation” section, 8 items in the “Exam Technique” section, 22 items in the “Report” section, and 5 items in the “Conclusion” section. Overall, 55 items were included in the final version of the SR. The overall mean of the scores of the experts and the sum of scores for the structured report were 4.5 (range 1–5) and 631 (mean value 67.54, STD 7.53), respectively, in the first round. The items of the structured report with higher accordance in the first round were primary lesion features, lymph nodes, metastasis and conclusions. The overall mean of the scores of the experts and the sum of scores for staging in the structured report were 4.7 (range 4–5) and 807 (mean value 70.11, STD 4.81), respectively, in the second round. The Cronbach’s alpha (Cα) correlation coefficient was 0.89 in the first round and 0.92 in the second round for staging in the structured report. Conclusions: The wide implementation of SR is critical for providing referring physicians and patients with the best quality of service, and for providing researchers with the best quality of data in the context of the big data exploitation of the available clinical data. Implementation is complex, requiring mature technology to successfully address pending user-friendliness, organizational and interoperability challenges.
Journal Article
Computed Tomography Structured Reporting in the Staging of Lymphoma: A Delphi Consensus Proposal
by
De Filippo, Massimo
,
Rengo, Marco
,
Lacasella, Giorgia Viola
in
Classification
,
Clinical medicine
,
Disease
2021
Structured reporting (SR) in radiology is becoming increasingly necessary and has been recognized recently by major scientific societies. This study aims to build structured CT-based reports for lymphoma patients during the staging phase to improve communication between radiologists, members of multidisciplinary teams, and patients. A panel of expert radiologists, members of the Italian Society of Medical and Interventional Radiology (SIRM), was established. A modified Delphi process was used to develop the SR and to assess a level of agreement for all report sections. The Cronbach’s alpha (Cα) correlation coefficient was used to assess internal consistency for each section and to measure quality analysis according to the average inter-item correlation. The final SR version was divided into four sections: (a) Patient Clinical Data, (b) Clinical Evaluation, (c) Imaging Protocol, and (d) Report, including n = 13 items in the “Patient Clinical Data” section, n = 8 items in the “Clinical Evaluation” section, n = 9 items in the “Imaging Protocol” section, and n = 32 items in the “Report” section. Overall, 62 items were included in the final version of the SR. A dedicated section of significant images was added as part of the report. In the first Delphi round, all sections received more than a good rating (≥3). The overall mean score of the experts and the sum of score for structured report were 4.4 (range 1–5) and 1524 (mean value of 101.6 and standard deviation of 11.8). The Cα correlation coefficient was 0.89 in the first round. In the second Delphi round, all sections received more than an excellent rating (≥4). The overall mean score of the experts and the sum of scores for structured report were 4.9 (range 3–5) and 1694 (mean value of 112.9 and standard deviation of 4.0). The Cα correlation coefficient was 0.87 in this round. The highest overall means value, highest sum of scores of the panelists, and smallest standard deviation values of the evaluations in this round reflect the increase of the internal consistency and agreement among experts in the second round compared to first round. The accurate statement of imaging data given to referring physicians is critical for patient care; the information contained affects both the decision-making process and the subsequent treatment. The radiology report is the most important source of clinical imaging information. It conveys critical information about the patient’s health and the radiologist’s interpretation of medical findings. It also communicates information to the referring physicians and records this information for future clinical and research use. The present SR was generated based on a multi-round consensus-building Delphi exercise and uses standardized terminology and structures, in order to adhere to diagnostic/therapeutic recommendations and facilitate enrolment in clinical trials, to reduce any ambiguity that may arise from non-conventional language, and to enable better communication between radiologists and clinicians.
Journal Article
Using Generative AI to Extract Structured Information from Free Text Pathology Reports
by
Jian, Wen-Shan
,
Shahid, Fahad
,
Chang, Yung-Chun
in
Applications programs
,
Artificial Intelligence
,
Automation
2025
Manually converting unstructured text pathology reports into structured pathology reports is very time-consuming and prone to errors. This study demonstrates the transformative potential of generative AI in automating the analysis of free-text pathology reports. Employing the ChatGPT Large Language Model within a Streamlit web application, we automated the extraction and structuring of information from 33 unstructured breast cancer pathology reports from Taipei Medical University Hospital. Achieving a 99.61% accuracy rate, the AI system notably reduced the processing time compared to traditional methods. This not only underscores the efficacy of AI in converting unstructured medical text into structured data but also highlights its potential to enhance the efficiency and reliability of medical text analysis. However, this study is limited to breast cancer pathology reports and was conducted using data obtained from hospitals associated with a single institution. In the future, we plan to expand the scope of this research to include pathology reports for other cancer types incrementally and conduct external validation to further substantiate the robustness and generalizability of the proposed system. Through this technological integration, we aimed to substantiate the capabilities of generative AI in improving both the speed and reliability of data processing. The outcomes of this study affirm that generative AI can significantly transform the handling of pathology reports, promising substantial advancements in biomedical research by facilitating the structured analysis of complex medical data.
Journal Article
Leveraging GPT-4o for Automated Extraction and Categorization of CAD-RADS Features From Free-Text Coronary CT Angiography Reports: Diagnostic Study
by
Chen, Youmei
,
Qin, Jie
,
Muhetaier, Abudushalamu
in
Advanced Data Analytics in eHealth
,
Algorithms
,
Artificial Intelligence
2025
Despite the Coronary Artery Reporting and Data System (CAD-RADS) providing a standardized approach, radiologists continue to favor free-text reports. This preference creates significant challenges for data extraction and analysis in longitudinal studies, potentially limiting large-scale research and quality assessment initiatives.
To evaluate the ability of the generative pre-trained transformer (GPT)-4o model to convert real-world coronary computed tomography angiography (CCTA) free-text reports into structured data and automatically identify CAD-RADS categories and P categories.
This retrospective study analyzed CCTA reports from January 2024 and July 2024. A subset of 25 reports was used for prompt engineering to instruct the large language models (LLMs) in extracting CAD-RADS categories, P categories, and the presence of myocardial bridges and noncalcified plaques. Reports were processed using the GPT-4o API (application programming interface) and custom Python scripts. The ground truth was established by radiologists based on the CAD-RADS 2.0 guidelines. Model performance was assessed using accuracy, sensitivity, specificity, and F1-score. Intrarater reliability was assessed using Cohen κ coefficient.
Among 999 patients (median age 66 y, range 58-74; 650 males), CAD-RADS categorization showed accuracy of 0.98-1.00 (95% CI 0.9730-1.0000), sensitivity of 0.95-1.00 (95% CI 0.9191-1.0000), specificity of 0.98-1.00 (95% CI 0.9669-1.0000), and F1-score of 0.96-1.00 (95% CI 0.9253-1.0000). P categories demonstrated accuracy of 0.97-1.00 (95% CI 0.9569-0.9990), sensitivity from 0.90 to 1.00 (95% CI 0.8085-1.0000), specificity from 0.97 to 1.00 (95% CI 0.9533-1.0000), and F1-score from 0.91 to 0.99 (95% CI 0.8377-0.9967). Myocardial bridge detection achieved an accuracy of 0.98 (95% CI 0.9680-0.9870), and noncalcified coronary plaques detection showed an accuracy of 0.98 (95% CI 0.9680-0.9870). Cohen κ values for all classifications exceeded 0.98.
The GPT-4o model efficiently and accurately converts CCTA free-text reports into structured data, excelling in CAD-RADS classification, plaque burden assessment, and detection of myocardial bridges and calcified plaques.
Journal Article
Precision Structuring of Free-Text Surgical Record for Enhanced Stroke Management: A Comparative Evaluation of Large Language Models
by
Wei, Jianyong
,
Zhu, Yueqi
,
Jin, Yidong
in
acute ischemia stroke
,
Chatbots
,
Clinical decision making
2024
Mechanical thrombectomy (MTB) is a critical procedure for acute ischemic stroke (AIS) patients. However, the free-text format of MTB surgical records limits the formulation of effective postoperative patient management and rehabilitation plans. This study compares the efficacy of large language models (LLMs) in structuring data from these free-text MTB surgical record.
This retrospective study collected a total of 382 MTB surgical records from a tertiary hospital. An initial analysis of 30 surgical record from these records provided a guiding prompt for LLMs, focusing on basic and advanced characteristics, such as occlusion locations, thrombectomy maneuvers, reperfusion status, and intraoperative complications. Six LLMs-ChatGPT, GPT-4, GeminiPro, ChatGLM4, Spark3, and QwenMax-were assessed against data extracted by neuroradiologists and a junior physician for comparison. The all 382 surgical records were used to test the performance of LLMs. The performance of the LLMs was quantified using Accuracy, Sensitivity, Specificity, AUC, and MSE as an additional metric for advanced characteristics.
All LLMs showed high performance in characteristic extraction, achieving an average accuracy of 95.09 ± 4.98% across 48 items, and 78.05 ± 4.2% overall. GLM4 and GPT-4 were most accurate in advanced characteristics extraction, with accuracies of 84.03% and 82.20%, respectively. The processing time for LLMs averaged 73.10 ± 10.86 seconds of six models, significantly faster than the 427.88 seconds for manual extraction by physicians.
LLMs, particularly GLM4 and GPT-4, efficiently and accurately structured both general and advanced characteristics from MTB surgical record, outperforming manual extraction methods and demonstrating potential for enhancing clinical data management in AIS treatment.
Journal Article
Artificial Intelligence-Driven Structurization of Diagnostic Information in Free-Text Pathology Reports
by
Frazier, Shellaine
,
Shyu, Chi-Ren
,
Kholod, Olha
in
Free-text pathology reports
,
information extraction
,
n-ary modeling
2020
Background: Free-text sections of pathology reports contain the most important information from a diagnostic standpoint. However, this information is largely underutilized for computer-based analytics. The vast majority of NLP-based methods lack a capacity to accurately extract complex diagnostic entities and relationships among them as well as to provide an adequate knowledge representation for downstream data-mining applications. Methods: In this paper, we introduce a novel informatics pipeline that extends open information extraction (openIE) techniques with artificial intelligence (AI) based modeling to extract and transform complex diagnostic entities and relationships among them into Knowledge Graphs (KGs) of relational triples (RTs). Results: Evaluation studies have demonstrated that the pipeline’s output significantly differs from a random process. The semantic similarity with original reports is high (Mean Weighted Overlap of 0.83). The precision and recall of extracted RTs based on experts’ assessment were 0.925 and 0.841 respectively (P <0.0001). Inter-rater agreement was significant at 93.6% and inter-rated reliability was 81.8%. Conclusion: The results demonstrated important properties of the pipeline such as high accuracy, minimality and adequate knowledge representation. Therefore, we conclude that the pipeline can be used in various downstream data-mining applications to assist diagnostic medicine.
Journal Article