Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
200 result(s) for "mHealth in a Clinical Setting"
Sort by:
Mobile Apps Designed for Patients With Polycystic Ovary Syndrome: Content Analysis Using the Mobile App Rating Scale
Digital health interventions, especially mobile apps, have become instrumental in helping women at risk of polycystic ovary syndrome (PCOS), increasing their understanding of the condition, improving self-care, and fostering empowerment. However, their rapid proliferation has brought about significant challenges regarding quality assessment and evidence-based determination. Therefore, establishing reliable quality assessment methods is essential to assist patients with PCOS in identifying effective and trustworthy mobile health tools. This study was designed to assess the content and quality of mobile apps developed for patients with PCOS using the Mobile App Rating Scale (MARS) to provide insights into their strengths, limitations, and areas needing improvement. In this descriptive-analytical study conducted in June 2024, a comprehensive search was performed to identify English and Persian mobile apps related to PCOS through the Café Bazaar and Google Play Store platforms, using both direct search methods and auxiliary tools such as AppAgg and AppBrain. Two trained reviewers (AR and NN) independently reviewed the apps using the MARS tool. The interrater reliability was measured using the intraclass correlation coefficient test. The quality of each app was scored across 4 dimensions: engagement, functionality, aesthetics, and information quality. Of the initial 199 apps identified, 15 met the inclusion criteria after screening and updates. The interrater agreement rate was 85%, which is considered acceptable. The apps' overall quality was sufficient, as assessed using the MARS, with a mean score of 3.6 (SD 0.52) of 5. Functionality and aesthetics emerged as the highest-scoring dimensions, highlighting user-friendliness and visual appeal (n=10). In contrast, engagement following information quality received the lowest average score, indicating limited interactivity and gaps in providing evidence-based information. The Ask PCOS app achieved the highest overall score, performing exceptionally well in subjective quality (4.75) and app-specific quality (4.33), reflecting its strong capacity to positively impact users' knowledge, attitudes, and behaviors related to PCOS. Uvi Health and Ask PCOS scored highest in engagement (4.2), while PCOS & PCOD Diet & Remedies led in functionality (5), and Uvi Health topped aesthetics (5). The findings revealed that even though many available PCOS-related apps demonstrate strengths in technical performance and design, critical limitations persist regarding user engagement and the credibility of the information provided. The predominance of commercially affiliated apps without academic or clinical oversight was identified as a key contributing factor to these shortcomings. These results underscore the need for future app development to incorporate more user-engaging features, reliable evidence-based content, and personalization strategies to enhance user engagement and support effective PCOS self-management. Addressing these limitations and leveraging the capabilities of existing mobile devices are essential steps toward improving the overall quality and impact of mobile health interventions for individuals with PCOS.
Development and Clinical Evaluation of a Large Language Model–Based System for Generating Patient-Friendly Echocardiography Reports: Two-Stage Retrospective Validation and Prospective Survey Study
Standard echocardiography reports use complex terminology, limiting patient comprehension and exacerbating preconsultation anxiety. Large language models (LLMs) can transform technical data into patient-friendly narratives by incorporating longitudinal comparisons with prior examinations. This study aims to develop an LLM-based patient-friendly echocardiography reporting system and evaluate its professional safety, patient comprehension, and impact on short-term anxiety. This study consisted of 2 stages. In the retrospective development stage, 60 patients were included. Clinical diagnosis, hospitalization records, and serial echocardiographic data were integrated as model inputs using DeepSeek-V3.2. Generated reports followed a standardized 4-module structure. Report quality was independently evaluated by 2 clinicians and an external LLM (Kimi 2.5, Moonshot AI) across 4 domains: data accuracy, information completeness, appropriateness of interpretation, and reasonableness of recommendations. In the prospective clinical evaluation stage, 100 patients undergoing echocardiography and 85 family members were enrolled. Participants received both conventional and LLM-generated patient-friendly reports. A 5-point Likert scale assessed helpfulness in understanding results, effectiveness in addressing concerns, helpfulness in improving disease-related knowledge, and anxiety relief. Anxiety was measured using the STAI-6 (6-item short-form State-Trait Anxiety Inventory) at 3 time points: after echocardiography, after conventional report release, and after reading the patient-friendly report. All 60 reports in the retrospective stage were successfully generated. Professional evaluation showed high overall quality scores from both clinicians and the external LLM, with no significant difference between evaluators (mean total scores 18.15, SD 1.36 vs 18.28, SD 1.26; P=.55). One hallucination event was identified. In the prospective stage, all 100 patients received patient-friendly reports. Both patients and family members rated the reports highly, with no significant between-group difference in total scores (17.61, SD 1.60 vs 17.62, SD 1.03; P=.95). Subgroup analyses showed greater perceived benefit among older patients and outpatients (both P<.001); these subgroup findings should be considered exploratory given the lack of adjustment for multiple comparisons. Patients with chronic heart failure, reduced left ventricular ejection fraction (≤40%), and left ventricular enlargement (>55 mm) reported higher scores for addressing concerns (all P<.001). Anxiety scores increased significantly after conventional report release and decreased significantly after reading the patient-friendly report (both P<.001). Older patients (>60 y) and outpatients showed significantly higher anxiety change rates than their counterparts (both P<.001). The reduction in anxiety was positively correlated with subjective anxiety relief ratings (r=0.531; P<.001). The LLM-based patient-friendly echocardiography reporting system with longitudinal comparison demonstrated good feasibility and promising preliminary clinical usefulness. While maintaining high professional quality, it was associated with improved patient understanding of echocardiographic findings and was associated with reduced short-term anxiety, particularly among older adults and outpatients. Causality cannot be inferred from this nonrandomized sequential design, and longer-term outcomes remain to be evaluated.
Long-Term Mobile-Based Glycemic Intervention for Secondary Prevention in Patients With Diabetes Undergoing Surgical Revascularization: Multicenter Randomized Controlled Trial
Despite the growing amount of patients who underwent coronary artery bypass grafting (CABG) in low- and middle-income countries like China, their glucose control was suboptimal, likely due to poor adherence to healthy lifestyles and preventive medications. Mobile health tools facilitating secondary prevention seem promising, but evidence focusing on this high-risk population is scarce. This study aimed to evaluate the significance of mobile health tools in long-term glycemic management for post-CABG patients with comorbid diabetes mellitus. GUIDEME (glycemic control using mini program-based intervention in patients with diabetes undergoing coronary artery bypass to promote self-management) is a multicenter, open-label, closed-user group, randomized controlled trial, in which 1066 patients with diabetes who had recently undergone CABG were enrolled and allocated into 2 groups. Patients in the control group received conventional health education before discharge, whereas those in the intervention group additionally received automatic delivery of bite-sized health education and medication reminders through a smartphone app during the 6 months after discharge. The primary end point was a change in glycosylated hemoglobin (HbA1c) from baseline to 6 months. Among the 1066 eligible participants enrolled, a total of 1038 (97.4%) had completed the follow-up, while 1000 (93.8%) had 6-month HbA1c results available. Although only 79 (14.9%) patients in the intervention group were defined as active users, a greater reduction of HbA1c in the intervention group was observed (adjusted between-group mean difference -0.13, 95% CI -0.25 to -0.01; P=.04). The intervention group also had a high proportion of good medication adherence (96.1% vs 93.2%, P=.04). There was no difference between the 2 groups regarding the secondary end points. Health education and medication reminders based on smartphone app achieved a statistically significant but modest between-group difference in HbA1c, the clinical relevance of which remains uncertain.
LLM-Generated Lay-Language Protocols for Molecular Tumor Board Patients: Evaluation of Quality and Clinical Usability
Molecular Tumor Boards (MTBs) generate highly technical recommendations. The language used in their protocols is rarely accessible to patients. Lay-language patient protocols could support patient-clinician communication, yet manual production is difficult to sustain in high-volume oncology settings. Large language models (LLMs) may offer scalable drafting assistance, yet clinical usability remains largely uninvestigated under real-world deployment constraints. Existing evaluations rely predominantly on synthetic data or closed-source models that are incompatible with strict data protection requirements. This study evaluated whether open-weight LLMs can provide clinically usable drafting support for German MTB patient protocols under real-world deployment constraints and developed a transferable evaluation framework for patient-facing text generation. Eight open-weight LLMs were evaluated under zero-shot (A1) and one-shot (A2) prompting with constrained decoding, which ensures section-schema compliance. Automatic evaluation used ROUGE-1 (Recall-Oriented Understudy for Gisting Evaluation), BERTScore-F1 (Bidirectional Encoder Representations From Transformers Score), Wiener Sachtextformel version 4, and DistilBERT (Distilled Version of Bidirectional Encoder Representations From Transformers)-based complexity using a corpus of 316 MTB protocols and 47 expert-written patient protocols. For expert evaluation, 7 medical oncologists evaluated 50 protocols from the best-performing model across 3 International Organization for Standardization 9241-11 usability dimensions using fine-grained error annotation, perceived postediting effort (PPEE), and net promoter score. Critical errors were defined as bearing the risk of patient harm. Llama-3.3-70B-Instruct achieved the strongest automatic performance. Across models, A2 significantly improved most automatic metrics compared to A1. However, expert usability evaluation of Llama-3.3-70B-Instruct showed the opposite picture: the proportion of protocols containing at least 1 critical error doubled under A2 (10/25, 40% vs 5/25, 20%) compared with A1, and the dominant error type shifted from language (40/108, 37%) errors to factual errors (69/145, 48%). Overall, 16% (230/1420) of the annotated paragraphs contained errors. Median PPEE was 2 (IQR 2.0-3.0; low), and median net promoter score was 7 (IQR 5.0-9.0). Detractors (46/100, 46%) outweighed promoters (29/100, 29%), which suggests hesitation toward routine adoption. These differences in expert evaluation between A2 and A1 were directionally consistent but did not reach individual statistical significance for the paired samples (n=25). Prompting strategies that improve automatic metrics can simultaneously increase the number of critical errors. Surface-level metric gains were, therefore, insufficient proxies for clinical safety. This was observed as a consistent directional pattern for a single model, but generalization to other models remains to be investigated. Nonetheless, the low paragraph-level error rate and favorable PPEE suggest that structured open-weight LLM generation may be a useful drafting support in a clinician-supervised setting. The proposed evaluation framework provides a text-quality-focused basis for future assessment of patient-facing LLM applications in real-world clinical settings.
Telehealth Scale and Artificial Intelligence Adoption Tiers Across Clinical and Operational Domains in US Hospitals: Cross-Sectional Study
Telehealth expansion and artificial intelligence (AI) adoption are often described as parallel dimensions of health system digital transformation. However, whether telehealth scale is associated with hospital AI adoption and whether this relationship varies across hospital settings remain unclear. This study examined the association of telehealth scale with clinical and operational AI adoption tiers in US hospitals and assessed whether these patterns differed by telehealth reporting behavior and geography. This cross-sectional study included 6173 US acute care hospitals using linked 2024 American Hospital Association Annual Survey and Information Technology Supplement data and 2023 Healthcare Cost Report Information System data. Telehealth scale was parameterized using log-transformed telehealth volume, a telehealth nonreporting indicator, and a reported-zero telehealth indicator. Clinical and operational AI adoption tiers were derived from hospital-reported AI capability items and classified into 3 tiers. Both outcomes were modeled using multioutput gradient-boosted tree classifiers, and model behavior was interpreted using Shapley additive explanations, partial dependence plots, and stratified analyses by the Core-Based Statistical Area category. Telehealth volume was the strongest predictor of both clinical and operational AI adoption tiers and had a larger contribution to the clinical AI model. Telehealth nonreporting was common, occurring in 57% (3521/6173) of hospitals, and was concentrated among hospitals in the lowest clinical AI adoption tier, accounting for 91.4% (3145/3441) of hospitals with no reported clinical AI adoption. Higher telehealth volume was associated with a steep increase in predicted clinical AI adoption tiers at lower telehealth volumes, followed by a plateau at higher volumes. At similar telehealth volumes, rural hospitals showed weaker telehealth-attributed contributions to predicted clinical AI adoption tiers than metropolitan hospitals. Supplementary analyses suggested that telehealth reporting status and telehealth intensity reflected related but distinct structural processes. Telehealth scale was strongly associated with hospital AI adoption tiers, especially clinical AI adoption tiers. These findings suggest that telehealth capacity may serve as a practical hospital-level marker of broader digital readiness for AI adoption, but the cross-sectional design does not establish whether telehealth expansion precedes or causes AI adoption. Hospitals with telehealth nonreporting and rural hospitals may face additional structural barriers that limit the translation of digital capacity into AI maturity. Policies to reduce inequities in hospital AI adoption may therefore need to pair telehealth expansion with implementation support, interoperability capacity, and organizational resources.
High Patient Willingness to Grant Broad Consent for Real-World Data Use in Rheumatology—Implications for Real-World Data Platform Governance: Cross-Sectional Study
Medical real-world data (RWD) are often siloed across organizations, making them inaccessible for research. Unlocking these data could advance clinical research and patient care. The pan-European Data Nexus platform (DNP) links RWD, facilitating its use, for example by artificial intelligence (AI) tools, to support the generation of real-world evidence. In Europe, particularly Germany, the secondary use of health data is governed by stringent regulatory requirements, including informed consent. This study evaluated the informed broad consent form for the RWD (Data Nexus) platform, predicated on the principles of the Medical Informatics Initiative in Germany, and contextualized its implications for future data governance and regulatory use of RWD. The broad consent form was developed for the DNP and cross-sectionally distributed to consecutive rheumatology outpatients during routine follow-up at a tertiary center. Analyses included rates of agreement to the predefined broad consent items. A zero-inflated model (using R) was used to predict response rates. From July 2023 to May 2024, 74.9% (292/390) of the patients signed the broad consent form and consented to DNP data donation. Median age was 56 (IQR 43.0-65.0) years, 72.5% (211/291) were female, and median disease duration was 12 (IQR 4.0-22.0) years. Diagnoses included rheumatoid arthritis (96/291, 33.3%), psoriatic arthritis (30/291, 10.3%), spondyloarthritis (15/291, 5.2%), systemic lupus erythematosus (90/291, 30.9%), systemic sclerosis (14/291, 4.8%), and other conditions (16/291, 5.5%). Patients also answered 8 yes/no broad consent items, with an average of \"yes\" and \"no\" responses of 7.5 (SD 1.5) and 0.2 (SD 0.5), respectively. Missing responses averaged 0.4 (SD 1.4). Of all participants, 78.4% (228/291) agreed to all broad consent items. Approval rates for individual items exceeded 86%, indicating strong patient acceptance of secondary use of RWD under a structured governance framework, possibly reflecting trust, perceived benefit, low perceived risk, and governance confidence. Importantly, patients agreed to new techniques such as AI-based analysis of their donated RWD and, despite the social and ethical sensitivity, data distribution to third parties, including commercial industry. Consent rates were also high for the use of valuable omics and genomics data from biomaterial donations. Women showed higher consent rates, whereas educational attainment was not an indicator of response behavior. This is the first study assessing the willingness of patients with inflammatory rheumatic diseases to grant broad consent for secondary data use in an innovative RWD platform, implementing a modern framework in routine care and considering detailed patient preferences. Patients demonstrated high willingness to grant broad consent, providing key real-world evidence on patients' actual consent behavior for implementing RWD platform integration-such as the European Health Data Space. The DNP supports scalable, General Data Protection Regulation-compliant data sharing, enabling real-world and AI-driven research while preserving patient trust.
Automated Waitlists for Ambulatory Appointment Scheduling: Multisite, Mixed Methods Evaluation
Delays in access to ambulatory care are associated with adverse health outcomes, diminished patient experience, and increased system inefficiencies. US health systems have adopted automated waitlists, a technology-enabled tool that notifies patients of earlier appointment availability. Existing case reports suggest benefits. However, there is limited evidence on the efficacy of adopting, implementing, and sustaining an automated waitlist to improve access. The objective of this study was to evaluate automated waitlists to identify determinants that influence their adoption and ongoing use. Findings may inform health care organizations seeking to improve access to appointments in the ambulatory setting. A convergent, multisite, mixed methods study was conducted. Data were collected through a survey of 127 health systems, with 90 reporting data about automated waitlist usage, accompanied by a criterion-based purposive study of 10 participating health systems. Both qualitative and quantitative data were collected from the 10 participating systems. The Consolidated Framework for Implementation Research was used to report the determinants of the intervention's performance. Automated waitlists provide benefits to US health systems. High-performing health systems reported that 38.8% (IQR 36.2%-45.7%) of appointments offered by the automated waitlist were filled. Participants reported a lower missed appointment rate (3.1%, IQR 2.5%-4.8%) for appointments scheduled via the automated waitlist, as compared to all appointments (6.6%, IQR 4.1%-9.9%). Flexible configuration serves as a key facilitator of adoption and maintenance. External pressures, including peer benchmarking and high patient demand, accelerated implementation. Insurance and records requirements, seasonality, care plan configuration, and digital inequities controlled which patients could benefit from the intervention. Additionally, specialty gatekeeping and clinician capacity constraints limited impact. Organizational context strongly shaped effectiveness, with cross-functional governance structures, leadership endorsement, and cultures of iteration enabling sustained use. Automated waitlists may offer a solution to health care organizations striving to positively impact access to care. The effectiveness of automated waitlists depends primarily on the implementation process and inner-setting organizational determinants. Rather than functioning as a stand-alone technical solution, automated waitlists are most impactful when integrated as a dynamic component of system-level scheduling infrastructure.
A Supervised Fine-Tuned Large Language Model for Lifestyle Management in Patients With Prostate Cancer: Development and Evaluation Study
Lifestyle interventions for patients with prostate cancer have been shown to improve treatment adherence and quality of life. However, there remains a lack of large language models (LLMs) capable of delivering individualized and professional lifestyle recommendations under clearly defined medical safety boundaries and controlled evidence sources. This study aimed to develop and evaluate a supervised fine-tuned LLM-PCaPLMM_SFT (Prostate Cancer Patient Lifestyle Management Model via Supervised Fine-Tuning)-to support health literacy improvement and lifestyle self-management among patients with prostate cancer. We searched English-language literature primarily from PubMed (February 2015 to February 2025) to build a structured lifestyle management knowledge base covering diet, physical activity, weight management, medication adherence, and psychological support. We used a retrieval-augmented generation pipeline to generate patient-style question-answer (QA) pairs from retrieved knowledge slices. Bilingual English-Chinese QA data were generated from English-language source evidence through patient-oriented reformulation and retrieval-augmented generation-based answer generation, and independent English and Chinese test sets were constructed to assess bilingual QA performance. We trained Baichuan2-7B-Chat using a 2-stage strategy, consisting of continued pretraining, followed by supervised fine-tuning with low-rank adaptation. Model outputs were evaluated in 2 double-blind rounds by referee LLMs (Qwen3-Max and DeepSeek-R1) and compared with GPT-3.5-Turbo and the base Baichuan2-7B-Chat using 2500 queries across 5 lifestyle scenarios. Additionally, 3 domain experts conducted a blinded review of 50 QA samples (10 per scenario). We used the Mann-Whitney U test with effect size r, and Benjamini-Hochberg false discovery rate correction, and examined consistency using intraclass correlation coefficients. Based on 2211 included publications, we constructed the PCaPLMM_SFT-Train dataset. The knowledge base yielded >150,000 structured knowledge slices. After 2 rounds of review, we obtained 42,330 single-turn QA pairs and 3008 multiturn dialogues, and the supervised fine-tuning phase used 45,338 structured QA samples. In the dual-round referee LLM assessment, PCaPLMM_SFT consistently outperformed Baichuan2-7B-Chat across dimensions and showed comparable or superior performance to GPT-3.5-Turbo across 5 lifestyle scenarios. Consistency analyses indicated moderate to good agreement between referee models across rounds, supporting the robustness of the comparative evaluation. PCaPLMM_SFT demonstrates the feasibility of constructing a medical lifestyle-focused LLM by integrating structured medical knowledge, QA-style training data, and a multilayer evaluation system. This framework provides a reproducible methodological foundation for evidence-based health education and lifestyle management and establishes groundwork for future evaluation in real-world health management settings.
Communication Barriers in Patient-Provider Interactions in Health Care: Scoping Review
Effective communication is crucial for high-quality health care, but systemic barriers still disrupt patient-provider interactions. Research shows that communication failures are a major reason for preventable medical errors. These issues are linked to around 30% of malpractice claims and over 1744 deaths each year in the United States. In addition, hospitals lose about US $12 billion each year because of miscommunication. Despite the critical nature of this issue, the literature remains fragmented. Most studies focus on specific communication problems rather than examining how these barriers work together in patient-provider settings. This scoping review aims to map the literature thoroughly to (1) identify and categorize key types of communication barriers, (2) evaluate their impacts on patient experience, clinical decision-making, and health outcomes, and (3) examine intervention strategies designed to mitigate these challenges. The findings aim to inform the development of adaptive, technology-enabled solutions to improve patient-provider communication in increasingly diverse and digitally integrated health care settings. A scoping review was conducted in accordance with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines across four databases: PubMed, IEEE Xplore, CINAHL, and Current Contents Connect. This covered the period from January 1, 2004, to April 30, 2026. The search focused on peer-reviewed, English papers that examined language barriers, cultural barriers, psychological barriers, and mental model differences in clinical settings. From an initial set of 6233 records, 253 studies met our stringent inclusion criteria. This review divides communication barriers into four categories: linguistic, cultural, psychological, and mental model differences. It highlights four key findings. First, there has been a notable increase in scholarly interest in communication barriers over the last decade, showing greater awareness in clinical fields. Second, although the four barriers have been known, little is known about the specific communication habits or physiological signs defining each barrier. Third, while researchers have linked individual barriers to adverse outcomes, evidence shows that various barriers often happen at the same time. This creates a combined effect that increases patient distress. Fourth, despite the introduction of promising solutions such as interpreter services, cultural training, and AI tools, a lack of basic knowledge about real-time indicators slows the development of intelligent systems. Unlike previous reviews, which typically examine single communication barriers, this scoping review offers a new approach by mapping how linguistic, cultural, psychological, and cognitive barriers intersect. It provides a unified framework that connects these complex communication failures to specific technological solutions. In practice, this guides the combining of AI tools with real-time physiological monitoring to develop adaptive, patient-centered communication strategies. Ultimately, incorporating these responsive technologies into clinical practice is crucial for lowering medical errors, addressing health disparities, and ensuring fair, patient-centered care.
Prompt-Sensitive Decision Behavior of Large Language Models in Intensive Care Unit Mortality Prediction for Spontaneous Intracerebral Hemorrhage: Comparative Benchmarking Study
Large language models (LLMs) are increasingly being explored for clinical decision support. However, whether inference-only LLM outputs can be interpreted as reliable quantitative risk estimates in structured clinical prediction remains unclear. This study aimed to evaluate the predictive performance and decision-making behavior of inference-only LLMs in a structured clinical prediction task and compare their outputs with those of an outcome-trained machine learning model. We conducted a controlled benchmarking study using identical structured clinical inputs from patients admitted to intensive care units with spontaneous intracerebral hemorrhage. An outcome-trained extreme gradient boosting model was compared with predictions generated by a general-purpose LLM using four prompting strategies: zero-shot, few-shot, chain-of-thought, and combined few-shot plus chain-of-thought prompting. Performance was evaluated using discrimination metrics, threshold-dependent classification behavior, and concordance between Shapley Additive Explanations-derived feature importance rankings and LLM-derived feature prioritization. The independent testing cohort included 435 patients, of whom 86 (19.7%) experienced in-hospital mortality. The outcome-trained machine learning model demonstrated superior discriminative performance compared with all LLM-based approaches. LLM predictions achieved moderate discrimination but exhibited substantial variability in threshold-dependent classification behavior across prompting strategies. At a fixed probability threshold of 0.5, LLM approaches consistently demonstrated high sensitivity and lower specificity, whereas operating characteristics varied considerably when thresholds were optimized using the Youden index. The optimal thresholds for LLM-based approaches ranged from 0.74 to 0.88, compared with 0.1555 for the extreme gradient boosting model. Concordance between Shapley Additive Explanations-derived attribution and LLM-derived feature prioritization was modest, suggesting only partial alignment between empirically learned predictor structure and language-based reasoning patterns. In this structured clinical prediction setting, inference-only LLM outputs demonstrated prompt-sensitive decision behavior despite moderate discriminative performance. These findings suggest that LLM-generated probability outputs should be interpreted cautiously when used for quantitative clinical risk estimation. A complementary framework integrating outcome-trained predictive models with LLM-assisted reasoning may provide a more reliable direction for future clinical decision support systems.