Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
      More Filters
      Clear All
      More Filters
      Source
    • Language
58 result(s) for "Ishihara, Shunichi"
Sort by:
Stylometry can reveal artificial intelligence authorship, but humans struggle: A comparison of human and seven large language models in Japanese
The purpose of this study was to estimate the artificial intelligence (AI) detection potential using stylometric analysis in Study 1 and examine the AI detection abilities of humans in Study 2. In Study 1, we compared 100 human-written public comments with 350 texts generated by seven large language models (LLMs) (ChatGPT [GPT-4o and o1], Claude3.5, Gemini, Microsoft Copilot, Llama3.1, and Perplexity) using multidimensional scaling (MDS) to visualize differences by focusing on three stylometric features (phrase patterns, part-of-speech bigrams, and unigrams of function words). In general, each stylometric feature can distinguish between LLM-generated and human-written texts. In particular, three integrated stylometric features achieved perfect discrimination on MDS dimensions. Interestingly, only Llama3.1 exhibited distinct characteristics compared with the other six LLMs. The random forest classifier also achieved 99.8% accuracy. In Study 2, we performed an online survey to assess the Japanese participants’ AI detection abilities by presenting LLM-generated and human-written texts, as used in Study 1. 403 participants tackled “AI or Human” judgment task and estimated their own confidence, revealing that overall human AI-detection ability was limited. Moreover, in our materials, more advanced ChatGPT(o1), plausibly reflecting relatively greater fluency and polish, tends to mislead the participants to believe “human-written” texts compared with ChatGPT(GPT-4o) and improves their confidence for their own judgments. Furthermore, an additional comment from the survey suggested that participants primarily relied on superficial impressions based on phraseology, expression, the ends of words, conjunctions, and punctuation marks in judgments. These findings have important implications for various scenarios, including public policy, education, and marketing, where the rapid and reliable detection of AI-generated content is increasing.
Can we spot fake public comments generated by ChatGPT(-3.5, -4)?: Japanese stylometric analysis expose emulation created by one-shot learning
Public comments are an important opinion for civic when the government establishes rules. However, recent AI can easily generate large quantities of disinformation, including fake public comments. We attempted to distinguish between human public comments and ChatGPT-generated public comments (including ChatGPT emulated that of humans) using Japanese stylometric analysis. Study 1 conducted multidimensional scaling (MDS) to compare 500 texts of five classes: Human public comments, GPT-3.5 and GPT-4 generated public comments only by presenting the titles of human public comments (i.e., zero-shot learning, GPT zero ), GPT-3.5 and GPT-4 emulated by presenting sentences of human public comments and instructing to emulate that (i.e., one-shot learning, GPT one ). The MDS results showed that the Japanese stylometric features of the public comments were completely different from those of the GPT zero -generated texts. Moreover, GPT one -generated public comments were closer to those of humans than those generated by GPT zero . In Study 2, the performance levels of the random forest (RF) classifier for distinguishing three classes (human, GPT zero , and GPT one texts). RF classifiers showed the best precision for the human public comments of approximately 90%, and the best precision for the fake public comments generated by GPT (GPT zero and GPT one ) was 99.5% by focusing on integrated next writing style features: phrase patterns, parts-of-speech (POS) bigram and trigram, and function words. Therefore, the current study concluded that we could discriminate between GPT-generated fake public comments and those written by humans at the present time.
Score-based likelihood ratios for linguistic text evidence with a bag-of-words model
•A description of estimating strength of authorship attribution evidence with a score-based likelihood ratio approach.•Efficacy of the likelihood ratio framework for authorship attribution text evidence.•Efficacy of a bag-of-words mode with the relative frequencies of words in the likelihood ratio framework. The likelihood ratio paradigm for quantifying the strength of evidence has been researched in many fields of forensic science. Within this paradigm, score-based approaches for estimating likelihood ratios are becoming more prevalent in the forensic science literature. In this study, a score-based approach for estimating likelihood ratios is implemented for linguistic text evidence. Text data are represented via a bag-of-words model with the Z-score normalised relative frequencies of selected most-frequent words (the number of the most-frequent words = N), and the Euclidean, Manhattan and Cosine distance measures are trialled as the score-generating functions for comparing paired text samples. The score-to-likelihood-ratio conversion model was built using a common source method, and the best fitting model was selected from the parametric models of the Normal, Log-normal, Gamma and Weibull distributions. With the Amazon Product Data Authorship Verification Corpus, two groups of documents (each group including documents of approximately 700, 1400 and 2100 words) were synthesised for each author, allowing 720 same-author comparisons and 517,680 different-author comparisons to test the validity of the system. A series of experiments was conducted using combinations of the following conditions: the three score functions, the different values of N for the feature vector and the different document lengths. The validity of the system was assessed using the log-likelihood-ratio cost (Cllr), and the strength of the derived likelihood ratios was charted in the form of Tippett plots. It was demonstrated that 1) the Cosine measure consistently outperforms the other measures—the best performance is achieved with N = 260, regardless of the document length (e.g., Cllr values of 0.70640, 0.45314 and 0.30692, respectively, for 700, 1400 and 2100 words)—and 2) the derived likelihood ratios are very well calibrated irrespective of the distance measures and document lengths. A follow-up experiment showed that the described score-based approach is relatively robust and stable for a limited quantity of background data. The derived likelihood ratios that were estimated separately to the three distance measures were logistic regression fused; and the fusion achieved a further improvement in performance—for example, a Cllr of 0.23494 for 2100 words. This study demonstrates the possibility of designing likelihood ratio–based systems that discriminate between same-author and different-author documents.
Retrospective non-inferiority study of stereotactic radiosurgery for more than ten brain metastases
Aim This study aimed to investigate the clinical benefits of stereotactic radiosurgery (SRS) in patients with > 10 brain metastases (BM) compared to patients with 2–10 BM. Methods The study included multiple BM patients who underwent SRS between 2014 and 2022, excluding patients who underwent whole brain radiotherapy, had a Karnofsky Performance Status score < 60, suspected leptomeningeal disease, or a single BM lesion. Patients were divided into two groups (2–10 and > 10 BM groups) and matched 2:1 based on propensity scores. The primary endpoint was overall survival (OS) in the matched dataset, with intracranial progression-free survival (PFS) as the secondary endpoint. Non-inferiority was established if the upper limit of the 95% confidence interval (CI) of the adjusted hazard ratio was below 1.3. Results Of the 1042 patients identified, 434 met eligibility criteria. After propensity score matching, 240 patients were analyzed (160 in the BM 2–10 group and 80 in the > 10 BM group). The median OS was 18.2 months in the 2–10 BM group and 19.4 months in the > 10 BM group (P = 0.60). The adjusted hazard ratio was 0.86 (95% CI: 0.59–1.24), indicating non-inferiority. PFS was not significantly different between the groups (4.8 months vs. 4.8 months, P = 0.94). The number of BM did not significantly impact OS or PFS. Conclusions SRS for selected patients with > 10 BM was non-inferior in terms of OS compared to those with 2–10 BM in a propensity score-matched dataset.
Effects of cardiac motion on dose distribution during stereotactic arrhythmia radioablation treatment: A simulation and phantom study
Purpose Cardiac motion may degrade dose distribution during stereotactic arrhythmia radioablation using the CyberKnife system, a robotic radiosurgery system. This study evaluated the dose distribution changes using a self‐made cardiac dynamic platform that mimics cardiac motion. Methods The cardiac dynamic platform was operated with amplitudes of 5 and 3.5 mm along the superior–inferior (SI) and left–right (LR) directions, respectively. The respiratory motion tracking of the CyberKnife system was applied when respiratory motion, simulated using a commercial platform, was introduced. The accuracy of respiratory motion tracking was evaluated by the correlation error between infrared markers and a fiducial marker. The dose distribution was compared with and without cardiac motion. The evaluations included error in the centroid analysis of the irradiated dose distribution, dose profile analysis in the SI and LR directions, and dose distribution analysis comparing the irradiated and planned dose distributions. Results Cardiac motion increased the correlation error in the direction of motion. Cardiac motion displaced the centroid by up to 0.23 and 0.19 mm in the SI and LR directions, respectively. Cardiac motion blurring caused the distance of the isodose lines to become smaller (bigger) at higher (lower) doses in the SI direction. The gamma pass rate was reduced by cardiac motion but exceeded 94.1% with 1 mm/3% for all conditions. Respiratory motion tracking was also effective under cardiac motion. The cardiac motion slightly varied the dose at the edges of the irradiation volume. Conclusion While cardiac motion increased respiratory tracking correlation errors, its effects on dose distribution were limited in this study. Further studies using motion phantoms that are close to a human or individual patient are necessary for a more detailed understanding of the effects of cardiac motion.
Validation in Forensic Text Comparison: Issues and Opportunities
It has been argued in forensic science that the empirical validation of a forensic inference system or methodology should be performed by replicating the conditions of the case under investigation and using data relevant to the case. This study demonstrates that the above requirement for validation is also critical in forensic text comparison (FTC); otherwise, the trier-of-fact may be misled for their final decision. Two sets of simulated experiments are performed: one fulfilling the above validation requirement and the other overlooking it, using mismatch in topics as a case study. Likelihood ratios (LRs) are calculated via a Dirichlet-multinomial model, followed by logistic-regression calibration. The derived LRs are assessed by means of the log-likelihood-ratio cost, and they are visualized using Tippett plots. Following the experimental results, this paper also attempts to describe some of the essential research required in FTC by highlighting some central issues and challenges unique to textual evidence. Any deliberations on these issues and challenges will contribute to making a scientifically defensible and demonstrably reliable FTC available.
Comparison of single- and multi-isocenter planning with Dynamic WaveArc for multiple brain metastases
Dynamic WaveArc (DWA) is a technique used for continuous, non-coplanar volumetric-modulated arc therapy on the Vero4DRT platform. This study aimed to evaluate the application of single-isocenter DWA (SI-DWA) for treating multiple brain metastases by comparing dose distribution and irradiation time with multi-isocenter DWA (MI-DWA) through retrospective treatment planning. Treatment plans were developed for SI-DWA and MI-DWA in 14 cases with 3–5 brain metastases. Parameters assessed included target dose indices, such as conformity index (CI) of the planning target volume (PTV), volumes of normal brain excluding gross tumor volumes (GTVs) receiving a single dose equivalent of 14 Gy (V14), V30%, V20%, V10%, volumes of normal brain, including GTVs receiving a single dose equivalent of 12 Gy (V12), D2% for other organs at risk, and beam-on time. SI-DWA showed inferior CI, V14, and V12 values for lesions with PTV volumes <1 cc, whereas it performed equivalently to MI-DWA for lesions with PTV volumes ≥1 cc. SI-DWA resulted in higher volumes of normal brain receiving low doses compared to MI-DWA. SI-DWA exhibited significantly shorter beam-on times than MI-DWA. In conclusion, SI-DWA is an effective method for treating multiple brain metastases with PTV volumes ≥1 cc, offering an index of radiation-induced brain necrosis comparable with MI-DWA while allowing for shorter irradiation times.
Grammar as a behavioral biometric: using cognitively motivated grammar models for authorship verification
Authorship Verification (AV) is a key area of research in digital text forensics, which addresses the fundamental question of whether two texts were written by the same person. Numerous computational approaches have been proposed over the last two decades in an attempt to address this challenge. However, existing AV methods often suffer from high complexity, low explainability, and especially from a lack of clear scientific justification. We propose a simpler method based on modeling the grammar of an author following Cognitive Linguistics principles. These models are used to calculate λ G (LambdaG): the ratio of the likelihoods of a document given the candidate’s grammar versus given a reference population’s grammar. Our empirical evaluation, conducted on 12 datasets and compared against seven baseline methods, demonstrates that LambdaG achieves superior performance, including against several neural network-based AV methods. LambdaG is also robust to small variations in the composition of the reference population and provides interpretable visualizations, enhancing its explainability. We argue that its effectiveness is due to the method’s compatibility with Cognitive Linguistics theories, predicting that a person’s grammar is a behavioral biometric.
Enhanced urethral identification for radiotherapy planning using fat-suppressed 3D T2-weighted magnetic resonance imaging
This study proposes a fat-suppressed three-dimensional T2-weighted (3D-T2W) sequence on magnetic resonance imaging to enhance prostatic urethral identification in radiotherapy planning. Conventional 3D-T2W and the proposed sequence were obtained to evaluate prostatic urethral identification in 13 male patients. The proposed sequence demonstrated significantly higher Dice similarity coefficients compared to conventional 3D-T2W sequence (p = 0.001) and superior contrast-to-noise ratios. The proposed sequence also achieved significantly better visibility scores in visual assessment (p = 0.001). The proposed technique uses fat suppression in a standard 3D-T2W sequence, making it a simple and clinically applicable method that does not require specialized sequence designs. Our findings suggest that this approach could be a valuable noninvasive method for enhancing prostatic urethral identification, although further research with larger sample sizes and optimization of acquisition parameters is needed.
The importance of choosing the right strategy to treat small cell carcinoma of the cervix: a comparative analysis of treatments
Background Standard treatments for small cell carcinoma of the cervix (SCCC) have not been established. In this study, we aimed to estimate the optimal treatment strategy for SCCC. Methods This was a multicenter retrospective study. Medical records of patients with pathologically proven SCCC treated between 2003 and 2016 were retrospectively analyzed. Overall survival (OS) was plotted using the Kaplan-Meier method. Log-rank tests and Cox regression analysis were used to assess the differences in survival according to stage, treatment strategy, and chemotherapy regimen. Results Data of 78 patients were collected, and after excluding patients without immunohistopathological staining, 65 patients were evaluated. The median age of the included patients was 47 (range: 24–83) years. The numbers of patients with International Federation of Gynecology and Obstetrics (FIGO) 2018 stages I-IIA, IIB-IVA, IVB were 23 (35%), 34 (52%), and 8 (12%), respectively. Of 53 patients who had undergone chemotherapy, 35 and 18 received SCCC and non-SCCC regimens as their first-line chemotherapy regimen, respectively. The 5-year OS for all patients was 49%, while for patients with FIGO stages I-IIA, IIB-IVA, IVB, it was 60, 50, and 0%, respectively. The 5-year OS rates for patients who underwent treatment with SCCC versus non-SCCC regimens were 59 and 13% ( p  < 0.01), respectively. This trend was pronounced in locally advanced stages. Multivariate analysis showed that FIGO IVB at initial diagnosis was a significant prognostic factor in all patients. Among the 53 patients who received chemotherapy, the SCCC regimen was associated with significantly better 5-year OS in both the uni- and multivariate analyses. Conclusion Our results suggest that the application of an SCCC regimen such as EP or IP as first-line chemotherapy for patients with locally advanced SCCC may play a key role in OS. These findings need to be validated in future nationwide, prospective clinical studies.