Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
3 result(s) for "Johnson, Braidon"
Sort by:
Advancing plant metabolic research by using large language models to expand databases and extract labeled data
Premise Recently, plant science has seen transformative advances in scalable data collection for sequence and chemical data. These large datasets, combined with machine learning, have demonstrated that conducting plant metabolic research on large scales yields remarkable insights. A key next step in increasing scale has been revealed with the advent of accessible large language models, which, even in their early stages, can distill structured data from the literature. This brings us closer to creating specialized databases that consolidate virtually all published knowledge on a topic. Methods Here, we first test different combinations of prompt engineering techniques and language models in the identification of validated enzyme–product pairs. Next, we evaluate the application of automated prompt engineering and retrieval‐augmented generation to identify compound–species associations. Finally, we build and determine the accuracy of a multimodal language model–based pipeline that transcribes images of tables into machine‐readable formats. Results When tuned for each specific task, these methods perform with high (80–90%) or modest (50%) accuracies for enzyme–product pair identification and table image transcription, but with lower false‐negative rates than previous methods (decreasing from 55% to 40%) for compound–species pair identification. Discussion We enumerate several suggestions for researchers working with language models, among which is the importance of the user's domain‐specific expertise and knowledge.
Advancing Plant Metabolic Research By Using Large Language Models To Expand Databases And Extract Labelled Data
Premise: Recently, plant science has seen transformative advances in scalable data collection for sequence and chemical data. These large datasets, combined with machine learning, revealed that conducting plant metabolic research on large scales yields remarkable insights. A key next step in increasing scale has been revealed with the advent of accessible large language models, which, even in their early stages, can distill structured data from literature. This brings us closer to creating specialized databases that consolidate virtually all published knowledge on a topic. Methods: Here, we first test different prompt engineering technique / language model combinations in the identification of validated enzyme-product pairs. Next, we evaluate automated prompt engineering and retrieval augmented generation applied to identifying compound-species associations. Finally, we build and determine the accuracy of a multimodal language model-based pipeline that transcribes images of tables into machine-readable formats. Results: When tuned for each specific task, these methods perform with high accuracies (80-90 percent for enzyme-product pair identification and table image transcription), or with modest accuracies (50 percent) but lower false-negative rates than previous methods (down to 40 percent from 55 percent) for compound-species pair identification. Discussion: We enumerate several suggestions for working with language models as researchers, among which is the importance of the user's domain-specific expertise and knowledge.Competing Interest StatementThe authors have declared no competing interest.Footnotes* https://github.com/thebustalab/ai_in_phytochemistry
Phylochemical mapping of natural products onto the plant tree of life using text mining and large language models
Plants produce a staggering array of chemicals that are the basis for organismal function and diversity and also provide essential human nutrients and medicine. However, it is poorly defined how these compounds have evolved and are distributed across the diverse lineages of the plant kingdom, hindering a systematic view and understanding of plant chemical diversity. Recent advances in plant genome/transcriptome sequencing have provided a well-defined molecular phylogeny of plants, on which the presence of diverse natural products can be mapped to systematically determine their phylogenetic distribution. Here, we built a proof-of-concept workflow via which previously reported diverse tyrosine-derived plant natural products were mapped on to the plant tree of life. Plant chemical-species associations were mined from literature, filtered, evaluated through manual inspection of over 2,500 scientific articles, and mapped onto the plant phylogeny. The resulting 'phylochemical' map confirmed several highly lineage-specific compound class distributions, such as betalain pigments and Amaryllidaceae alkaloids. The map also highlighted several lineages enriched in dopamine-derived compounds, including the orders Caryophyllales, Liliales, and Fabales. Additionally, the application of large language models using our manually curated data as a ground truth set showed that post-mining manual processing steps can largely be automated with a low false positive rate. Our study demonstrates that a workflow combining text mining with language model-based processing can generate broader phylochemical maps, which will serve as a critical community resource to uncover key evolutionary events that underlie plant chemical diversity and enable system-level views of nature's millions of years of chemical experimentation.Competing Interest StatementThe authors have declared no competing interest.Footnotes* https://github.com/thebustalab/phylochemical_mapping