Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
12
result(s) for
"Brockhauser, Sandor"
Sort by:
Comparing End-to-End Machine Learning Methods for Spectra Classification
by
Hegedűs, Péter
,
Sun, Yue
,
Brockhauser, Sandor
in
Algorithms
,
Classification
,
convolutional neural network (CNN)
2021
In scientific research, spectroscopy and diffraction experimental techniques are widely used and produce huge amounts of spectral data. Learning patterns from spectra is critical during these experiments. This provides immediate feedback on the actual status of the experiment (e.g., time-resolved status of the sample), which helps guide the experiment. The two major spectral changes what we aim to capture are either the change in intensity distribution (e.g., drop or appearance) of peaks at certain locations, or the shift of those on the spectrum. This study aims to develop deep learning (DL) classification frameworks for one-dimensional (1D) spectral time series. In this work, we deal with the spectra classification problem from two different perspectives, one is a general two-dimensional (2D) space segmentation problem, and the other is a common 1D time series classification problem. We focused on the two proposed classification models under these two settings, the namely the end-to-end binned Fully Connected Neural Network (FCNN) with the automatically capturing weighting factors model and the convolutional SCT attention model. Under the setting of 1D time series classification, several other end-to-end structures based on FCNN, Convolutional Neural Network (CNN), ResNets, Long Short-Term Memory (LSTM), and Transformer were explored. Finally, we evaluated and compared the performance of these classification models based on the High Energy Density (HED) spectra dataset from multiple perspectives, and further performed the feature importance analysis to explore their interpretability. The results show that all the applied models can achieve 100% classification confidence, but the models applied under the 1D time series classification setting are superior. Among them, Transformer-based methods consume the least training time (0.449 s). Our proposed convolutional Spatial-Channel-Temporal (SCT) attention model uses 1.269 s, but its self-attention mechanism performed across spatial, channel, and temporal dimensions can suppress indistinguishable features better than others, and selectively focus on obvious features with high separability.
Journal Article
Application of self-supervised approaches to the classification of X-ray diffraction spectra during phase transitions
by
Hegedűs, Péter
,
Gelisio, Luca
,
Sun, Yue
in
639/301/119/2795
,
639/705/117
,
Humanities and Social Sciences
2023
Spectroscopy and X-ray diffraction techniques encode ample information on investigated samples. The ability of rapidly and accurately extracting these enhances the means to steer the experiment, as well as the understanding of the underlying processes governing the experiment. It improves the efficiency of the experiment, and maximizes the scientific outcome. To address this, we introduce and validate three frameworks based on self-supervised learning which are capable of classifying 1D spectral curves using data transformations preserving the scientific content and only a small amount of data labeled by domain experts. In particular, in this work we focus on the identification of phase transitions in samples investigated by x-ray powder diffraction. We demonstrate that the three frameworks, based either on relational reasoning, contrastive learning, or a combination of the two, are capable of accurately identifying phase transitions. Furthermore, we discuss in detail the selection of data augmentation techniques, crucial to ensure that scientifically meaningful information is retained.
Journal Article
Predicting the X-ray lifetime of protein crystals
by
Garman, Elspeth F.
,
Zeldin, Oliver B.
,
Bremridge, John
in
Animals
,
Axes of rotation
,
Biological Sciences
2013
Radiation damage is a major cause of failure in macromolecular crystallography experiments. Although it is always best to evenly illuminate the entire volume of a homogeneously diffracting crystal, limitations of the available equipment and imperfections in the sample often require a more sophisticated targeting strategy, involving microbeams smaller than the crystal, and translations of the crystal during data collection. This leads to a highly inhomogeneous distribution of absorbed X-rays (i.e., dose). Under these common experimental conditions, the relationship between dose and time is nonlinear, making it difficult to design an experimental strategy that optimizes the radiation damage lifetime of the crystal, or to assign appropriate dose values to an experiment. We present, and experimentally validate, a predictive metric diffraction-weighted dose for modeling the rate of decay of total diffracted intensity from protein crystals in macromolecular crystallography, and hence we can now assign appropriate “dose” values to modern experimental setups. Further, by taking the ratio of total elastic scattering to diffraction-weighted dose, we show that it is possible to directly compare potential data-collection strategies to optimize the diffraction for a given level of damage under specific experimental conditions. As an example of the applicability of this method, we demonstrate that by offsetting the rotation axis from the beam axis by 1.25 times the full-width half maximum of the beam, it is possible to significantly extend the dose lifetime of the crystal, leading to a higher number of diffracted photons, better statistics, and lower overall radiation damage.
Journal Article
Gold Standard for macromolecular crystallography diffraction data
by
Bernstein, Herbert J.
,
Brockhauser, Sandor
,
Santoni, Gianluca
in
BASIC BIOLOGICAL SCIENCES
,
Chemical Sciences
,
Construction standards
2020
Macromolecular crystallography (MX) is the dominant means of determining the three-dimensional structures of biological macromolecules. Over the last few decades, most MX data have been collected at synchrotron beamlines using a large number of different detectors produced by various manufacturers and taking advantage of various protocols and goniometries. These data came in their own formats: sometimes proprietary, sometimes open. The associated metadata rarely reached the degree of completeness required for data management according to Findability, Accessibility, Interoperability and Reusability (FAIR) principles. Efforts to reuse old data by other investigators or even by the original investigators some time later were often frustrated. In the culmination of an effort dating back more than two decades, a large portion of the research community concerned with high data-rate macromolecular crystallography (HDRMX) has now agreed to an updated specification of data and metadata for diffraction images produced at synchrotron light sources and X-ray free-electron lasers (XFELs). This `Gold Standard' will facilitate the processing of data sets independent of the facility at which they were collected and enable data archiving according to FAIR principles, with a particular focus on interoperability and reusability. This agreed standard builds on the NeXus/HDF5 NXmx application definition and the International Union of Crystallography (IUCr) imgCIF/CBF dictionary, and it is compatible with major data-processing programs and pipelines. Just as with the IUCr CBF/imgCIF standard from which it arose and to which it is tied, the NeXus/HDF5 NXmx Gold Standard application definition is intended to be applicable to all detectors used for crystallography, and all hardware and software developers in the field are encouraged to adopt and contribute to the standard.
Journal Article
Shared metadata for data-centric materials science
2023
The expansive production of data in materials science, their widespread sharing and repurposing requires educated support and stewardship. In order to ensure that this need helps rather than hinders scientific work, the implementation of the FAIR-data principles (
Findable, Accessible, Interoperable, and Reusable
) must not be too narrow. Besides, the wider materials-science community ought to agree on the strategies to tackle the challenges that are specific to its data, both from computations and experiments. In this paper, we present the result of the discussions held at the workshop on “Shared Metadata and Data Formats for Big-Data Driven Materials Science”. We start from an operative definition of metadata, and the features that a FAIR-compliant metadata schema should have. We will mainly focus on computational materials-science data and propose a constructive approach for the
FAIRification
of the (meta)data related to ground-state and excited-states calculations, potential-energy sampling, and generalized workflows. Finally, challenges with the
FAIRification
of experimental (meta)data and materials-science ontologies are presented together with an outlook of how to meet them.
Journal Article
Good Usability Practices in Scientific Software Development
by
Silva, Raniere
,
Fangohr, Hans
,
Miller, Jonah
in
Computation
,
Machine tool industry
,
Product design
2017
Scientific software often presents very particular requirements regarding usability, which is often completely overlooked in this setting. As computational science has emerged as its own discipline, distinct from theoretical and experimental science, it has put new requirements on future scientific software developments. In this paper, we discuss the background of these problems and introduce nine aspects of good usability. We also highlight best practices for each aspect with an emphasis on applications in computational science.
Open and FAIR Raman spectroscopy. Paving the way for artificial intelligence
by
Polli, Dario
,
Gorenflot, Julien
,
Kochev, Nikolay
in
Analytical Chemistry
,
Physical Chemistry
,
Review
2025
Raman spectroscopy is an increasingly powerful and fast-growing analytical technique across diverse disciplines, from materials science and chemistry to biology and medicine, thanks to advances in Raman instrumentation and greatly supported by the flourishing of chemometrics and artificial intelligence (AI). However, the full potential of this technique is often hampered by challenges related to data acquisition, processing, interpretation, and sharing. This review paper addresses how a concerted effort towards digitalization, incorporating principles of Open Science and FAIR data (Findable, Accessible, Interoperable, and Reusable), is essential to develop and implement robust, standardized, and accessible digital workflows in the field of Raman spectroscopy and thereby unlock the full power of Raman spectroscopy in combination with AI. We explore the current landscape of digital tools and open resources in Raman spectroscopy, highlight existing solutions as well as critical gaps. In this regard, we assess the trends in Raman spectroscopy hardware and control software as well as the role of artificial intelligence and machine learning in improving data collection, automating data analysis, extracting meaningful insights, and enabling predictive modelling. Furthermore, we discuss the importance of standardized data formats, metadata schemas, and ontologies to ensure database federation and interoperability as well as to facilitate collaborative research. We also provide lists of existing open hardware, open databases and standards. Finally, we propose a roadmap toward an open and FAIR ecosystem for Raman spectroscopy, emphasizing the need for sustainable infrastructure, collaborative development, and community involvement.
pynxtools: A Python framework for generating and validating NeXus files in experimental data workflows
by
Draxl, Claudia
,
Weber, Heiko B
,
Kühbach, Markus
in
Data conversion
,
Data management
,
Data storage
2025
Scientific data across physics, materials science, and materials engineering often lacks adherence to FAIR principles (Barker et al., 2022; Jacobsen et al., 2020; M. D. Wilkinson et al., 2016; S. R. Wilkinson et al., 2025) due to incompatible instrument-specific formats and diverse standardization practices. pynxtools is a Python software development framework with a command line interface (CLI) that standardizes data conversion for scientific experiments in materials science to the NeXus format (Klosowski et al., 1997; Könnecke, 2006; Könnecke et al., 2015) across diverse scientific domains. NeXus defines data storage specifications for different experimental techniques through application definitions. pynxtools provides a fixed, versioned set of NeXus application definitions that ensures convergence and alignment in data specifications across, among others, atom probe tomography, electron microscopy, optical spectroscopy, photoemission spectroscopy, scanning probe microscopy, and X-ray diffraction. Through its modular plugin architecture pynxtools provides conversion of data and metadata from instruments and electronic lab notebooks to these unified definitions, while performing validation to ensure data correctness and NeXus compliance. pynxtools can be integrated directly into Research Data Management Systems (RDMS) to facilitate parsing and normalization. We detail one example for the RDM system NOMAD. By simplifying the adoption of NeXus, the framework enables true data interoperability and FAIR data management across multiple experimental techniques.