Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
51
result(s) for
"Vilenchik, Dan"
Sort by:
From physical to social interactions: The relative entropy model
2020
Complex social systems at various scales of analysis (e.g. dyads, families, tribes, etc.) are formed and maintained through verbal interactions. Therefore, the ability to (1) model these interactions and (2) to use models of interaction for identifying significant relations may be of interest to the social sciences. Adopting the perspective of social physics, we present a general approach for modeling interactions through relative entropy. For illustrating the benefits of the approach, we derive measures of “perspective-taking” and use them for identifying significant-romantic relations in a data set composed of the verbal interactions taken place at the famous TV series “Sex and the City”. Using these measures, we show that significant-romantic relations can be identified with success. These results provide preliminary support for the benefits of using the proposed approach.
Journal Article
Feature Stability as a Trust Layer for Feature Selection: Resampling-Based Recurrence Profiles Beyond Predictive Performance
by
Elmakias, Itamar
,
Kolsky, Dor
,
Vilenchik, Dan
in
chance-corrected stability
,
Dimensional analysis
,
Feature selection
2026
Feature selection in high-dimensional studies is conventionally evaluated by the predictive performance of the features it returns, but predictive performance does not indicate whether the same subset would be selected again under a reasonable perturbation of the data. We propose a feature-stability profile, reported as a diagnostic layer beside predictive performance rather than in place of it. The profile is assembled from established quantities: per-feature selection frequency across repeated stratified resamples, a chance-corrected stability summary, recurrent sets reported across a sweep of descriptive cutoffs, a random-selection baseline, and the held-out predictive performance recorded on the same resamples. We examine it in two settings. In a controlled synthetic study, the informative support is planted by construction, so recovery can be measured directly; in an illustrative application across high-dimensional binary datasets, no such support exists, and a broader exploratory roster is reported as supporting results. Under planted support, predictive performance, subset stability, and support recovery can diverge rather than decline together: at an intermediate signal level, performance can remain relatively preserved while exact recovery falls, and some low exact overlap reflects substitution among redundant alternatives. On real data, recurrent features are treated as recurrent candidates, not recovered or validated features; selectors reaching near-equal area under the receiver operating characteristic curve (AUC) can differ about twofold in chance-corrected recurrence. The contribution is diagnostic and integrative, not a new selector, a new metric, a benchmark ranking, or an error-controlled procedure, making the reliability of a selected feature set visible rather than assumed.
Journal Article
The Missing Signal: Gradient Injection Recovers O(n2) Residual Decoding in GNNs
2026
Residual decoding algorithms repeatedly score the current elements, prune the lowest-scoring candidate, and update the residual state. Classical implementations are efficient because their scores are built from locally maintainable statistics such as degree: after a deletion, surviving scores change by local corrections, giving O(n2) total decoding cost. However, a Graph Neural Network (GNN) score does not automatically have this property. After nonlinear activations such as tanh or sigmoid, equal changes in the underlying statistic need not induce equal changes in the output, so no fixed-additive correction can generally update the score after a deletion. The standard adaptive alternative is to rerun the full network after every deletion, inflating decoding cost to O(n3). We show that, for Max-Clique, the gradient of an unsupervised clique objective provides an update-compatible residual signal. After each vertex deletion, every surviving gradient coordinate changes by a local, maintainable quantity, updatable in O(n) per step. Supplying this live gradient to a lightweight learned updater recovers O(n2) decoding at state-of-the-art quality on Easy and Medium planted clique regimes. Zero-shot evaluation on six real-world graph benchmarks confirms consistent gains over score-only updating.
Journal Article
The Adaptive Behavior of a Soccer Team: An Entropy-Based Analysis
2018
To optimize its performance, a competitive team, such as a soccer team, must maintain a delicate balance between organization and disorganization. On the one hand, the team should maintain organized patterns of behavior to maximize the cooperation between its members. On the other hand, the team’s behavior should be disordered enough to mislead its opponent and to maintain enough degrees of freedom. In this paper, we have analyzed this dynamic in the context of soccer games and examined whether it is correlated with the team’s performance. We measured the organization associated with the behavior of a soccer team through the Tsallis entropy of ball passes between the players. Analyzing data taken from the English Premier League (2015/2016), we show that the team’s position at the end of the season is correlated with the team’s entropy as measured with a super-additive entropy index. Moreover, the entropy score of a team significantly contributes to the prediction of the team’s position at the end of the season beyond the prediction gained by the team’s position at the end of the previous season.
Journal Article
Revisiting Information Cascades in Online Social Networks
by
Hadar, Ofer
,
Sidorov, Michael
,
Vilenchik, Dan
in
Algorithms
,
Analysis
,
Computational linguistics
2025
It is widely believed that a user’s activity pattern in Online Social Networks (OSNs) is strongly influenced by their friends or the users they follow. Building on this intuition, numerous models have been proposed over the years to predict information propagation in OSNs. Many of these models drew inspiration from the process of infectious spread within a population. While this approach is definitely plausible, it relies on knowledge of users’ social connections, which can be challenging to obtain due to privacy concerns. Moreover, while a significant body of work has focused on predicting macro-level features, such as the total cascade size, relatively little attention has been given to the prediction of micro-level features, such as the activity of an individual user. In this study we aim to address this gap by proposing a method to predict the activity of individual users in an OSN, relying solely on their interactions rather than prior knowledge of their social network. We evaluated our results on four large datasets, each comprising over 14 million tweets, recorded on X social network across four different topics over several month. Our method achieved a mean F1 score of 0.86, with a best result of 0.983.
Journal Article
An Oblivious Approach to Machine Translation Quality Estimation
2021
Machine translation (MT) is being used by millions of people daily, and therefore evaluating the quality of such systems is an important task. While human expert evaluation of MT output remains the most accurate method, it is not scalable by any means. Automatic procedures that perform the task of Machine Translation Quality Estimation (MT-QE) are typically trained on a large corpus of source–target sentence pairs, which are labeled with human judgment scores. Furthermore, the test set is typically drawn from the same distribution as the train. However, recently, interest in low-resource and unsupervised MT-QE has gained momentum. In this paper, we define and study a further restriction of the unsupervised MT-QE setting that we call oblivious MT-QE. Besides having no access no human judgment scores, the algorithm has no access to the test text’s distribution. We propose an oblivious MT-QE system based on a new notion of sentence cohesiveness that we introduce. We tested our system on standard competition datasets for various language pairs. In all cases, the performance of our system was comparable to the performance of the non-oblivious baseline system provided by the competition organizers. Our results suggest that reasonable MT-QE can be carried out even in the restrictive oblivious setting.
Journal Article
MolOptimizer: A Molecular Optimization Toolkit for Fragment-Based Drug Design
2024
MolOptimizer is a user-friendly computational toolkit designed to streamline the hit-to-lead optimization process in drug discovery. MolOptimizer extracts features and trains machine learning models using a user-provided, labeled, and small-molecule dataset to accurately predict the binding values of new small molecules that share similar scaffolds with the target in focus. Hosted on the Azure web-based server, MolOptimizer emerges as a vital resource, accelerating the discovery and development of novel drug candidates with improved binding properties.
Journal Article
Fast, accurate, and cost-effective poultry sex genotyping using real-time polymerase chain reaction
2023
According to The Organization for Economic Co-operation and Development (OECD), demand for poultry meat and eggs consumption is growing consistently since poultry meat and eggs are readily available and cheap source for nutritional protein. As such, there is pressing demand from industry for improved protocols to determine chicken sex, especially in layer industry since only females can lay eggs. Extensive efforts are being dedicated to avoiding male chicks culling by developing in-ovo sexing detection methods. Any established in-ovo detection method will need to be validated by embryo genotyping. Therefore, there is a growing demand for fast, inexpensive, and precise method for proper discrimination between males and females in the poultry science community. Our aim with this study was to develop an accurate, high-throughput protocol for sex determination using small volumes of blood. We designed primers targeting the Hint-W gene within the W chromosome clearly distinguishing between males and females. In the interest of establishing an efficient protocol without the need for gel electrophoresis, crude DNA extraction without further purification was coupled with qPCR. We validated the accuracy of our method using established protocols and gonad phenotyping and tested our protocol with four different chicken breeds, day-nine embryos, day-old chicks and adult chicken. In summary, we developed a fast, cost-effective, and accurate method for the genotyping of sex chromosomes in chicken.
Journal Article
Modulation of RNA primer formation by Mn(II)-substituted T7 DNA primase
by
Arthanari, Haribabu
,
Akabayov, Sabine R.
,
Froimovici, Roy
in
101/6
,
631/45/607/1172
,
631/57/2272
2017
Lagging strand DNA synthesis by DNA polymerase requires RNA primers produced by DNA primase. The N-terminal primase domain of the gene 4 protein of phage T7 comprises a zinc-binding domain that recognizes a specific DNA sequence and an RNA polymerase domain that catalyzes RNA polymerization. Based on its crystal structure, the RNA polymerase domain contains two Mg(II) ions. Mn(II) substitution leads to elevated RNA primer synthesis by T7 DNA primase. NMR analysis revealed that upon binding Mn(II), T7 DNA primase undergoes conformational changes near the metal cofactor binding site that are not observed when the enzyme binds Mg(II). A machine-learning algorithm called linear discriminant analysis (LDA) was trained by using the large collection of Mn(II) and Mg(II) binding sites available in the protein data bank (PDB). Application of the model to DNA primase revealed a preference in the enzyme’s second metal binding site for Mn(II) over Mg(II), suggesting that T7 DNA primase activity modulation when bound to Mn(II) is based on structural changes in the enzyme.
Journal Article
DO SEMIDEFINITE RELAXATIONS SOLVE SPARSE PCA UP TO THE INFORMATION LIMIT?
2015
Estimating the leading principal components of data, assuming they are sparse, is a central task in modern high-dimensional statistics. Many algorithms were developed for this sparse PCA problem, from simple diagonal thresholding to sophisticated semidefinite programming (SDP) methods. A key theoretical question is under what conditions can such algorithms recover the sparse principal components? We study this question for a single-spike model with an ℓ₀-sparse eigenvector, in the asymptotic regime as dimension p and sample size n both tend to infinity. Amini and Wainwright [Ann. Statist. 37 (2009) 2877-2921] proved that for sparsity levels k ≥ Ω(n/log p), no algorithm, efficient or not, can reliably recover the sparse eigenvector. In contrast, for $k\\, \\leqslant \\,O(\\sqrt {n/\\log \\,p)} $, diagonal thresholding is consistent. It was further conjectured that an SDP approach may close this gap between computational and information limits. We prove that when $k \\geqslant \\,\\Omega (\\sqrt n )$, the proposed SDP approach, at least in its standard usage, cannot recover the sparse spike. In fact, we conjecture that in the single-spike model, no computationally-efficient algorithm can recover a spike of ℓ₀-sparsity $k \\geqslant \\,\\Omega (\\sqrt n )$. Finally, we present empirical results suggesting that up to sparsity levels $k = \\Omega (\\sqrt {n)} $, recovery is possible by a simple covariance thresholding algorithm.
Journal Article