Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
24
result(s) for
"Probabilistic reversal learning"
Sort by:
Effect of lysergic acid diethylamide (LSD) on reinforcement learning in humans
2023
The non-selective serotonin 2A (5-HT
) receptor agonist lysergic acid diethylamide (LSD) holds promise as a treatment for some psychiatric disorders. Psychedelic drugs such as LSD have been suggested to have therapeutic actions through their effects on learning. The behavioural effects of LSD in humans, however, remain incompletely understood. Here we examined how LSD affects probabilistic reversal learning (PRL) in healthy humans.
Healthy volunteers received intravenous LSD (75
g in 10 mL saline) or placebo (10 mL saline) in a within-subjects design and completed a PRL task. Participants had to learn through trial and error which of three stimuli was rewarded most of the time, and these contingencies switched in a reversal phase. Computational models of reinforcement learning (RL) were fitted to the behavioural data to assess how LSD affected the updating ('learning rates') and deployment of value representations ('reinforcement sensitivity') during choice, as well as 'stimulus stickiness' (choice repetition irrespective of reinforcement history).
Raw data measures assessing sensitivity to immediate feedback ('win-stay' and 'lose-shift' probabilities) were unaffected, whereas LSD increased the impact of the strength of initial learning on perseveration. Computational modelling revealed that the most pronounced effect of LSD was the enhancement of the reward learning rate. The punishment learning rate was also elevated. Stimulus stickiness was decreased by LSD, reflecting heightened exploration. Reinforcement sensitivity differed by phase.
Increased RL rates suggest LSD induced a state of heightened plasticity. These results indicate a potential mechanism through which revision of maladaptive associations could occur in the clinical application of LSD.
Journal Article
The feedback-related negativity (FRN) revisited: New insights into the localization, meaning and network organization
by
Iannaccone, Reto
,
Stämpfli, Philipp
,
Drechsler, Renate
in
Adult
,
Anticipation, Psychological - physiology
,
Biological and medical sciences
2014
Changes in response contingencies require adjusting ones assumptions about outcomes of behaviors. Such adaptation processes are driven by reward prediction error (RPE) signals which reflect the inadequacy of expectations. Signals resembling RPEs are known to be encoded by mesencephalic dopamine neurons projecting to the striatum and frontal regions. Although regions that process RPEs, such as the dorsal anterior cingulate cortex (dACC), have been identified, only indirect evidence links timing and network organization of RPE processing in humans. In electroencephalography (EEG), which is well known for its high temporal resolution, the feedback-related negativity (FRN) has been suggested to reflect RPE processing. Recent studies, however, suggested that the FRN might reflect surprise, which would correspond to the absolute, rather than the signed RPE signals. Furthermore, the localization of the FRN remains a matter of debate.
In this simultaneous EEG–functional magnetic resonance imaging (fMRI) study, we localized the FRN directly using the superior spatial resolution of fMRI without relying on any spatial constraint or other assumption. Using two different single-trial approaches, we consistently found a cluster within the dACC. One analysis revealed additional activations of the salience network. Furthermore, we evaluated the effect of signed RPEs and surprise signals on the FRN amplitude. We considered that both signals are usually correlated and found that only surprise signals modulate the FRN amplitude. Last, we explored the pathway of RPE signals using dynamic causal modeling (DCM). We found that the surprise signals are directly projected to the source region of the FRN. This finding contradicts earlier theories about the network organization of the FRN, but is in line with a recent theory stating that dopamine neurons also encode surprise-like saliency signals.
Our findings crucially advance the understanding of the FRN. We found compelling evidence that the FRN originates from the dACC. Furthermore, we clarified the functional role of the FRN, and determined the role of the dACC within the RPE network. These findings should enable us to study the processing of surprise and adjustment signals in the dACC in healthy and also in psychiatric patients.
•The feedback-related negativity (FRN) is associated with surprise signals.•The FRN is rather associated with absolute than signed reward prediction errors.•EEG-informed fMRI consistently locates the FRN in the dorsal anterior cingulum.•Surprise signals are directly projected to the dorsal anterior cingulum.
Journal Article
Ketamine produces no detectable long-term positive or negative effects on cognitive flexibility or reinforcement learning of male rats
by
Walsh, Stephen J
,
Shahan, Timothy A
,
Nist, Anthony N
in
Acute effects
,
Cognitive ability
,
Environmental changes
2024
RationalePatients with major depressive disorder (MDD) often experience abnormalities in behavioral adaptation following environmental changes (i.e., cognitive flexibility) and tend to undervalue positive outcomes but overvalue negative outcomes. The probabilistic reversal learning task (PRL) is used to study these deficits across species and to explore drugs that may have therapeutic value. Selective serotonin-reuptake inhibitors (SSRIs) have limited effectiveness in treating MDD and produce inconsistent effects in non-human versions of the PRL. As such, ketamine, a novel and potentially rapid-acting therapeutic, has begun to be examined using the PRL. Two previous studies examining the effects of ketamine in the PRL have shown conflicting results and only examined short-term effects of ketamine.ObjectiveThis experiment examined PRL performance across a 2-week period following a single exposure to a ketamine dose that varied across groups.MethodsAfter five sessions of PRL training, groups of rats received an injection of either 0, 10, 20 or 30 mg/kg ketamine. One-hour post-injection, rats engaged in the PRL, and subsequently sessions continued daily for 2 weeks. Traditional behavioral and computational reinforcement learning-derived measures were examined.ResultsResults showed that ketamine had acute effects 1-h post-injection, including a significant decrease in the value of the punishment learning rate. Beyond 1 h, ketamine produced no detectable improvements nor decrements in performance across 2 weeks.ConclusionOverall, the present results suggest that the range of ketamine doses examined do not have long-term positive or negative effects on cognitive flexibility or reward processing in healthy rats as measured by the PRL.
Journal Article
How uncertain are you? Disentangling expected and unexpected uncertainty in pupil-linked brain arousal during reversal learning
2023
During decision making, we are continuously faced with two sources of uncertainty regarding the links between stimuli, our actions, and outcomes. On the one hand, our expectations are often probabilistic, that is, stimuli or actions yield the expected outcome only with a certain probability (expected uncertainty). On the other hand, expectations might become invalid due to sudden, unexpected changes in the environment (unexpected uncertainty). Several lines of research show that pupil-linked brain arousal is a sensitive indirect measure of brain mechanisms underlying uncertainty computations. Thus, we investigated whether it is involved in disentangling these two forms of uncertainty. To this aim, we measured pupil size during a probabilistic reversal learning task. In this task, participants had to figure out which of two response options led to reward with higher probability, whereby sometimes the identity of the more advantageous response option was switched. Expected uncertainty was manipulated by varying the reward probability of the advantageous choice option, whereas the level of unexpected uncertainty was assessed by using a Bayesian computational model estimating change probability and resulting uncertainty. We found that both aspects of unexpected uncertainty influenced pupil responses, confirming that pupil-linked brain arousal is involved in model updating after unexpected changes in the environment. Furthermore, high level of expected uncertainty impeded the detection of sudden changes in the environment, both on physiological and behavioral level. These results emphasize the role of pupil-linked brain arousal and underlying neural structures in handling situations in which the previously established contingencies are no longer valid.
Journal Article
A review on exploration–exploitation trade-off in psychiatric disorders
by
Vahabie, Abdol-Hossein
,
Abbaszade, Sajjad
,
Jami, Ali
in
Adaptability
,
Addictions
,
Addictive behaviors
2025
Balancing exploration and exploitation is a crucial aspect of adaptive decision-making, but psychiatric disorders can disrupt this balance in various ways, shedding light on their neurocognitive roots and guiding targeted interventions. In this systematic review, we aimed to delineate potential exploration–exploitation impairments across psychiatric disorders. Through a thorough search on PubMed, we identified forty-six relevant studies employing tasks probing exploration–exploitation balances, which we synthesized to reveal distinct patterns. These disorders are clustered into three categories: addictive patterns, emotional/cognitive disturbances, and neurological (neurodevelopmental and neurodegenerative) disorders. Our findings show that anxiety and mood disorders often enhance exploratory behaviors, while depression impact decision stability and reward sensitivity. In contrast, schizophrenia, OCD (Obsessive–Compulsive Disorder), and ADHD (Attention-Deficit/Hyperactivity Disorder) are characterized by excessive switching and difficulties in balancing exploration and exploitation, leading to impaired learning and adaptability. Additionally, disorders with addictive-like features disrupt optimal decision-making strategies by either heightening exploration or causing maladaptive persistence, thus skewing the balance away from effective decision-making. Individuals exhibiting addiction-like or compulsive behaviors often demonstrate imbalances in the explore-exploit trade-off, resulting in suboptimal decision-making characterized by reduced exploration, flawed foraging strategies, and impulsive or perseverative choices despite adverse outcomes. This suggests that such disorders may originate from dysfunctional foraging processes applied to decision-making. In sum, different patterns of exploration–exploitation balance in different disorders are crucial in understanding the difficulties in learning and decision making of neuropsychiatric disorders. This suggests that such disorders may stem from dysregulated decision-making processes, where uncertainty plays a central role. Dysfunctions in dopaminergic and noradrenergic pathways appear to disrupt the brain's representation of uncertainty, thereby altering exploratory behavior. In sum, the varying patterns of exploration–exploitation balance across different disorders are critical for understanding the challenges in learning and decision-making associated with neuropsychiatric conditions.
Journal Article
Individuals with anxiety and depression use atypical decision strategies in an uncertain world
2024
Previous studies on reinforcement learning have identified three prominent phenomena: (1) individuals with anxiety or depression exhibit a reduced learning rate compared to healthy subjects; (2) learning rates may increase or decrease in environments with rapidly changing (i.e. volatile) or stable feedback conditions, a phenomenon termed learning rate adaptation ; and (3) reduced learning rate adaptation is associated with several psychiatric disorders. In other words, multiple learning rate parameters are needed to account for behavioral differences across participant populations and volatility contexts in this flexible learning rate (FLR) model. Here, we propose an alternative explanation, suggesting that behavioral variation across participant populations and volatile contexts arises from the use of mixed decision strategies. To test this hypothesis, we constructed a mixture-of-strategies (MOS) model and used it to analyze the behaviors of 54 healthy controls and 32 patients with anxiety and depression in volatile reversal learning tasks. Compared to the FLR model, the MOS model can reproduce the three classic phenomena by using a single set of strategy preference parameters without introducing any learning rate differences. In addition, the MOS model can successfully account for several novel behavioral patterns that cannot be explained by the FLR model. Preferences for different strategies also predict individual variations in symptom severity. These findings underscore the importance of considering mixed strategy use in human learning and decision-making and suggest atypical strategy preference as a potential mechanism for learning deficits in psychiatric disorders.
Journal Article
Serotonin Modulates Sensitivity to Reward and Negative Feedback in a Probabilistic Reversal Learning Task in Rats
by
Theobald, David E
,
Dalley, Jeffrey W
,
Mar, Adam C
in
631/378/2649
,
631/378/548/1964
,
631/92/436/2388
2010
Depressed patients show cognitive deficits that may depend on an abnormal reaction to positive and negative feedback. The precise neurochemical mechanisms responsible for such cognitive abnormalities have not yet been clearly characterized, although serotoninergic dysfunction is frequently associated with depression. In three experiments described here, we investigated the effects of different manipulations of central serotonin (5-hydroxytryptamine, 5-HT) levels in rats performing a probabilistic reversal learning task that measures response to feedback. Increasing or decreasing 5-HT tone differentially affected behavioral indices of cognitive flexibility (reversals completed), reward sensitivity (win-stay), and reaction to negative feedback (lose-shift). A single low dose of the selective serotonin reuptake inhibitor citalopram (1 mg/kg) resulted in fewer reversals completed and increased lose-shift behavior. By contrast, a single higher dose of citalopram (10 mg/kg) exerted the opposite effect on both measures. Repeated (5 mg/kg, daily, 7 days) and subchronic (10 mg/kg, b.i.d., 5 days) administration of citalopram increased the number of reversals completed by the animals and increased the frequency of win-stay behavior, whereas global 5-HT depletion had the opposite effect on both indices. These results show that boosting 5-HT neurotransmission decreases negative feedback sensitivity and increases reward (positive feedback) sensitivity, whereas reducing it has the opposite effect. However, these effects depend on the nature of the manipulation used: acute manipulations of the 5-HT system modulate negative feedback sensitivity, whereas long-lasting treatments specifically affect reward sensitivity. These results parallel some of the findings in humans on effects of 5-HT manipulations and are relevant to hypotheses of altered response to feedback in depression.
Journal Article
Chronic cocaine but not chronic amphetamine use is associated with perseverative responding in humans
2008
Rationale
Chronic drug use has been associated with increased impulsivity and maladaptive behaviour, but the underlying mechanisms of this impairment remain unclear. We investigated the ability to adapt behaviour according to changes in reward contingencies, using a probabilistic reversal-learning task, in chronic drug users and controls.
Materials and methods
Five groups were compared: chronic amphetamine users (
n
= 30); chronic cocaine users (
n
= 27); chronic opiate users (
n
= 42); former drug users of psychostimulants and opiates (
n
= 26); and healthy non-drug-taking control volunteers (
n
= 25). Participants had to make a forced choice between two alternative stimuli on each trial to acquire a stimulus–reward association on the basis of degraded feedback and subsequently to reverse their responses when the reward contingencies changed.
Results
Chronic cocaine users demonstrated little behavioural change in response to the change in reward contingencies, as reflected by perseverative responding to the previously rewarded stimulus. Perseverative responding was observed in cocaine users regardless of whether they completed the reversal stage successfully. Task performance in chronic users of amphetamines and opiates, as well as in former drug users, was not measurably impaired.
Conclusion
Our findings provide convincing evidence for response perseveration in cocaine users during probabilistic reversal-learning. Pharmacological differences between amphetamine and cocaine, in particular their respective effects on the 5-HT system, may account for the divergent task performance between the two psychostimulant user groups. The inability to reverse responses according to changes in reinforcement contingencies may underlie the maladaptive behaviour patterns observed in chronic cocaine users but not in chronic users of amphetamines or opiates.
Journal Article
Ketamine decreases sensitivity of male rats to misleading negative feedback in a probabilistic reversal-learning task
2017
Rationale
Depression is characterized by an excessive attribution of value to negative feedback. This imbalance in feedback sensitivity can be measured using the probabilistic reversal-learning (PRL) task. This task was initially designed for clinical research, but introduction of its rodent version provides a new and much needed translational paradigm to evaluate potential novel antidepressants.
Objectives
In the present study, we aimed at evaluating the effects of a compound showing clear antidepressant properties—ketamine (KET)—on the sensitivity of rats to positive and negative feedback in the PRL paradigm.
Methods
We trained healthy rats in an operant version of the PRL task. For successful completion of the task, subjects had to learn to ignore infrequent and misleading feedback, arising from the probabilistic (80:20) nature of the discrimination. Subsequently, we evaluated the effect of KET (5, 10, and 20 mg/kg) on feedback sensitivity 1, 24, and 48 h after administration.
Results
We report that acute administration of the highest dose of KET (20 mg/kg) rapidly and persistently decreases the proportion of lose–shift responses made by rats after receiving negative feedback.
Conclusion
Present results suggest that KET decreases negative feedback sensitivity and that changes in this basic neurocognitive function might be one of the factors responsible for its antidepressant action.
Journal Article
Communicated beliefs about action-outcomes: The role of initial confirmation in the adoption and maintenance of unsupported beliefs
2018
As agents seeking to learn how to successfully navigate their environments, humans can both obtain knowledge through direct experience, and second-hand through communicated beliefs. Questions remain concerning how communicated belief (or instruction) interacts with first-hand evidence integration, and how the former can bias the latter. Previous research has revealed that people are more inclined to seek out confirming evidence when they are motivated to uphold the belief, resulting in confirmation bias. The current research explores whether merely communicated beliefs affect evidence integration over time when it is not of interest to uphold the belief, and all evidence is readily available. In a novel series of on-line experiments, participants chose on each trial which of two options to play for money, being exposed to outcomes of both. Prior to this, they were exposed to favourable communicated beliefs regarding one of two options. Beliefs were either initially supported or undermined by subsequent probabilistic evidence (probabilities reversed halfway through the task, rendering the options equally profitable overall). Results showed that while communicated beliefs predicted initial choices, they only biased subsequent choices when supported by initial evidence in the first phase of the experiment. Findings were replicated across contexts, evidence sequence lengths, and probabilistic distributions. This suggests that merely communicated beliefs can prevail even when not supported by long run evidence, and in the absence of a motivation to uphold them. The implications of the interaction between communicated beliefs and initial evidence for areas including instruction effects, impression formation, and placebo effects are discussed.
Journal Article