Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
25,996
result(s) for
"Optimal control systems"
Sort by:
A Probabilistic Approach to Classical Solutions of the Master Equation for Large Population Equilibria
by
Chassagneux, Jean-François
,
Delarue, François
,
Crisan, Dan
in
Probability theory and stochastic processes -- Special processes -- Interacting random processes; statistical mechanics type models; percolation theory msc
,
Probability theory and stochastic processes -- Stochastic analysis -- Applications of stochastic analysis (to PDE, etc.) msc
,
Stochastic analysis
2022
We analyze a class of nonlinear partial differential equations (PDEs) defined on
Unified reinforcement Q-learning for mean field game and control problems
by
Laurière, Mathieu
,
Fouque, Jean-Pierre
,
Angiuli, Andrea
in
Algorithms
,
Asymptotic properties
,
Machine learning
2022
We present a Reinforcement Learning (RL) algorithm to solve infinite horizon asymptotic Mean Field Game (MFG) and Mean Field Control (MFC) problems. Our approach can be described as a unified two-timescale Mean Field Q-learning: The same algorithm can learn either the MFG or the MFC solution by simply tuning the ratio of two learning parameters. The algorithm is in discrete time and space where the agent not only provides an action to the environment but also a distribution of the state in order to take into account the mean field feature of the problem. Importantly, we assume that the agent cannot observe the population’s distribution and needs to estimate it in a model-free manner. The asymptotic MFG and MFC problems are also presented in continuous time and space, and compared with classical (non-asymptotic or stationary) MFG and MFC problems. They lead to explicit solutions in the linear-quadratic (LQ) case that are used as benchmarks for the results of our algorithm.
Journal Article
Recurrent neural networks for stochastic control problems with delay
2021
Stochastic control problems with delay are challenging due to the path-dependent feature of the system and thus its intrinsic high dimensions. In this paper, we propose and systematically study deep neural network-based algorithms to solve stochastic control problems with delay features. Specifically, we employ neural networks for sequence modeling (e.g., recurrent neural networks such as long short-term memory) to parameterize the policy and optimize the objective function. The proposed algorithms are tested on three benchmark examples: a linear-quadratic problem, optimal consumption with fixed finite delay, and portfolio optimization with complete memory. Particularly, we notice that the architecture of recurrent neural networks naturally captures the path-dependent feature with much flexibility and yields better performance with more efficient and stable training of the network compared to feedforward networks. The superiority is even evident in the case of portfolio optimization with complete memory, which features infinite delay.
Journal Article
Neural network architectures using min-plus algebra for solving certain high-dimensional optimal control problems and Hamilton–Jacobi PDEs
2023
Solving high-dimensional optimal control problems and corresponding Hamilton–Jacobi PDEs are important but challenging problems in control engineering. In this paper, we propose two abstract neural network architectures which are, respectively, used to compute the value function and the optimal control for certain class of high-dimensional optimal control problems. We provide the mathematical analysis for the two abstract architectures. We also show several numerical results computed using the deep neural network implementations of these abstract architectures. A preliminary implementation of our proposed neural network architecture on FPGAs shows promising speedup compared to CPUs. This work paves the way to leverage efficient dedicated hardware designed for neural networks to solve high-dimensional optimal control problems and Hamilton–Jacobi PDEs.
Journal Article
Convergence results for an averaged LQR problem with applications to reinforcement learning
by
Falcone, Maurizio
,
Pesare, Andrea
,
Palladino, Michele
in
Algorithms
,
Conditional probability
,
Control theory
2021
In this paper, we will deal with a linear quadratic optimal control problem with unknown dynamics. As a modeling assumption, we will suppose that the knowledge that an agent has on the current system is represented by a probability distribution π on the space of matrices. Furthermore, we will assume that such a probability measure is opportunely updated to take into account the increased experience that the agent obtains while exploring the environment, approximating with increasing accuracy the underlying dynamics. Under these assumptions, we will show that the optimal control obtained by solving the “average” linear quadratic optimal control problem with respect to a certain π converges to the optimal control driven related to the linear quadratic optimal control problem governed by the actual, underlying dynamics. This approach is closely related to model-based reinforcement learning algorithms where prior and posterior probability distributions describing the knowledge on the uncertain system are recursively updated. In the last section, we will show a numerical test that confirms the theoretical results.
Journal Article
Parameter calibration with stochastic gradient descent for interacting particle systems driven by neural networks
2022
We propose a neural network approach to model general interaction dynamics and an adjoint-based stochastic gradient descent algorithm to calibrate its parameters. The parameter calibration problem is considered as optimal control problem that is investigated from a theoretical and numerical point of view. We prove the existence of optimal controls, derive the corresponding first-order optimality system and formulate a stochastic gradient descent algorithm to identify parameters for given data sets. To validate the approach, we use real data sets from traffic and crowd dynamics to fit the parameters. The results are compared to forces corresponding to well-known interaction models such as the Lighthill–Whitham–Richards model for traffic and the social force model for crowd motion.
Journal Article
Logarithmic regret in online linear quadratic control using Riccati updates
by
Gharesifard, Bahman
,
Linder, Tamas
,
Akbari, Mohammad
in
Algorithms
,
Control systems
,
Control theory
2022
An online policy learning problem of linear control systems is studied. In this problem, the control system is known and linear, and a sequence of quadratic cost functions is revealed to the controller in hindsight, and the controller updates its policy to achieve a sublinear regret, similar to online optimization. A modified online Riccati algorithm is introduced that under some boundedness assumption leads to logarithmic regret bound. In particular, the logarithmic regret for the scalar case is achieved without boundedness assumption. Our algorithm, while achieving a better regret bound, also has reduced complexity compared to earlier algorithms which rely on solving semi-definite programs at each stage.
Journal Article
A discrete method to solve fractional optimal control problems
2015
We present a method to solve fractional optimal control problems, where the dynamic control system depends on integer order and Caputo fractional derivatives. Our approach consists in approximating the initial fractional order problem with a new one that involves integer order derivatives only. The latter problem is then discretized, by application of finite differences, and solved numerically. We illustrate the effectiveness of the procedure with an example.
Journal Article
An application-oriented approach to dual control with excitation for closed-loop identification
by
Ebadat, Afrooz
,
Larsson, Christian A.
,
Rojas, Cristian R.
in
Algorithms
,
Automatic
,
Behavioral research
2016
Identification of systems operating in closed loop is an important problem in industrial applications, where model-based control is used to an increasing extent. For model-based controllers, plant changes over time eventually result in a mismatch between the dynamics of any initial model in the controller and the actual plant dynamics. When the mismatch becomes too large, control performance suffers and it becomes necessary to re-identify the plant to restore performance. Often the available data are not informative enough when the identification is performed in closed loop and extra excitation needs to be injected. This paper considers the problem of generating such excitation with the least possible disruption to the normal operations of the plant. The methods explicitly take time domain constraints into account. The formulation leads to optimal control problems which are in general very difficult optimization problems. Computationally tractable solutions based on Markov decision processes and model predictive control are presented. The performance of the suggested algorithms is illustrated in two simulation examples comparing the novel methods and algorithms available in the literature.
Journal Article
A Compartmental Approach to Modeling the Measles Disease: A Fractional Order Optimal Control Model
by
Al Basir, Fahad
,
Chatterjee, Amar Nath
,
Sharma, Santosh Kumar
in
Basic converters
,
basic reproduction number
,
Birth rate
2024
Measles is the most infectious disease with a high basic reproduction number (R0). For measles, it is reported that R0 lies between 12 and 18 in an endemic situation. In this paper, a fractional order mathematical model for measles disease is proposed to identify the dynamics of disease transmission following a declining memory process. In the proposed model, a fractional order differential operator is used to justify the effect and success rate of vaccination. The total population of the model is subdivided into five sub-compartments: susceptible (S), exposed (E), infected (I), vaccinated (V), and recovered (R). Here, we consider the first dose of measles vaccination and convert the model to a controlled system. Finally, we transform the control-induced model to an optimal control model using control theory. Both models are analyzed to find the stability of the system, the basic reproduction number, the optimal control input, and the adjoint equations with the boundary conditions. Also, the numerical simulation of the model is presented along with using the analytical findings. We also verify the effective role of the fractional order parameter alpha on the model dynamics and changes in the dynamical behavior of the model with R0=1.
Journal Article