Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
825 result(s) for "Bellman theory"
Sort by:
Stochastic finite-time partial stability, partial-state stabilization, and finite-time optimal feedback control
In many practical applications, stability with respect to part of the system’s states is often necessary with finite-time convergence to the equilibrium state of interest. Finite-time partial stability involves dynamical systems whose part of the trajectory converges to an equilibrium state in finite time. In this paper, we address finite-time partial stability in probability and uniform finite-time partial stability in probability for nonlinear stochastic dynamical systems. Specifically, we provide Lyapunov conditions involving a Lyapunov function that is positive definite and decrescent with respect to part of the system state and satisfies a differential inequality involving fractional powers for guaranteeing finite-time partial stability in probability. In addition, we show that finite-time partial stability in probability leads to uniqueness of solutions in forward time and we establish necessary and sufficient conditions for almost sure continuity of the settling-time operator of the nonlinear stochastic dynamical system. Finally, we develop a unified framework to address the problem of optimal nonlinear analysis and feedback control design for finite-time partial stochastic stability and finite-time, partial-state stochastic stabilization. Finite-time partial stability in probability of the closed-loop nonlinear system is guaranteed by means of a Lyapunov function that is positive definite and decrescent with respect to part of the system state and can clearly be seen to be the solution to the steady-state form of the stochastic Hamilton–Jacobi–Bellman equation guaranteeing both finite-time, partial-state stability and optimality. The overall framework provides the foundation for extending stochastic optimal linear–quadratic controller synthesis to nonlinear–nonquadratic optimal finite-time, partial-state stochastic stabilization.
Machine Learning Approximation Algorithms for High-Dimensional Fully Nonlinear Partial Differential Equations and Second-order Backward Stochastic Differential Equations
High-dimensional partial differential equations (PDEs) appear in a number of models from the financial industry, such as in derivative pricing models, credit valuation adjustment models, or portfolio optimization models. The PDEs in such applications are high-dimensional as the dimension corresponds to the number of financial assets in a portfolio. Moreover, such PDEs are often fully nonlinear due to the need to incorporate certain nonlinear phenomena in the model such as default risks, transaction costs, volatility uncertainty (Knightian uncertainty), or trading constraints in the model. Such high-dimensional fully nonlinear PDEs are exceedingly difficult to solve as the computational effort for standard approximation methods grows exponentially with the dimension. In this work, we propose a new method for solving high-dimensional fully nonlinear second-order PDEs. Our method can in particular be used to sample from high-dimensional nonlinear expectations. The method is based on (1) a connection between fully nonlinear second-order PDEs and second-order backward stochastic differential equations (2BSDEs), (2) a merged formulation of the PDE and the 2BSDE problem, (3) a temporal forward discretization of the 2BSDE and a spatial approximation via deep neural nets, and (4) a stochastic gradient descent-type optimization procedure. Numerical results obtained using TensorFlow in Python illustrate the efficiency and the accuracy of the method in the cases of a 100-dimensional Black–Scholes–Barenblatt equation, a 100-dimensional Hamilton–Jacobi–Bellman equation, and a nonlinear expectation of a 100-dimensional G -Brownian motion.
Power-optimization of multistage non-isothermal chemical engine system via Onsager equations, Hamilton-Jacobi-Bellman theory and dynamic programming
The research on the output rate performance limit of the multi-stage energy conversion system based on modern optimal control theory is one of the hot spots of finite time thermodynamics. The existing research mainly focuses on the multi-stage heat engine system with pure heat transfer and the multi-stage isothermal chemical engine (ICE) system with pure mass transfer, while the multi-stage non ICE system with heat and mass transfer coupling is less involved. A multistage endoreversible non-isothermal chemical engine (ENICE) system with a finite high-chemical-potential (HCP) source (driving fluid) and an infinite low-chemical-potential sink (environment) is researched. The multistage continuous system is treated as infinitesimal ENICEs located continuously. Each infinitesimal ENICE is assumed to be a single-stage ENICE with stationary reservoirs. Extending single-stage results, the maximum power output (MPO) of the multistage system is obtained. Heat and mass transfer processes between the reservoir and working fluid are assumed to obey Onsager equations. For the fixed initial time, fixed initial fluid temperature, and fixed initial concentration of key component (CKC) in the HCP source, continuous and discrete models of the multistage system are optimized. With given initial reservoir temperature, initial CKC, and total process time, the MPO of the multistage ENICE system is optimized with fixed and free final temperature and final concentration. If the final concentration and final temperature are free, there are optimal final temperature and optimal final concentration for the multistage ENICE system to achieve MPO; meanwhile, there are low limit values for final fluid temperature and final concentration. Special cases for multistage endoreversible Carnot heat engines and ICE systems are further obtained. For the model in this paper, the minimum entropy generation objective is not equivalent to MPO objective.
Memristive Bellman solver for decision-making
The Bellman equation, with a resource-consuming solving process, plays a fundamental role in formulating and solving dynamic optimization problems. The realization of the Bellman solver with memristive computing-in-memory (MCIM) technology, is significant for implementing efficient dynamic decision-making. However, the iterative nature of the Bellman equation solving process poses a challenge for efficient implementation on MCIM systems, which excel at vector-matrix multiplication (VMM) operations but are less suited for iterative algorithms. In this work, by incorporating the temporal dimension and transforming the solution into recurrent dot product operations, a memristive Bellman solver (MBS) is proposed, facilitating the implementation of the Bellman equation solving process with efficient MCIM technology. The MBS effectively reduces the iteration numbers and which further enhanced by approximated solutions leveraging memristor noise. Finally, the path planning tasks are used to verify the feasibility of the proposed MBS. The theoretical derivation and experimental results demonstrate that the MBS effectively reduces the iteration cycles, facilitating the solving efficiency. This work could be a sound of choice for developing high-efficiency decision-making systems. Bellman equation is widely applied in solving dynamic optimization problems. Here, the authors present a memristive Bellman solver reducing the computational complexity and improving energy efficiency for advanced decision-making process.
Optimal control for stochastic neural oscillators
This study develops an event-based, energy-efficient control strategy for desynchronizing coupled neuronal networks using optimal control theory. Inspired by phase resetting techniques in Parkinson’s disease treatment, we incorporate stochasticity of the system’s dynamics into deterministic models to address neural system intrinsic noise. We use an advanced computational solver for nonlinear stochastic partial differential equations to solve the stochastic Hamilton–Jacobi–Bellman equation via level set methods for a single neuron model; this allows us to find control inputs which drive the dynamics close to the system’s phaseless set. When applied to coupled neuronal networks, these inputs achieve effective randomization of neuronal spike timing, leading to significant network desynchronization. Compared to its deterministic counterpart, our stochastic method can achieve considerable energy savings. The event-based control minimizes unnecessary charge transfer, potentially extending implanted stimulator battery life while maintaining robustness against variations in neuronal coupling strengths and network heterogeneities. These findings highlight the potential for developing energy-efficient neurostimulation techniques with implications for deep brain stimulation protocols. The presented computational framework could also be applied to other domains for which stochastic optimal control problems are prevalent.
Optimal Bounds for POD Approximations of Infinite Horizon Control Problems Based on Time Derivatives
In this paper we consider the numerical approximation of infinite horizon problems via the dynamic programming approach. The value function of the problem solves a Hamilton–Jacobi–Bellman equation that is approximated by a fully discrete method. It is known that the numerical problem is difficult to handle by the so called curse of dimensionality. To mitigate this issue we apply a reduction of the order by means of a new proper orthogonal decomposition (POD) method based on time derivatives. We carry out the error analysis of the method using recently proved optimal bounds for the fully discrete approximations. Moreover, the use of snapshots based on time derivatives allows us to bound some terms of the error that could not be bounded in a standard POD approach. Some numerical experiments show the good performance of the method in practice.
Event-triggered optimal tracking control for strict-feedback nonlinear systems with non-affine nonlinear faults
This article studies the control ideas of the optimal backstepping technique, proposing an event-triggered optimal tracking control scheme for a class of strict-feedback nonlinear systems with non-affine and nonlinear faults. A simplified identifier-critic-actor framework is employed in the reinforcement learning algorithm to achieve optimal control. The identifier estimates the unknown dynamic functions, the critic evaluates the system performance, and the actor implements control actions, enabling modeling and control of anonymous systems for achieving optimal control performance. In this paper, a simplified reinforcement learning algorithm is designed by deriving update rules from the negative gradient of a simple positive function related to the Hamilton-Jacobi-Bellman equation, and it also releases the stringent persistent excitation condition. Then, a fault-tolerant control method is developed by applying filtered signals for controller design. Additionally, to address communication resource reduction, an event-triggered mechanism is employed for designing the actual controller. Finally, the proposed scheme’s feasibility is validated through theoretical analysis and simulation.
Policy Iteration for Exploratory Hamilton–Jacobi–Bellman Equations
We study the policy iteration algorithm (PIA) for entropy-regularized stochastic control problems on an infinite time horizon with a large discount rate, focusing on two main scenarios. First, we analyze PIA with bounded coefficients where the controls applied to the diffusion term satisfy a smallness condition. We demonstrate the convergence of PIA based on a uniform C 2 , α estimate for the value sequence generated by PIA, and provide a quantitative convergence analysis for this scenario. Second, we investigate PIA with unbounded coefficients but no control over the diffusion term. In this scenario, we first provide the well-posedness of the exploratory Hamilton–Jacobi–Bellman equation with linear growth coefficients and polynomial growth reward function. By such a well-posedess result we achieve PIA’s convergence by establishing a quantitative locally uniform C 1 , α estimates for the generated value sequence.
Meta-learning-based fault-tolerant attitude control of hypersonic flight vehicle with input constraints
In this study, a meta-learning-based fault-tolerant control (MLFTC) strategy is proposed for the accurate attitude tracking control of hypersonic flight vehicle (HFV) with actuator fault and input constraints. The proposed MLFTC combines the advantage of integral reinforcement learning (IRL) algorithm and meta-learning, can greatly reduce the calculation amount of IRL. By recalling the control purpose of HFV’s attitude system, a tracking error system is derived and the control objective is obtained. Then a neural networks-based IRL algorithm is developed to solve the Hamilton-Jacobi-Bellman equation of the tracking error system, and an approximate optimal control law is directly derived. With considering the actuator fault of HFV, a fault compensator is adopted to achieve optimal fault tolerant control law. The meta-learning ideas are also adopted to address the unknown and sudden faults of HFV, and the convergence speed of the proposed IRL-based FTC law can improved by meta-learning. The ultimately uniformly bounded of the proposed MLFTC is proven by Lyapunov method. Finally, simulation results show that the proposed MLFTC can achieve accurate attitude tracking control for HFV with different actuator faults.
Research on technological innovation decision-making considering government subsidies and corporate reputation
Aiming at the information asymmetry between pharmaceutical enterprises’ technological innovation decisions and government subsidy strategy, this paper establishes a differential game model consisting of the government and a single pharmaceutical company, proposes three different government subsidy strategies, and obtains an equilibrium solution with the help of the Hamilton-Jacobi-Bellman equation, taking into consideration of the transmission effect of the enterprise’s reputation. First, the innovation decisions of pharmaceutical firms without government subsidies are analysed, and based on this, the optimal strategies with government subsidies for non-cooperative pacts and cooperation between the government and enterprises are analysed separately. In addition, the effects of different subsidy strategies on the government’s investment efficiency, corporate reputation, and the choice of corporate innovation strategies are compared, and the results are verified by numerical analysis. Finally, based on the results of the study, references and suggestions are provided for the formulation of government subsidy policies as well as corporate innovation decisions. The results show that: government subsidies can effectively stimulate the innovation ability of pharmaceutical enterprises and improve their reputation; the more sensitive an enterprise’s reputation is to the coefficient of technological innovation, the more it can improve the enterprise’s innovation level; and the coordination contract of government-enterprise cooperation can realize the Pareto improvement of the benefits of the government and enterprises.