Abstract
Lithium ion batteries are crucial in the field of modern energy storage. Since their commercialization in the 1990s, advances in materials science and engineering have enhanced their capacity, safety, and lifespan. However, the dynamic characteristics of lithium-ion batteries are complex and require advanced charging and control strategies to optimize performance, safety, and lifespan. This article proposes a comparative analysis of three advanced control methods for lithium-ion battery charging: reinforcement learning, fuzzy logic control, and traditional proportional integral derivative (PID) control. Traditional charging methods are difficult to cope with the dynamic complexity of the battery, resulting in poor performance. This article uses MATLAB Simulink simulation to evaluate these intelligent control strategies to improve charging efficiency, speed, and battery life. The results show that reinforcement learning has strong adaptability, fuzzy logic can effectively handle nonlinear problems, PID control performance is reliable, and computational resource requirements are low.
1. Introduction
The limitations of lithium-ion batteries compared to traditional charging methods: Lithium ion batteries have the advantages of high energy density, long cycle life, and low self discharge rate, completely changing the way energy is stored and used, and becoming one of the fundamental technologies in modern society. However, traditional constant current and constant voltage (CC/CV) charging methods often fail to cope with the dynamic complexity of lithium-ion batteries, resulting in poor charging performance and possible degradation of battery performance over time.
Introduction to Intelligent Control Strategy
To address the challenges, this article compares and analyzes three main intelligent control methods for lithium-ion battery charging: reinforcement learning (RL), fuzzy logic control (FL), and traditional proportional integral derivative (PID) control.
The RL controller learns the optimal control strategy by interacting with the battery model and power electronic devices, using a small signal model to simplify the power electronic devices and enhance the training process. Training a neural network based on a reward function to penalize current and voltage spikes for a more stable charging process, with the aim of controlling the battery input voltage. By evaluating the response through multiple interactions and maximizing the reward value, while monitoring the battery status, rewards or punishments can be given based on the current value and aggressive control actions. It can be found that a charging strategy that minimizes charging time, energy consumption, and battery degradation while ensuring safe operation can be developed.
The FL controller provides a flexible and intuitive way to integrate expert knowledge and heuristic rules into the charging process by defining language variables such as charging status, temperature, and charging rate, and establishing inference rules. Two FL controllers were designed, one regulating voltage to maintain stability under different charging curves, and the other regulating current to avoid excessive spikes and maintain stable values, which can better handle the nonlinearity and inherent uncertainty in the dynamics of lithium-ion batteries, thereby improving charging performance and extending battery life.
The PID controller can balance factors such as charging time, energy efficiency, and battery health protection by adjusting and optimizing the charging curve.
2. Research methods for charging control strategies of lithium-ion batteries
Advantages and battery characteristics of MATLAB Simulink simulation platform: MATLAB Simulink provides a powerful platform for analyzing and optimizing lithium-ion battery charging systems. Lithium ion batteries have become the preferred choice for various applications due to their high energy density (ensuring a high energy to weight ratio, suitable for portable electronic devices and electric vehicles with limited space and weight), low self discharge rate (suitable for long-term energy storage), advanced recycling technology, and low environmental impact (more sustainable and environmentally friendly).
Modeling of average small signal converter: Effective design is required for lithium-ion battery energy transfer systems. In this study, an isolated DC/DC converter circuit is used (the forward converter is similar to a DC/DC buck converter topology and includes a transformer to provide electrical isolation and enhance battery safety). The circuit behavior is characterized as a second-order transfer function using the small signal model criterion, and the low signal (low Q) approximation is used to simplify the analysis, resulting in a converter model transfer function that includes operation and conversion ratio. The converter simulates the use of ideal components to improve efficiency, and analyzes two scenarios: no-load and with lithium-ion battery load. Under no-load conditions, the transfer function and converter voltage and current trends are similar, but the converter output oscillates. Under load conditions, the main difference lies in the stabilization time. The idealization of the transfer function makes the system respond faster but the final output value is consistent, and using the transfer function can significantly shorten the simulation time.

Control architecture description: The control strategy is based on the lithium-ion battery model in MATLAB Simulink, which provides detailed technical parameters for accurate simulation and analysis of battery performance. The CC/CV charging process includes a current control stage (where the current is set to a safe level and the battery voltage increases with charging, reaching a threshold before entering the voltage control stage) and a voltage control stage (where the voltage remains stable and the current gradually decreases until the battery is fully charged), which can prevent overcharging, reduce battery pressure, lower the risk of overheating, and extend battery life. The evaluated control strategies are summarized in Table 1 and will be introduced in detail later.
| Controller | Scenario 1 | Scenario 2 | Scenario 3 |
| Voltage | Reinforcement Learning (RL) | Sugeno Fuzzy PD | PID |
| Current | PI | Sugeno Fuzzy PD + I | PID |
Reinforcement learning architecture: The reinforcement learning controller learns the optimal strategy through interaction with the system to handle complex nonlinear dynamics. Adopting an actor critic scheme, the actor network selects actions and the critic network evaluates actions to provide feedback. Continuous Gaussian Actor Network (CGAN) is used to process continuous action spaces, and the optimal strategy is explored by outputting Gaussian distribution parameters. Its architecture includes multiple fully connected layers, and the activation function is mostly a Modified Linear Unit Function (RELU). Based on MATLAB reinforcement learning toolbox training, the reward function is calculated based on voltage error, control actions, and voltage observations to incentivize maintaining voltage within the expected range and penalize deviations. By setting constants, the maximum learning effect is ensured, and the maximum average reward is achieved after 200 rounds of training. However, slow battery response makes parameter tuning difficult.


Fuzzy architecture: Fuzzy proportional derivative (PD) controllers are more effective in handling system nonlinearity and uncertainty than traditional PD controllers, adapting to changes in battery characteristics to ensure stable and accurate voltage control. A fuzzy inference system is constructed using the Sugeno scheme, with voltage controlled by fuzzy PD and current controlled by fuzzy PD+I. The input is processed using normalized triangular membership functions to handle errors and their derivatives, and the output is processed using three Sugeno normalization functions to handle states. The input range is constrained to avoid saturation, and the input and output ranges are modified by relevant constants.

Classic PID architecture: The classic proportional integral derivative (PID) controller is widely used in battery charging due to its simplicity, effectiveness, and reliable performance. It is easy to implement and tune, and can adjust and optimize charging conditions in real time. It combines proportional, integral, and derivative actions to accurately regulate voltage and current. It has strong universality and is suitable for various battery and charging scenarios. It has low cost and low computational resource requirements, and is suitable for embedded systems and low-cost hardware implementation. Its architecture only uses two classic controllers to control voltage and current respectively (the internal structure is not detailed due to space limitations).
3. Performance evaluation of lithium-ion battery charging control strategy
Controller parameter tuning: The reinforcement learning reward function is determined through heuristic methods, and the constants are adjusted through repeated experiments to obtain (r1=200), (r2=-25), (p1=-10), and (p2=180). The fuzzy controller (PD+I) adjusts parameters manually through trial and error, with current control parameters of (P=20), (D=0. 00000 1), (I=2), and (K_D=0.297), and voltage control parameters of (P=15), (D=0.0001), and (K_D=0.315). The PID controller uses MATLAB PID Tuner tool and neural network to find the optimal parameters, including current control (P=15), (I {PID}=5), voltage control (P=22.5), (I=4.9), and (D {PID}=0.03).

Performance analysis of controller
Evaluation aspect: Evaluate the performance of the controller in terms of voltage and current regulation accuracy, response time, stability, and anti-interference.
| Controller | RMS Voltage [V] | RMS Current [A] | Charging Time | Simulation Time 6 s | Step Simulation Time |
| RL | 3.9347 | 0.3 | 10,017.57 s | 21,046.5 s | 3.5 ms |
| FuZZy PI+ D | 3.8601 | 0.3 | 18,401.40 s | 1961.4 s | 0.33 ms |
| PID | 3.8601 | 0.3 | 12,933.57 s | 87.90 s | 0.015 ms |
Reinforcement Learning Controller: A neural network-based reinforcement learning controller can achieve stable nominal voltage without overshoot when the current is zero (as shown in Figure 5b). However, when applying CC/CV control, voltage fluctuations occur due to a lack of understanding of the existing current, resulting in unstable and decreasing current during battery charging. The loading speed is the fastest, but there is ringing phenomenon, which can damage electronic components and shorten battery life. In practical applications, it is necessary to consider circulating current to improve performance.
Fuzzy controller: designed to maintain stable charging, with longer charging and stabilization times, but no spikes in current and voltage control conversion (as shown in Figure 5c). Although safer, it is slower and can be optimized by adjusting inference rules.
PID controller: It responds quickly to errors, but has significant interference and overcurrent phenomenon during voltage to current transition (as shown in Figure 5b). Its performance is moderate and does not rely on operator experience.
Performance indicator analysis: Analyze the voltage control action of the input DC/DC converter, calculate the root mean square (RMS) value, and find that the RMS value of the reinforcement learning controller is slightly higher, indicating that its control action changes more significantly and is sensitive to small changes. The evaluation of the RMS value of the battery current for all controllers resulted in a constant value of 0.3A, indicating that despite differences in control strategies and actions, all controllers were able to maintain the output current within the expected range. This means that although the voltage control actions of the reinforcement learning controller vary greatly, it does not affect the stability and consistency of the output current, which is crucial for the safe and efficient operation of the system.
4. Discussion on the results of lithium-ion battery charging control strategy
Controller simulation execution time and output variation: The results in Table 2 were obtained from simulations where each controller was configured for 6 seconds, with the reinforcement learning (RL) controller having the longest execution time of 21046 seconds. Compared with other controllers used to regulate the voltage and current of lithium-ion batteries, the RL controller has a larger output variation, and even in the face of small disturbances, ringing phenomena (high-frequency oscillation control actions) may occur. When applied to actual power electronic devices, it can cause overheating of battery and converter components, sensor noise, and shortened battery life, with a voltage variation of 0.02V. To improve the controller and reduce vibration, a low-pass filter can be added to the controller output, the neural network architecture can be modified, or a hybrid control method (such as PID or fuzzy controller to adjust the neural network output) can be used.
Controller adaptability and performance characteristics: RL controllers have high adaptability to different situations, as they are more sensitive to interference. However, output changes indicate the need to incorporate more parameters into agent training to improve their ability to adapt to more scenarios, increase charging efficiency, and avoid possible instability. Fuzzy controllers aim to avoid high current peaks and slow charging speeds due to their inference rules; The PID controller responds to errors by modifying the output signal and can adapt to dynamic changes.
Controller output performance evaluation: The root mean square (RMS) value was used to evaluate the output performance of three controllers, and the results showed that the average voltage maintained by all controllers was similar to the nominal voltage of the battery, which is crucial for preventing overcharging and avoiding battery overheating. In terms of current control, all controllers can reach the reference value in a very short response time without overshoot, and have good tolerance to external interference. In terms of voltage control, the controller can continuously reach the reference value, gradually reducing the current until it reaches zero to complete the charging cycle. However, it should be noted that RL based controllers may experience slight oscillations in the current during the final stage of charging, fluctuating between the current value and zero, and may require additional adjustments to improve stability.
5. Summary
PID controller
Advantages: Known for its simplicity and effectiveness, it performs well in regulating voltage and current, and can provide fast and stable response in many applications.
Disadvantage: Relatively weak adaptability in dealing with complex battery dynamic characteristics, making it difficult to achieve highly dynamic optimization.
Fuzzy controller
Advantages: By defining language variables and inference rules, experience is integrated into the control process, which can better handle battery nonlinearity and uncertainty, enabling the system to adapt to specific situations and perform stably under different operating conditions. Adjusting the inference rules according to application requirements can optimize charging performance to a certain extent.
Disadvantages: The design relies heavily on rules and experience, the development is complex, and the charging speed is relatively slow, which may not meet the high requirements for fast charging in application scenarios.
Controller based on reinforcement learning
Advantages: It can learn the optimal strategy through interaction with the system, has strong processing ability for small disturbances, can dynamically adapt to changing conditions, continuously optimize performance, and has high adaptability to different load conditions. In complex and ever-changing battery charging scenarios, it can effectively improve charging accuracy and efficiency, especially suitable for applications that require high charging flexibility and adaptability.
Disadvantages: Requires a large amount of computing resources, long training time, such as the longest execution time in this study. The training process is complex and requires careful design of reward functions and adjustment of neural network architecture, otherwise unstable phenomena such as ringing problems may occur. To improve accuracy, a more complex neural network architecture is required, which further increases the computational burden and simulation time.





