Technical gazette, Vol. 33 No. 5, 2026.
Original scientific paper
https://doi.org/10.17559/TV-20250916002994
Adaptive Firewall Policy Optimization Based on Multi-Objective Reinforcement Learning
Xiaoyu Zhao
; Geely University of China, Chengdu, Sichuan, 610000, China
Lei Bu
; Zhejiang Dahua Technology Co., Ltd., Chengdu, Sichuan, 610000, China
*
Shufang He
; Geely University of China, Chengdu, Sichuan, 610000, China
Xing Yang
; Geely University of China, Chengdu, Sichuan, 610000, China
* Corresponding author.
Abstract
Traditional static firewalls struggle to manage complex, dynamically evolving network traffic, often leading to rule redundancy, conflicts, and latency degradation. To address these challenges, this study proposes an intelligent firewall rule optimization framework based on Proximal Policy Optimization (PPO). The framework models the rule scheduling task as a multi-objective reinforcement learning problem, integrating system latency, throughput, and false positive rate into a composite reward function. A custom simulation environment compatible with OpenAI Gym is developed to represent firewall states and actions, while visualization tools such as reward curves, clip fraction, and action heatmaps enhance interpretability. Experimental evaluations demonstrate that the proposed PPO-based method outperforms baseline algorithms (Q-learning, DQN, A2C) in convergence speed, false positive reduction, and throughput improvement, achieving up to 60% faster stabilization and a 15% lower error rate. The approach offers a scalable and adaptive framework for real-time firewall policy management, contributing to the development of intelligent, self-optimizing network defense systems.
Keywords
firewall optimization; network security; performance evaluation; proximal policy optimization (PPO); reinforcement learning training interpretability
Hrčak ID:
350396
URI
Publication date:
31.8.2026.
Visits: 0 *