Conceptual

Cross-Layer Actor-Critic Reinforcement Learning for Adaptive Wireless Video Streaming

Formulates adaptive wireless video streaming as an infinite-horizon discounted Markov decision process that incorporates past and lower-layer network information, then solves it with an enhanced asynchronous advantage actor-critic (eA3C) that jointly trains policy and value networks. Adds continual-learning online-tuning schemes that adapt the trained policy to an individual user's real-time data, improving quality of experience.