We employed GTA-5 (Grand Theft Auto) and FIFA (Federation International Football Association) for collecting the games action dataset. We asked the players to play the games and record the same action from multiple views. Note that GTA and FIFA allow users to record the actions from multiple angles with real-looking scenes and different realistic camera motions.
In total, we collected seven human actions including cycling, fighting, soccer kicking, running, walking, shooting, and skydiving. Due to the availability of plenty of soccer kicking in FIFA games, we collected kicking from FIFA and the rest of the actions are collected from GTA-5. Although in our current approach we are only using aerial game video, for more complete dataset purposes, we captured both ground and aerial video pairs i.e., the same action frames captured from both aerial and ground cameras.
For each action, our dataset contains 200 videos (100 ground and 100 aerial) with a total of 1400 videos for seven actions. Note that most of the scenes and interactions in the video games are biased towards actions related to fighting, shooting, walking and running, etc. Therefore, employing game videos to improve action recognition in real-world videos is not trivial. Therefore, in this paper, we proposed a unified approach to combine games and real videos employing disjoint multitask learning.
