Conceptual

Detection-Fusion Model for Extracting Knowledge Graphs from Videos

A deep-learning model that annotates a video with a knowledge graph by first predicting the individuals (entities) present and then the relations between pairs of them, with an extension that incorporates background knowledge into graph construction — avoiding the pitfalls of natural-language video description.