Conceptual

Spatially-Guided Temporal Aggregation for Event-RGB Optical Flow Fusion

A cross-modal optical-flow framework that uses the spatially dense RGB frame modality to guide the aggregation of the temporally dense but spatially sparse event-camera modality: it forms an event-enhanced frame representation as a guide, directs how sparse event motion features are aggregated over time, adds a transformer module to propagate spatially rich frame information into the sparse event features, and uses a mix-fusion encoder to extract joint spatiotemporal features.