Conceptual

Masked Autoencoding on Event-Camera Point Streams for Action Recognition

Treats an event-camera stream as a 3D point cloud in (x, y, time) and pre-trains a transformer encoder-decoder by masking and reconstructing event patches, learning representations that transfer to action and gesture recognition. Introduces an inlier plane-fitting model that selects noise-free patch centers, replacing farthest-point sampling, with PointNet patch embeddings and a Chamfer-distance reconstruction loss.