MAEPose: learning human pose from raw mmWave video without labels
MAEPose pre-trains a video Transformer on unlabelled mmWave radar video with masked autoencoding, then decodes multi-frame joint heatmaps, reducing pose error by up to 22.1% against state-of-the-art baselines across three datasets.
Read the report
