Multi-Modal Multi-Task Masked Autoencoder: A Simple, Flexible, and Effective ViT Pre-Training Strategy
Source: Deephub Imba This article is about 1000 words long and is recommended to read in 4 minutes. This article introduces a simple, flexible, and effective pre-training strategy for ViT. MAE is a ViT that uses a self-supervised pre-training strategy, masking patches in the input image and then predicting the missing areas for sub-supervision and … Read more