Masked AutoDecoder is Effective Multi-Task Vision Generalist

Qiu, Han; Huang, Jiaxing; Gao, Peng; Lu, Lewei; Zhang, Xiaoqin; Lu, Shijian

Computer Science > Computer Vision and Pattern Recognition

arXiv:2403.07692 (cs)

[Submitted on 12 Mar 2024 (v1), last revised 14 Mar 2024 (this version, v2)]

Title:Masked AutoDecoder is Effective Multi-Task Vision Generalist

Authors:Han Qiu, Jiaxing Huang, Peng Gao, Lewei Lu, Xiaoqin Zhang, Shijian Lu

View PDF HTML (experimental)

Abstract:Inspired by the success of general-purpose models in NLP, recent studies attempt to unify different vision tasks in the same sequence format and employ autoregressive Transformers for sequence prediction. They apply uni-directional attention to capture sequential dependencies and generate task sequences recursively. However, such autoregressive Transformers may not fit vision tasks well, as vision task sequences usually lack the sequential dependencies typically observed in natural languages. In this work, we design Masked AutoDecoder~(MAD), an effective multi-task vision generalist. MAD consists of two core designs. First, we develop a parallel decoding framework that introduces bi-directional attention to capture contextual dependencies comprehensively and decode vision task sequences in parallel. Second, we design a masked sequence modeling approach that learns rich task contexts by masking and reconstructing task sequences. In this way, MAD handles all the tasks by a single network branch and a simple cross-entropy loss with minimal task-specific designs. Extensive experiments demonstrate the great potential of MAD as a new paradigm for unifying various vision tasks. MAD achieves superior performance and inference efficiency compared to autoregressive counterparts while obtaining competitive accuracy with task-specific models. Code will be released.

Comments:	Accepted by CVPR 2024
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2403.07692 [cs.CV]
	(or arXiv:2403.07692v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2403.07692

Submission history

From: Han Qiu [view email]
[v1] Tue, 12 Mar 2024 14:36:52 UTC (503 KB)
[v2] Thu, 14 Mar 2024 18:54:46 UTC (503 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Masked AutoDecoder is Effective Multi-Task Vision Generalist

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Masked AutoDecoder is Effective Multi-Task Vision Generalist

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators