Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing

Xu, Pengcheng; Jiang, Boyuan; Hu, Xiaobin; Luo, Donghao; He, Qingdong; Zhang, Jiangning; Wang, Chengjie; Wu, Yunsheng; Ling, Charles; Wang, Boyu

Computer Science > Computer Vision and Pattern Recognition

arXiv:2411.15843 (cs)

[Submitted on 24 Nov 2024 (v1), last revised 26 Nov 2024 (this version, v2)]

Title:Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing

Authors:Pengcheng Xu, Boyuan Jiang, Xiaobin Hu, Donghao Luo, Qingdong He, Jiangning Zhang, Chengjie Wang, Yunsheng Wu, Charles Ling, Boyu Wang

View PDF HTML (experimental)

Abstract:Leveraging the large generative prior of the flow transformer for tuning-free image editing requires authentic inversion to project the image into the model's domain and a flexible invariance control mechanism to preserve non-target contents. However, the prevailing diffusion inversion performs deficiently in flow-based models, and the invariance control cannot reconcile diverse rigid and non-rigid editing tasks. To address these, we systematically analyze the \textbf{inversion and invariance} control based on the flow transformer. Specifically, we unveil that the Euler inversion shares a similar structure to DDIM yet is more susceptible to the approximation error. Thus, we propose a two-stage inversion to first refine the velocity estimation and then compensate for the leftover error, which pivots closely to the model prior and benefits editing. Meanwhile, we propose the invariance control that manipulates the text features within the adaptive layer normalization, connecting the changes in the text prompt to image semantics. This mechanism can simultaneously preserve the non-target contents while allowing rigid and non-rigid manipulation, enabling a wide range of editing types such as visual text, quantity, facial expression, etc. Experiments on versatile scenarios validate that our framework achieves flexible and accurate editing, unlocking the potential of the flow transformer for versatile image editing.

Comments:	Project Page: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as:	arXiv:2411.15843 [cs.CV]
	(or arXiv:2411.15843v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2411.15843

Submission history

From: Pengcheng Xu [view email]
[v1] Sun, 24 Nov 2024 13:48:16 UTC (20,698 KB)
[v2] Tue, 26 Nov 2024 07:56:41 UTC (20,699 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators