Dense semantic labeling of sub-decimeter resolution images with convolutional neural networks

Volpi, Michele; Tuia, Devis

doi:10.1109/TGRS.2016.2616585

Computer Science > Computer Vision and Pattern Recognition

arXiv:1608.00775 (cs)

[Submitted on 2 Aug 2016 (v1), last revised 10 Oct 2016 (this version, v2)]

Title:Dense semantic labeling of sub-decimeter resolution images with convolutional neural networks

Authors:Michele Volpi, Devis Tuia

View PDF

Abstract:Semantic labeling (or pixel-level land-cover classification) in ultra-high resolution imagery (< 10cm) requires statistical models able to learn high level concepts from spatial data, with large appearance variations. Convolutional Neural Networks (CNNs) achieve this goal by learning discriminatively a hierarchy of representations of increasing abstraction.
In this paper we present a CNN-based system relying on an downsample-then-upsample architecture. Specifically, it first learns a rough spatial map of high-level representations by means of convolutions and then learns to upsample them back to the original resolution by deconvolutions. By doing so, the CNN learns to densely label every pixel at the original resolution of the image. This results in many advantages, including i) state-of-the-art numerical accuracy, ii) improved geometric accuracy of predictions and iii) high efficiency at inference time.
We test the proposed system on the Vaihingen and Potsdam sub-decimeter resolution datasets, involving semantic labeling of aerial images of 9cm and 5cm resolution, respectively. These datasets are composed by many large and fully annotated tiles allowing an unbiased evaluation of models making use of spatial information. We do so by comparing two standard CNN architectures to the proposed one: standard patch classification, prediction of local label patches by employing only convolutions and full patch labeling by employing deconvolutions. All the systems compare favorably or outperform a state-of-the-art baseline relying on superpixels and powerful appearance descriptors. The proposed full patch labeling CNN outperforms these models by a large margin, also showing a very appealing inference time.

Comments:	Accepted in IEEE Transactions on Geoscience and Remote Sensing, 2016
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1608.00775 [cs.CV]
	(or arXiv:1608.00775v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1608.00775
Related DOI:	https://doi.org/10.1109/TGRS.2016.2616585

Submission history

From: Michele Volpi Michele Volpi [view email]
[v1] Tue, 2 Aug 2016 11:33:44 UTC (8,738 KB)
[v2] Mon, 10 Oct 2016 15:07:33 UTC (9,342 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Dense semantic labeling of sub-decimeter resolution images with convolutional neural networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Dense semantic labeling of sub-decimeter resolution images with convolutional neural networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators