Home /Research /Learning Rich Features from RGB-D Images for Object Detection and\n Segmentation
PERCEPTION

Learning Rich Features from RGB-D Images for Object Detection and\n Segmentation

Saurabh Gupta, Ross Girshick, Pablo Arbeláez, Jitendra Malik

Year
2014
Citations
10
Access
Open access

Abstract

In this paper we study the problem of object detection for RGB-D images using\nsemantically rich image and depth features. We propose a new geocentric\nembedding for depth images that encodes height above ground and angle with\ngravity for each pixel in addition to the horizontal disparity. We demonstrate\nthat this geocentric embedding works better than using raw depth images for\nlearning feature representations with convolutional neural networks. Our final\nobject detection system achieves an average precision of 37.3%, which is a 56%\nrelative improvement over existing methods. We then focus on the task of\ninstance segmentation where we label pixels belonging to object instances found\nby our detector. For this task, we propose a decision forest approach that\nclassifies pixels in the detection window as foreground or background using a\nfamily of unary and binary tests that query shape and geocentric pose features.\nFinally, we use the output from our object detectors in an existing superpixel\nclassification framework for semantic scene segmentation and achieve a 24%\nrelative improvement over current state-of-the-art for the object categories\nthat we study. We believe advances such as those represented in this paper will\nfacilitate the use of perception in fields like robotics.\n

Keywords

Artificial intelligenceComputer visionComputer sciencePixelObject detectionSegmentationConvolutional neural networkEmbeddingObject (grammar)Pattern recognition (psychology)

Related papers

Browse all PERCEPTION papers