Chapter 6: Recognition
- The face recognition and detection datasets listed in Table 6.1 and Masi, Wu et al. (2018).
• The Caltech pedestrian detection benchmark (Dollár, Belongie, and Perona 2010) and person detection subtasks in datasets such as KITTI, http://www.cvlibs.net/datasets/kitti (Geiger, Lenz, and Urtasun 2012) and Cityscapes, https://www.cityscapes-dataset.com (Cordts, Omran et al. 2016)
- Table 6.2 lists datasets and benchmarks for image classification, general object detection, and segmentation. Two recent workshops that highlight the latest results on these datasets are the Robust Vision Challenge Zendel et al. (2020) (see Table C.1) and the COCO + LVIS Joint Recognition Challenge Kirillov, Lin et al. (2020).
Datasets and benchmarks for fine-grained category recognition can be found at the CVPR Workshop on Fine-Grained Visual Categorization, https://sites.google.com/view/fgvc8 as well as some of the papers on this topic discussed in Section 6.2.2.
- Table 6.3 lists some datasets for video understanding and action recognition.
- Table 6.4 lists some widely used datasets for vision and language research, which includes image captioning, dense annotation, visual question answering, and visual dialog.