EXTENDED BAG-OF-WORDS FORMALISM FOR IMAGE CLASSIFICATION

In this dissertation, we have addressed the problem of representing images based on their visual information. Our aim is content-based concept detection in images and videos, with a novel representation that enriches the Bag-of-Words model. Relying on the quantization of highly discriminant local descriptors by a codebook, and the aggregation of those quantized descriptors into a single pooled feature vector, the Bag-of-Words model has emerged as the most promising approach for image classification. We propose BossaNova, a novel image representation which offers a more information-preserving pooling operation based on a distance-to-codeword distribution. The experimental evaluations on many challenging image classification benchmarks, such as ImageCLEF Photo Annotation, MIRFLICKR, PASCAL VOC and 15-Scenes, have shown the advantage of BossaNova when compared to traditional techniques, even without using complex combinations of different local descriptors. An extension of our approach has also been studied. It concerns the combination of BossaNova representation with another representation very competitive based on Fisher Vectors. The results consistently reaches other state-of-the-art representations in many datasets. It also experimentally demonstrate the complementarity of the two approaches. This study allowed us to achieve, in the competition ImageCLEF 2012 Flickr Photo Annotation Task, the 2nd among the 28 visual submissions. Finally, we have explored our BossaNova representation in the challenging real-world application of pornography detection. Once again, the results validated the relevance of our approach compared to standard techniques on a real application.

Data and Resources

Additional Info

Field Value
Source https://theses.hal.science/tel-00958547
Author Avila, Sandra
Maintainer CCSD
Last Updated May 6, 2026, 01:41 (UTC)
Created May 6, 2026, 01:41 (UTC)
Identifier tel-00958547
Language en
Rights https://about.hal.science/hal-authorisation-v1/
contributor Machine Learning and Information Access (MLIA) ; Laboratoire d'Informatique de Paris 6 (LIP6) ; Université Pierre et Marie Curie - Paris 6 (UPMC)-Centre National de la Recherche Scientifique (CNRS)-Université Pierre et Marie Curie - Paris 6 (UPMC)-Centre National de la Recherche Scientifique (CNRS)
creator Avila, Sandra
date 2013-06-14T00:00:00
harvest_object_id 14eb1cae-bf8e-4fb5-9513-7a036ecdb617
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2025-08-12T00:00:00
set_spec type:THESE