Image Classification with the Fisher Vector: Theory and Practice

A standard approach to describe an image for classification and retrieval purposes is to extract a set of local patch descriptors, encode them into a high dimensional vector and pool them into an image-level signature. The most common patch encoding strategy consists in quantizing the local descriptors into a finite set of prototypical elements. This leads to the popular Bag-of-Visual words (BOV) representation. In this work, we propose to use the Fisher Kernel framework as an alternative patch encoding strategy: we describe patches by their deviation from an ''universal'' generative Gaussian mixture model. This representation, which we call Fisher Vector (FV) has many advantages: it is efficient to compute, it leads to excellent results even with efficient linear classifiers, and it can be compressed with a minimal loss of accuracy using product quantization. We report experimental results on five standard datasets -- PASCAL VOC 2007, Caltech 256, SUN 397, ILSVRC 2010 and ImageNet10K -- with up to 9M images and 10K classes, showing that the FV framework is a state-of-the-art patch encoding technique.

Data and Resources

Additional Info

Field Value
Source https://inria.hal.science/hal-00779493
Author Sanchez, Jorge, Perronnin, Florent, Mensink, Thomas, Verbeek, Jakob
Maintainer CCSD
Last Updated May 10, 2026, 18:12 (UTC)
Created May 10, 2026, 18:12 (UTC)
Identifier Report N°: RR-8209
Language en
Rights https://about.hal.science/hal-authorisation-v1/
contributor Facultad de Matemática, Astronomía y Física [Cordoba] (FaMAF) ; Universidad Nacional de Córdoba [Argentina]
creator Sanchez, Jorge
date 2013-05-10T00:00:00
harvest_object_id 6977a6a8-c159-44cb-8adf-b2b2614d8170
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2025-09-27T00:00:00
set_spec type:REPORT