This report deals with the scientific problem of objects visual recognition within an observed scene. We propose an approache going from object's model building to the definition of a strategy for its future recognition. From the representation point of view, this methodology can represente the structure of the object as well as its appearance from multiple features. These last ones are used as attentional clues during the recognition stage. In this framework, the recognition of the object consists in instantiate it in the current scene. The recognition task is an active process of hypothesis generation/verification driven by a focusing principle. Focus acts on four levels of the "