Towards a taking into account of several aspects of information needs in the models of the information retrieval : propagation of metadata on the World Wide Web

In this thesis, which is part and parcel of the more general context of web information retrieval, we consider the issue of thematic and non thematic page indexation, with particular focus on page typology. We suggest a page characterization method in two steps. The first one, named homogeneous corpus extraction, aims at connecting several pages sharing similar features. The second one, called semi-automatic metadata assignment within each homogeneous corpus, is based on propagation : to begin with, only a small proportion of all ressources is manually qualified, ressources information is then propagated to other ressources. Methodologically, the homogeneous corpus extraction is grounded on hypertext link analysis. More precisely, it uses the "co-citation" principle. This principle is a Web transposition of the well-known scientometry co-citation method.

Data and Resources

Additional Info

Field Value
Source https://theses.hal.science/tel-00839565
Author Prime-Claverie, Camille
Maintainer CCSD
Last Updated May 10, 2026, 12:58 (UTC)
Created May 10, 2026, 12:58 (UTC)
Identifier NNT: 2004EMSE0020
Language fr
Rights https://about.hal.science/hal-authorisation-v1/
contributor Département Réseaux, Information, Multimédia (RIM-ENSMSE) ; École des Mines de Saint-Étienne (Mines Saint-Étienne MSE) ; Institut Mines-Télécom [Paris] (IMT)-Institut Mines-Télécom [Paris] (IMT)-Centre G2I
creator Prime-Claverie, Camille
date 2004-11-26T00:00:00
harvest_object_id 7f64c029-8002-4f2b-a1e8-585324da1510
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2026-01-19T00:00:00
set_spec type:THESE