In this thesis, which is part and parcel of the more general context of web information retrieval, we consider the issue of thematic and non thematic page indexation, with particular focus on page typology. We suggest a page characterization method in two steps. The first one, named homogeneous corpus extraction, aims at connecting several pages sharing similar features. The second one, called semi-automatic metadata assignment within each homogeneous corpus, is based on propagation : to begin with, only a small proportion of all ressources is manually qualified, ressources information is then propagated to other ressources. Methodologically, the homogeneous corpus extraction is grounded on hypertext link analysis. More precisely, it uses the "co-citation" principle. This principle is a Web transposition of the well-known scientometry co-citation method.