Methodology for assessing consistency between multiple representations for spatial databases integration.<br />An approach combining the use of metadata and machine learning.

Nowadays most databases are run independently. An independence that leads to a series ofproblems: repeated efforts of maintenance and updating, difficulty in proceeding with an analysis atvarious levels and no guarantee of coherence between sources.Joint management of these sources requires them to be integrated in order to define the explicitlinks between the various bases and to provide a unified vision. Our thesis deals with this issue. Itconcentrates in particular on the means of relating data and of assessing coherence between multiplerepresentations. We have sought to systematically analyse each difference in representation betweenmatching data so as to determine whether it results from different criteria used for data capture or fromerrors in the capture itself, the aim being to ensure coherent data integration.In order to study the conformity of representations, we suggest exploiting existing databasespecifications. These documents describe specific selection and modelling rules for objects. They arereference metadata used to determine whether representations are equivalent or incoherent. But theiruse is insufficient since specifications described in a natural language can be imprecise or incomplete.So the data contained in the bases is a second interesting source of knowledge. If one uses machinelearning techniques to analyse how they tally, it becomes possible to establish evaluation rules thatenable a justification of the conformity of representations.The methodology we put forward is based upon these elements. It consists in a coherenceevaluation process and a knowledge acquisition proceeding. The process comprises several steps: dataenrichment, intra-base control, matching, inter-bases control, and the final assessment. Each of thesesteps exploits knowledge inferred from the specifications or induced from the data through learning.The benefit of using machine learning techniques is twofold: not only does it enable to acquireevaluation rules, it also reveals the discrepancy tolerated in the data when compared to the writtenspecifications.This approach has been carried out on NGI databases that showed different levels of detail.

Data and Resources

Additional Info

Field Value
Source https://theses.hal.science/tel-00085693
Author Sheeren, David
Maintainer CCSD
Last Updated May 9, 2026, 22:18 (UTC)
Created May 9, 2026, 22:18 (UTC)
Identifier tel-00085693
Language fr
Rights https://about.hal.science/hal-authorisation-v1/
contributor Conception Objet et Généralisation de l'Information Topographique (COGIT) ; Ecole nationale des sciences géographiques (ENSG) ; Institut géographique national [IGN] (IGN)-Institut géographique national [IGN] (IGN)
creator Sheeren, David
date 2005-05-20T00:00:00
harvest_object_id b11aae8d-f8b1-4d81-baef-7ec663907655
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2026-04-30T00:00:00
set_spec type:THESE