Bilingual lexicon extraction from comparable corpora

Most work in bilingual lexicon acquisition from comparable corpora are based on the distributional hypothesis that has been extended to the bilingual scenario. Hence, two words are more likely to be translation of each other if they appear in the same lexical contexts. This assumption presupposes a clear and rigorous definition of context and a thorough knowledge of contextual clues. However, the complexity and specificity of each language make the formulation of such a definition that ensures effective extraction of translation pairs in all cases not easy. All the difficulty lies in how to define, extract and compare these contexts in order to build reliable bilingual lexicons. We strive throughout the different chapters of this thesis to try to understand this notion of context, and then extend and adapt it to improve the quality of bilingual lexicons. The first part of contributions aims at improving the standard approach considered as a baseline in the community. Thus, we propose several ways to consider the context for better words characterization. In the second part of the contributions, we first present an approach that aims to improve the extended approach. Then, a method called Q-Align directly inspired from question/answering systems is presented. Finally, we present several mathematical transforms and thus multiple vector space representations to focus primarily on the ones we have chosen to develop a new alignment method.

Data and Resources

Additional Info

Field Value
Source https://theses.hal.science/tel-00946914
Author Hazem, Amir
Maintainer CCSD
Last Updated May 6, 2026, 09:33 (UTC)
Created May 6, 2026, 09:33 (UTC)
Identifier tel-00946914
Language fr
Rights https://about.hal.science/hal-authorisation-v1/
contributor Traitement Automatique du Langage Naturel (LS2N - équipe TALN) ; Laboratoire des Sciences du Numérique de Nantes (LS2N) ; Université de Nantes - UFR des Sciences et des Techniques (UN UFR ST) ; Université de Nantes (UN)-Université de Nantes (UN)-École Centrale de Nantes (ECN)-Centre National de la Recherche Scientifique (CNRS)-IMT Atlantique (IMT Atlantique) ; Institut Mines-Télécom [Paris] (IMT)-Institut Mines-Télécom [Paris] (IMT)-Université de Nantes - UFR des Sciences et des Techniques (UN UFR ST) ; Université de Nantes (UN)-Université de Nantes (UN)-École Centrale de Nantes (ECN)-Centre National de la Recherche Scientifique (CNRS)-IMT Atlantique (IMT Atlantique) ; Institut Mines-Télécom [Paris] (IMT)-Institut Mines-Télécom [Paris] (IMT)
creator Hazem, Amir
date 2013-10-11T00:00:00
harvest_object_id b11d417b-0249-46fa-881a-4580058db95c
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2026-03-31T00:00:00
set_spec type:THESE