Classification of coding and non-coding RNAs

The work described in this thesis is part of the analysis of biological phenomena using computers, id est bioinformatics. More precisely, we are interested in nucleic sequence analysis. In this context, our work is splitted in two parts: identification of coding sequences and identification of non-coding sequences that share a common structure such as non-coding RNAs. The main feature of our methods, protea and carnac, is to deal with poorly conserved sequences without the need to align them. Our methods rely on the same comparative analysis scheme to detect evolutionary patterns that are globally coherent between all sequences. \protea and \carnac have been submitted on several reference benchmarks and have reached significative results. We also present two collaborative projects that involve protea and carnac. magnolia is multiple alignement software designed to align nucleic sequences according to their conserved function predicted by protea and/or carnac. The second collaborative project is a software pipeline to automatically annotate genomes by comparative genomics.

Data and Resources

Additional Info

Field Value
Source https://theses.hal.science/tel-00401991
Author Fontaine, Arnaud
Maintainer CCSD
Last Updated May 10, 2026, 15:18 (UTC)
Created May 10, 2026, 15:18 (UTC)
Identifier tel-00401991
Language fr
Rights https://about.hal.science/hal-authorisation-v1/
contributor Laboratoire d'Informatique Fondamentale de Lille (LIFL) ; Université de Lille, Sciences et Technologies-Institut National de Recherche en Informatique et en Automatique (Inria)-Université de Lille, Sciences Humaines et Sociales-Centre National de la Recherche Scientifique (CNRS)
creator Fontaine, Arnaud
date 2009-03-31T00:00:00
harvest_object_id 5201a6e9-67d6-4a96-ac97-3b7730d87e57
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2025-02-26T00:00:00
set_spec type:THESE