Technical fault management in grid computing

The construction of grid computing is one of the major research on networked computer systems . The main construction of a grid computing is to provide the concepts and system software components suitable to aggregate computing resources ( processors, memory , and also network ) within a grid of data processing for make (eventually) a global IT infrastructure simulations, data processing and industrial process control . This infrastructure can potentially be used in all fields of scientific research, industrial research and operational activities ( new processes and products , instrumentation, etc. . ) , in the evolution of information systems , web and multimedia. The production quality grids assume a mastery of problems of reliability, enhanced security through better access control and better protection against attacks, fault tolerance and failure prevention, all these properties to result in grid infrastructure computer safe operation. In this thesis we propose to conduct research into the problems of automated fault management, the main objective is to hide as much as possible such failures, ultimately making them transparent to applications, so that from the point of view applications, the grid infrastructure operates almost continuously . We have developed a new self- adaptive hierarchical algorithm to ensure fault tolerance in computational grids. This protocol is based on the hierarchical architecture of grid computing. In each cluster, we defined a coordinator called leader, whose role is to coordinate intra-cluster and ensure the role of intermediary between processes belonging to different clusters process. To save the state of inter-cluster process, the adaptive protocol uses pessimistic message logging protocol based on the issuer. Inside the cluster, the protocol used depends on the frequency of messages. From a maximum threshold determined by the density of communications frequency, non-blocking coordinated checkpoint protocol is used. If the number of messages in the cluster is low , messages are saved using the pessimistic message logging protocol.

Data and Resources

Additional Info

Field Value
Source https://theses.hal.science/tel-00931839
Author Ndiaye, Ndeye Massata
Maintainer CCSD
Last Updated May 7, 2026, 09:34 (UTC)
Created May 7, 2026, 09:34 (UTC)
Identifier tel-00931839
Language fr
Rights https://about.hal.science/hal-authorisation-v1/
contributor Large-Scale Distributed Systems and Applications (Regal) ; Laboratoire d'Informatique de Paris 6 (LIP6) ; Université Pierre et Marie Curie - Paris 6 (UPMC)-Centre National de la Recherche Scientifique (CNRS)-Université Pierre et Marie Curie - Paris 6 (UPMC)-Centre National de la Recherche Scientifique (CNRS)-Inria Paris-Rocquencourt ; Institut National de Recherche en Informatique et en Automatique (Inria)-Institut National de Recherche en Informatique et en Automatique (Inria)
creator Ndiaye, Ndeye Massata
date 2013-09-17T00:00:00
harvest_object_id 1e7d985d-6615-4238-bc1d-7ff649f7c6b9
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2025-08-20T00:00:00
set_spec type:THESE