Approximation of OLAP queries on data warehouses

We study the approximate answers to OLAP queries on data warehouses. We consider the relative answers to OLAP queries on a schema, as distributions with the L1 distance and approximate the answers without storing the entire data warehouse. We first introduce three specific methods: the uniform sampling, the measure-based sampling and the statistical model. We introduce also an edit distance between data warehouses with edit operations adapted for data warehouses. Then, in the OLAP data exchange, we study how to sample each source and combine the samples to approximate any OLAP query. We next consider a streaming context, where a data warehouse is built by streams of different sources. We show a lower bound on the size of the memory necessary to approximate queries. In this case, we approximate OLAP queries with a finite memory. We describe also a method to discover the statistical dependencies, a new notion we introduce. We are looking for them based on the decision tree. We apply the method to two data warehouses. The first one simulates the data of sensors, which provide weather parameters over time and location from different sources. The second one is the collection of RSS from the web sites on Internet.

Data and Resources

Additional Info

Field Value
Source https://theses.hal.science/tel-00905292
Author Cao, Phuong Thao
Maintainer CCSD
Last Updated May 8, 2026, 05:24 (UTC)
Created May 8, 2026, 05:24 (UTC)
Identifier NNT: 2013PA112091
Language en
Rights https://about.hal.science/hal-authorisation-v1/
contributor Laboratoire de Recherche en Informatique (LRI) ; Université Paris-Sud - Paris 11 (UP11)-CentraleSupélec-Centre National de la Recherche Scientifique (CNRS)
creator Cao, Phuong Thao
date 2013-06-20T00:00:00
harvest_object_id 326a5c0a-67a4-4459-b28e-b2e0bfd70c03
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2026-03-31T00:00:00
set_spec type:THESE