Two comparable corpora of German newspaper text gathered on the web: Bild & Die Zeit

This technical report documents the creation of two comparable corpora of German newspaper text, focused on the daily tabloid Bild and the weekly newspaper Die Zeit. Two specialized crawlers and corpus builders were designed in order to crawl the domain names bild.de and zeit.de with the objective of gathering as many complete articles as possible. A high content quality was made possible by the specially designed boilerplate removal and metadata recording code. As a result, two separate corpora were created. Currently, the last version for Bild is from 2011 and the last version for Die Zeit is from early 2013. The corpora feature a total of respectively 60 476 and 134 222 articles. Whereas the crawler designed for Bild has been discontinued due to frequent layout changes on the website, the other one concerning Die Zeit is still actively maintained, its code has been made available under an open source license.

Data and Resources

Additional Info

Field Value
Source https://shs.hal.science/halshs-00844541
Author Barbaresi, Adrien
Maintainer CCSD
Last Updated May 10, 2026, 08:42 (UTC)
Created May 10, 2026, 08:42 (UTC)
Identifier halshs-00844541
Language en
Rights https://about.hal.science/hal-authorisation-v1/
contributor Interactions, Corpus, Apprentissages, Représentations (ICAR) ; École normale supérieure de Lyon (ENS de Lyon) ; Université de Lyon-Université de Lyon-Université Lumière - Lyon 2 (UL2)-Centre National de la Recherche Scientifique (CNRS)
creator Barbaresi, Adrien
date 2013-07-15T00:00:00
harvest_object_id db6bd878-12b7-46e6-bdce-01d6aa257595
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2025-10-13T00:00:00
set_spec type:UNDEFINED