Cross-lingual alignment and completion of Wikipedia templates

by Gosse Bouma, Sergio Duarte Torres and Zahurul Islam. 

For many languages, the size of Wikipedia is an order of magnitude smaller than the English Wikipedia. We present a method for cross-lingual alignment of template and infobox attributes in Wikipedia. The alignment is used to add and complete templates and infoboxes in one language with information derived from Wikipedia in another language. We show that alignment between English and Dutch Wikipedia is accurate and that the result can be used to expand the number of template attribute-value pairs in Dutch Wikipedia by 50%. Furthermore, the alignment provides valuable information for normalization of template and attribute names and can be used to detect potential inconsistencies. Read the paper.

About The Author

Visit Us On FacebookVisit Us On LinkedinCheck Our Feed