Dataset accompanying:

Fisun, Roman (2026) Zusammengesetzte Indefinitpronomen in den slavischen Sprachen. Universität Regensburg

Version

1.0

Description

This dataset contains the input matrices and R code used for the hierarchical cluster analyses of the contextual distributions of 51 compound indefinite pronoun series presented in Chapter 11.

Contents

cluster_context.R
R script for the cluster analyses of semantic maps of types 1 and 2. The script calculates Hellinger and Spearman distances, performs Ward.D2 and average-linkage clustering, determines the optimal number of clusters using silhouette widths, and compares the cluster solutions using the Adjusted Rand Index.

context_matrix_map_1.csv
Series-by-context frequency matrix used for the analysis of semantic maps of type 1.

context_matrix_map_2.csv
Series-by-context frequency matrix used for the analysis of semantic maps of type 2.

For Croatian, the contexts 6A and 11A are represented under the original context labels 6 and 11, respectively; categories 6B and 11B are not included.

Data structure

* Rows represent the 51 compound indefinite pronoun series.
* The first column contains the series labels.
* The remaining columns represent semantic or distributional contexts.
* Cells contain non-negative context frequencies.

Usage

Load both CSV files into R as `matrix_map_1` and `matrix_map_2`, then run `cluster_context.R`.

Project website

https://fisun.org/ZIPs

Author

Roman Fisun
University of Regensburg
https://fisun.org

License

CC BY 4.0

Contact

roman@fisun.org
