--- title: "Testing a cluster count with matched nulls" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Testing a cluster count with matched nulls} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") has_mclust <- requireNamespace("mclust", quietly = TRUE) ``` ## The question A clustering method applied to continuous data will report some number of clusters whether or not any exist. So when an analysis announces that a dataset contains, say, four types, the number may describe the population, or only the shape of the data under that method. `matchednull` separates the two. ## The idea: a null twin `copula_null()` builds a synthetic twin of your data that keeps everything innocent and removes the one thing in question: * every variable's distribution is preserved **exactly** (the twin reuses the real values, reshuffled); * the correlation matrix is preserved to within sampling error; * all remaining dependence is Gaussian, so the twin contains **no cluster structure by construction**. If your pipeline finds as many clusters in the twins as in the real data, those clusters were the data's shape, not its people. ```{r} library(matchednull) set.seed(1) x <- matrix(rnorm(400 * 3), 400, 3) %*% chol(matrix(c(1, .5, .3, .5, 1, .4, .3, .4, 1), 3, 3)) twin <- copula_null(x) all(sort(twin[, 1]) == sort(x[, 1])) # margins: identical round(cor(x) - cor(twin), 2) # correlations: close ``` ## The test `matched_null_test()` takes your data and **your own pipeline**, wrapped as a function that returns one number (typically the selected number of clusters). It runs the identical pipeline on the real data and on `R` twins, and asks whether the real answer stands out. ```{r, eval = has_mclust} suppressPackageStartupMessages(library(mclust)) pick_k <- function(d) Mclust(d, G = 1:4, modelNames = "VVV", verbose = FALSE)$G # 1. Typeless data: the test should stay quiet. set.seed(7) matched_null_test(x, pick_k, R = 30) ``` ```{r, eval = has_mclust} # 2. Two genuine types, hidden in the dependence structure # (identical margins, opposite correlation orientation). set.seed(42) z <- sample(2, 400, replace = TRUE) X <- matrix(rnorm(400 * 4), 400, 4) L1 <- chol(matrix(c(1, .85, .85, 1), 2, 2)) L2 <- chol(matrix(c(1, -.85, -.85, 1), 2, 2)) X[z == 1, 1:2] <- X[z == 1, 1:2] %*% L1 X[z == 1, 3:4] <- X[z == 1, 3:4] %*% L1 X[z == 2, 1:2] <- X[z == 2, 1:2] %*% L2 X[z == 2, 3:4] <- X[z == 2, 3:4] %*% L2 set.seed(7) matched_null_test(X, pick_k, R = 30) ``` The first call returns *null-like*: whatever clustering the pipeline reports is reproduced by twins with no types. The second returns *exceeds the null*: the grouping lives in structure the twins cannot carry. ## Plugging in your own pipeline `cluster_fn` can wrap anything: a k-means heuristic, a published typology's exact workflow, or any scalar measure of clustering strength. The null is defined at the level of the data, not the pipeline. Two practical notes: * **Set a seed before calling** — reproducibility is deliberately left to the caller. * **Check the margins first.** The one regime the null cannot flag is types so separated that they are visible in the margins and covariance themselves (strongly bimodal variables); inspect the margins for pronounced multimodality before relying on the count test, and tie-break granular (Likert-type) scales before formal unimodality tests. ## Reference The method, its positive controls, and its false-positive calibration are described in the accompanying paper: *Types Without Taxa: A Covariance-Matched-Null Multiverse Test of Categorical versus Continuous Personality Structure*.