From non-spatial univariate to spatial multivariate
University of Birmingham
Lund University

\[f(y) = \pi_0\, \mathcal{N}(y; \mu_0,\sigma_0^2) + \pi_1\, \mathcal{N}(y; \mu_1,\sigma_1^2)\]
White, Baudemont et al. (2026). Whose Line Is it Anyway? Defining Seropositivity Cutoffs for Infectious Disease Surveillance. JID. doi.org/10.1093/infdis/jiaf537
Seroreactivity: degree of antibody reactivity, irrespective of any threshold.
Latent \(T \in [0,1]\)

Taken from Alberts B, Johnson A, Lewis J, Morgan D, Raff M, Roberts K, Walter P (2015). Molecular Biology of the Cell, 6th edition. New York: Garland Science.
Recall the two ways of modelling \(T\):
\[ 1) \quad T \sim \text{Beta}(\alpha,\beta) \]
\[ 2) \quad T \sim \bigl(1-\pi(a)\bigr)\,\text{Beta}(t;\alpha_1,\beta_1) + \pi(a)\,\text{Beta}(t;\alpha_2,\beta_2) \]
Through the shape: \(\alpha(a)=\alpha_0 a^{\gamma}\), \(\beta(a)=\beta_0 a^{\delta}\)
the single Beta’s own shape parameters vary with age
Through the mixing weights: \(\pi(a) = 1-\exp(-\lambda a)\)
age shifts weight from the low- to the high-seroreactivity component
Rachuonyo South District, 2011 — IgG against AMA1 and MSP1 (Bousema et al., 2013, 2016).
Strategy: fix conditional parameters \((\mu_0,\mu_1,\sigma_0^2,\sigma_1^2)\) from the youngest age group, then re-estimate only \(T\)’s distribution within each successive age band.
A change point \(\tau\) separates two regimes, with the mean of \(T\) kept continuous across it.
Below \(\tau\) — mechanistic, with a degenerate mass at \(T=0\):
\[ f_T(t;a) = \begin{cases} 1-\pi(a), & t=0\\[4pt] \pi(a)\,\alpha_2\, t^{\alpha_2-1}, & 0<t<1 \end{cases} \qquad \pi(a) = 1-\exp(-\lambda a) \]
Above \(\tau\) — empirical Beta with logit-linear mean:
\[ T \sim \text{Beta}\big(\mu(a)\phi,\,[1-\mu(a)]\phi\big), \qquad \text{logit}\{\mu(a)\} = \eta_0 + \eta_1\log(a) \]
A single Beta throughout, with both shape parameters varying with age:
\[ T \sim \text{Beta}\big(\alpha(a),\,\beta(a)\big), \qquad \alpha(a) = \alpha_0\, a^{\gamma}, \qquad \beta(a) = \beta_0\, a^{\delta(a)} \]
with a change point \(\zeta\) in the exponent of \(\beta\):
\[ \delta(a) = \begin{cases} \delta_1, & a \le \zeta\\[4pt] \delta_1+\delta_2, & a > \zeta \end{cases} \qquad \mathbb{E}[T] = \left(1 + \tfrac{\beta_0}{\alpha_0}\, a^{\delta_1+\delta_2-\gamma}\right)^{-1} \]
BIC comparison, age-stratified, both fitted with MLE (positive \(\Delta\)BIC favours the latent model):
| Antigen | Age band | \(\Delta\)BIC |
|---|---|---|
| AMA1 | \(6\)–\(10\) | 45 |
| AMA1 | \(20\)–\(30\) | 700 |
| AMA1 | \(45+\) | 1394 |
| MSP1 | \(<6\) | 42 |
| MSP1 | \(20\)–\(30\) | 259 |
| MSP1 | \(45+\) | 459 |
Dependence in multi-antigen serology operates at two levels:
No existing framework handles cross-antigen correlation, age dependence and spatial dependence together for non-Gaussian, multimodal outcomes.
For individual \(i\) and antigen \(k \in \{1,2\}\):
The question the principles answer: where in this hierarchy does each source of dependence belong?
The observation model describes assay characteristics, not biology
Conditional on \(\mathbf{T}_i\), the two antibody measurements are independent:
\[ p(\mathbf{y}_i \mid \mathbf{T}_i = \mathbf{t}) = \prod_{k=1}^{2} \mathcal{N}\big\{y_{i,k};\, \mu_k(t_{i,k}),\, \sigma_k^2(t_{i,k})\big\} \]
Implication: all cross-antigen dependence is carried by the joint density \(g(\mathbf{t}\mid\mathbf{d}_i)\) — nothing is left in the observation model
When it fails: antibody cross-reactivity would induce correlation in \(\mathbf{Y}\) unrelated to \(\mathbf{T}\), requiring a residual covariance term instead
Exposure determines the probability of mounting a high response
Individual immunology determines the magnitude, given exposure
Implication: a mixture for \(\mathbf{T}_i\) is the natural choice, since it provides two distinct places to put them
\[ \mathbf{T}_i \mid \mathbf{Z}_i = \mathbf{z} \;\sim\; \mathcal{N}_{(0,1)^2}\big\{ \mathbf{m}(\mathbf{d}_i;\mathbf{z}),\, \Sigma_T \big\} \]
exposure → the mixing probabilities \(\pi_i(\mathbf{z})\); magnitude → the component locations \(\mathbf{m}(\mathbf{d}_i;\mathbf{z})\)
Consequence: spatial variation, being variation in the force of infection, enters only the mixing probabilities
Shared infection event, within an individual
→ \(\delta(a_i)\) in \(\Pr(Z_{i,2}=1 \mid Z_{i,1}=z_1)\), the association between component labels
Shared within-host processes, within an individual
→ \(\rho_T\), the correlation in \(\Sigma_T\) once the labels are fixed
Shared exposure environment, between individuals
→ \(\rho_S\), the cross-correlation of the fields \(\{S_1(\mathbf{x}), S_2(\mathbf{x})\}\)


Mass concentrates on the diagonal, with an off-diagonal tail.
| Quantity | Estimate | 95% CI |
|---|---|---|
| Spatial range, AMA1 | 473 m | (353, 639) |
| Spatial range, MSP1 | 437 m | (318, 577) |
| Cross-field \(\rho_S\) | 0.587 | (0.453, 0.711) |
| Within-component \(\rho_T\) | 0.717 | (0.691, 0.750) |
| Age effect \(\delta_1\) | 0.286 | (0.211, 0.339) |


Coarsened total variation distance for the joint distribution:
| Age band | Univariate | Joint |
|---|---|---|
| \([1,5]\) | 0.295 | 0.189 |
| \((20,40]\) | 0.225 | 0.083 |
| All ages | 0.194 | 0.078 |
Joint modelling matters if we are answering research questions that involve joint properties of antibodies
The sequential construction extends directly by adding conditional terms \(\delta_{jk}(a_i)\). The harder questions are epidemiological.
Does the pathogen structure matter? Whether the antigens belong to one pathogen or several has direct implications for how cross-antigen association should be modelled
Disease-agnostic approach. Impose no prior structure on the correlations and let the data determine the full pattern of dependence
Assumption-driven approach. Impose a parsimonious structure informed by biology, for example allowing association only among antigens sharing a pathogen or a transmission route
The principles adapt. P3 as stated assumes a shared pathogen and life-cycle stage; for antigens from distinct pathogens the shared-infection pathway no longer applies and \(\delta(a)\) should be left unconstrained
Computation is the binding constraint. The approximations used here are adequate for a bivariate outcome but unlikely to scale to tens of antibody responses
Extending to a handful of antigens from the same pathogen is relatively straightforward. A genuinely multivariate formulation would need to be tailored to the epidemiological context.
References
Giorgi, E. and Wallin, J. (2026). A flexible class of latent variable models for the analysis of antibody response data. Biostatistics, Volume 27, Issue 1, January 2026, kxag008, https://doi.org/10.1093/biostatistics/kxag008
Giorgi, E. and Wallin, J. (2026). Bivariate geostatistical latent variable models for the analysis of antibody density data. Pre-print available at https://arxiv.org/abs/2608.27534
Let’s Connect
🔗 giorgistat.github.io
📧 e.giorgi@bham.ac.uk
📍Department of Health Sciences, University of Birmingham, Birmingham, UK
