Skip to content
Multivariate Lab

Assignment 3 · submitted 22 Oct 2023

North or South? A rainfall classifier story

Each weather station is described by its average rainfall on every day of the year, and labelled as being in the North or the South of Australia. With 365 predictors and only 150 training stations, the textbook classifiers break down, so the assignment explored three ways of reducing the dimension first.

train stations
150
train stations
test stations
41
test stations
daily values
365
daily values

The data

Two rainfall regimes

Northern stations get most of their rain in the summer monsoon (December–March) and much less in winter; southern stations get moderate rain all year with a winter peak. Only class-level summaries are shown here; the individual station records stay in the course files.

northern wet season
  • North mean (n = 58)
  • South mean (n = 92)
  • middle 50% of stations

hover, tap or focus and use ← → for daily values

Question 1

Why not just fit QDA or logistic regression?

Without regularisation, quadratic discriminant analysis needs two mean vectors and two full covariance matrices; logistic regression needs one coefficient per day plus an intercept.

134,321
parameters for QDA
2 × 365 means + 2 × 66,795 covariance entries + 1 prior
366
parameters for logistic regression
365 slopes + 1 intercept
150
training stations
fewer observations than parameters
QDA cannot even start
Each class covariance Σ^k\hat\Sigma_k is 365 × 365 but estimated from 58 or 92 stations, so it is singular; R's qda() stops with “some group is too small for ‘qda’”.
Logistic regression separates perfectly
With more predictors than stations the classes are linearly separable, the likelihood has no maximum, coefficients run off to ±∞ with huge standard errors, and glm warns that fitted probabilities are numerically 0 or 1.

Question 2

PCA, then logistic regression on q components

Replace the 365 days by the first q principal components of the standardised curves, fit logit⁡P(G=1)=β0+∑k≤qβkYk\operatorname{logit} P(G=1)=\beta_0+\sum_{k\le q}\beta_k Y_k, and choose q∈{1,…,30}q\in\{1,\dots,30\} by hand-written leave-one-out cross-validation.

Loading principal component scores…

Question 3(1)–(2)

Random forest on all 365 days

Trees do not mind p > n. Each of the 5,000 trees sees a bootstrap sample of stations and, at every split, a random subset of m days; m is tuned by the out-of-bag error on a grid around c=365c=\sqrt{365}.

Loading random-forest results…
Interpretation: the monsoon decides
The most important days cluster in late January and early February, the heart of the northern wet season. On those days northern stations receive heavy rain while southern stations are in their dry summer, so a single day's rainfall already separates most stations; the forest picks them out without being told anything about seasons.

Question 3(3)

One tree on 50 PLS components

Partial least squares builds components that, unlike PCA, are chosen to covary with the label. A single classification tree on the first 50 PLS components only needs two of them: C1, which is strongly negatively correlated (about −0.9) with December–March rainfall, so wet-summer stations score low; and C2, which tracks winter rainfall and rescues a handful of borderline stations. C1 is therefore built from the same late-January days the random forest ranked highest.

Loading PLS components…

Summary

Test errors, as submitted and corrected

Misclassified test stations out of 41, re-run in R 4.6 from the original code and from the corrected code.

The as-submitted column reproduces the 2023 PDF exactly (4, 3 and 16 errors). With 41 test stations, differences of one or two errors are within noise; the PLS-tree change is not.