Assignment 3 · submitted 22 Oct 2023
North or South? A rainfall classifier story
Each weather station is described by its average rainfall on every day of the year, and labelled as being in the North or the South of Australia. With 365 predictors and only 150 training stations, the textbook classifiers break down, so the assignment explored three ways of reducing the dimension first.
- train stations
- 150
- train stations
- test stations
- 41
- test stations
- daily values
- 365
- daily values
The data
Two rainfall regimes
Northern stations get most of their rain in the summer monsoon (December–March) and much less in winter; southern stations get moderate rain all year with a winter peak. Only class-level summaries are shown here; the individual station records stay in the course files.
- North mean (n = 58)
- South mean (n = 92)
- middle 50% of stations
hover, tap or focus and use ← → for daily values
Question 1
Why not just fit QDA or logistic regression?
Without regularisation, quadratic discriminant analysis needs two mean vectors and two full covariance matrices; logistic regression needs one coefficient per day plus an intercept.
qda() stops with “some group is too small for ‘qda’”.glm warns that fitted probabilities are numerically 0 or 1.Question 2
PCA, then logistic regression on q components
Replace the 365 days by the first q principal components of the standardised curves, fit , and choose by hand-written leave-one-out cross-validation.
Question 3(1)–(2)
Random forest on all 365 days
Trees do not mind p > n. Each of the 5,000 trees sees a bootstrap sample of stations and, at every split, a random subset of m days; m is tuned by the out-of-bag error on a grid around .
Question 3(3)
One tree on 50 PLS components
Partial least squares builds components that, unlike PCA, are chosen to covary with the label. A single classification tree on the first 50 PLS components only needs two of them: C1, which is strongly negatively correlated (about −0.9) with December–March rainfall, so wet-summer stations score low; and C2, which tracks winter rainfall and rescues a handful of borderline stations. C1 is therefore built from the same late-January days the random forest ranked highest.
Summary
Test errors, as submitted and corrected
Misclassified test stations out of 41, re-run in R 4.6 from the original code and from the corrected code.
PCA + logistic
- As submitted (2023)
- 4 errors (q = 27)
- Corrected
- 1 error (q = 3)
real held-out LOOCV; test set standardised; probability threshold
Random forest
- As submitted (2023)
- 3 errors (m = 153)
- Corrected
- 3 errors (m = 57)
<instead of<-(the integer grid alone changes nothing)Tree on 50 PLS comps
- As submitted (2023)
- 16 errors
- Corrected
- 1 error
test set divided by the training SDs before projection
The as-submitted column reproduces the 2023 PDF exactly (4, 3 and 16 errors). With 41 test stations, differences of one or two errors are within noise; the PLS-tree change is not.