Skip to content
Multivariate Lab

Assignment 1 · Problems 1(c)–2 · submitted 11 Sep 2023

Wheat seeds PCA explorer

Seventy X-ray-scanned kernels from each of three wheat varieties (Kama, Rosa and Canadian), each described by seven shape measurements. Principal components turn seven correlated measurements into a few uncorrelated summaries, and here two of them are enough to see the three varieties pull apart.

kernels
210
kernels
measurements
7
measurements
varieties
3
varieties

The data

Seven measurements, one label

The eighth column is the variety code 1–3. It is a category, not a measurement, so it is left out of the PCA and only used to colour the plots (Problem 2(a)).

Problem 2

Principal components, three ways

PCA finds orthogonal directions γ1,γ2,…\gamma_1,\gamma_2,\dots of decreasing variance λ1≥λ2≥⋯\lambda_1\ge\lambda_2\ge\cdots. The cumulative proportion ψk=∑j≤kλj/∑jλj\psi_k=\sum_{j\le k}\lambda_j/\sum_j\lambda_j tells how much of the total variance the first k components keep. Switch the variant to see what changes when the analysis is done on the covariance matrix (as asked), on the correlation matrix, or in the mixed way it was submitted.

Which analysis?

Every chart below is recomputed in your browser (Jacobi eigen-decomposition, 210 × 7).

What the question asked: PCA of S
On the raw scale, area (X1) and perimeter (X2) have by far the largest variances, so PC1 alone explains 82.9% of the total variance and is essentially a size component. This is the usual argument for standardising first, but it is a different analysis from the one requested.
Scree plot and cumulative variance ψ
Eigenvalues of S · bars: share of variance, line: ψ_k
PC1: λ = 10.79, 82.94% of variancePC2: λ = 2.129, 16.36% of variancePC3: λ = 0.07363, 0.57% of variancePC4: λ = 0.01289, 0.10% of variancePC5: λ = 0.002748, 0.02% of variancePC6: λ = 0.001570, 0.01% of variancePC7: λ = 0.00002966, 0.00% of varianceψ1 = 82.94%ψ2 = 99.30%ψ3 = 99.87%ψ4 = 99.97%ψ5 = 99.99%ψ6 = 100.00%ψ7 = 100.00%
PC1234567
λ10.82.130.07360.01290.002750.001573.0e-5
share82.9%16.4%0.6%0.1%0.0%0.0%0.0%
ψ82.9%99.3%99.9%100.0%100.0%100.0%100.0%
PC1 vs PC2 scores by variety
centred X × eigenvectors of S
  • Kama (variety 1)
  • Rosa (variety 2)
  • Canadian (variety 3)
Correlation circle: PC1 and PC2
ρ(Xᵢ, PCⱼ) = γᵢⱼ √λⱼ / √sᵢᵢ
X1 Area: (0.998, 0.051)X2 Perimeter: (0.995, 0.063)X3 Compactness: (0.599, -0.179)X4 Kernel length: (0.953, 0.101)X5 Kernel width: (0.966, 0.009)X6 Asymmetry: (-0.279, 0.960)X7 Groove length: (0.862, 0.244)X1X2X3X4X5X6X7
Loadings γ (weights of each variable)
first three principal components; bar length = |γᵢⱼ|
Variableγ·1γ·2γ·3
X1 Area
0.884
0.101
-0.265
X2 Perimeter
0.395
0.056
0.283
X3 Compactness
0.004
-0.003
-0.059
X4 Kernel length
0.129
0.031
0.400
X5 Kernel width
0.111
0.002
-0.319
X6 Asymmetry
-0.128
0.989
-0.064
X7 Groove length
0.129
0.082
0.762

The coefficients do not depend on the kernel i; the scores PCᵢ = γᵀ(Xᵢ − μ) do. Signs of eigenvectors are arbitrary (here the largest entry is made positive; the as-submitted view uses the 2023 signs).

Problem 1(d)

S=ΓΛΓ⊤S=\Gamma\Lambda\Gamma^{\top} for the wheat data

The covariance-PCA eigenvalues are exactly the entries of Λ\Lambda below. Computed here with a Jacobi eigen-solver in TypeScript and checked against R's eigen(cov(X)) to 10 decimal places in the test suite.

Sample covariance S (7 × 7)

Cell shade ∝ |sᵢⱼ|. Area (X1) has variance 8.47 while compactness (X3) has 5.6e-4: four orders of magnitude apart.

X1X2X3X4X5X6X7
X18.4663.7780.0421.2251.067-1.0041.235
X23.7781.7060.0160.5630.466-0.4270.572
X30.0420.0165.6e-40.0040.007-0.0120.003
X41.2250.5630.0040.1960.144-0.1140.203
X51.0670.4660.0070.1440.143-0.1470.139
X6-1.004-0.427-0.012-0.114-0.1472.261-0.008
X71.2350.5720.0030.2030.139-0.0080.242

Λ: eigenvalues of S

  1. λ110.7933
  2. λ22.1295
  3. λ30.0736
  4. λ40.0129
  5. λ52.748e-3
  6. λ61.570e-3
  7. λ72.966e-5

Γ: orthonormal eigenvectors (columns)

γ1γ2γ3γ4γ5γ6γ7
X10.8840.101-0.2650.1990.137-0.281-0.025
X20.3950.0560.283-0.579-0.5750.3020.066
X30.004-0.003-0.0590.0580.0530.0450.994
X40.1290.0310.400-0.4360.7870.1130.001
X50.1110.002-0.3190.2340.1450.896-0.082
X6-0.1280.989-0.064-0.0250.002-0.0030.001
X70.1290.0820.7620.613-0.0880.1100.009

Check: max |ΓΛΓᵀ − S| = 1.8e-14 (R's round(S - S.sample) in the 2023 answer printed a matrix of zeros).

Reading the results
With the covariance matrix the first component is dominated by area and perimeter, the two variables with the largest raw variances. After standardising, three components keep about 98.7% of the variance (ψ₃), matching the submitted “keep 3 PCs” conclusion. Along PC1 (size) Rosa kernels are the largest and Canadian the smallest, with Kama in between; PC2 sets Canadian further apart because its kernels are the least compact and the most asymmetric.