A reproducible journey through ANOVA, linear models and experimental design using a realistic analytical-chemistry problem.
We have four analytical methods and three environmental matrices.
Our response is relative bias.
At first, the experiment is perfectly balanced and the observations are independent.
Then we deliberately make the problem harder.
Why?
Because real statistical problems rarely arrive labelled "two-way ANOVA" or "ANCOVA". They emerge when the scientific question becomes more precise.
Which analytical method should we use?
At first this seems to be a comparison of four means.
But progressively we discover that:
- not all differences are scientifically relevant;
- the matrix matters;
- the effect of the method depends on the matrix;
- matrix composition is a continuous variable;
- the design can be unbalanced;
- model assumptions can fail;
- observations can become dependent.
Each chapter introduces a statistical model because the previous model has become insufficient to answer the next scientific question.
01 Four methods ↓ One-way ANOVA
02 Which differences matter? ↓ Contrasts
03 Does matrix matter? ↓ Two-way ANOVA
04 Does method performance depend on matrix? ↓ Interaction
05 What happens when the design is unbalanced? ↓ Type I / II / III
06 Is matrix composition the real driver? ↓ ANCOVA
07 Can we trust the model? ↓ Diagnostics
08 What if assumptions fail? ↓ Robustness
09 What if we have several responses? ↓ MANOVA / MANCOVA
10 What if observations are not independent? ↓ Mixed models
The sequence is deliberately cumulative: later models build on questions and concepts introduced earlier.
The repository contains simulated analytical chemistry data designed to reproduce common experimental structures encountered when comparing analytical methods.
The main dataset contains measurements obtained by applying four analytical methods to samples belonging to three matrices. Each sample is analysed by all four methods, reating a repeated-measures structure that becomes important in the later episodes.
The data are synthetic and are intended for educational purposes. They should not be interpreted as results from a real analytical method comparison.
The project is developed as a reproducible Quarto project.
R package dependencies are managed with renv, with the project environment recorded in renv.lock.
To reproduce the analysis:
- Clone the repository.
- Open the project in RStudio or another R environment.
- Restore the project library with:
renv::restore()- Render the desired Quarto document, or render the complete project.
Using renv ensures that the package versions used for the analysis
can be restored rather than relying on whatever versions
happen to be installed on the user's system.
.
├── data/ # Input datasets
├── analyses/ # Quarto documents for the individual episodes
├── posts/ # Quarto documents for LinkedIn posts
├── renv/ # Project-specific renv infrastructure
├── renv.lock # Locked R package versions
└── ...
The exact directory structure may evolve as the project develops.
The analyses are written in R and documented with Quarto.
The code favours transparent base R and lightweight packages where practical.
data.table is used for data manipulation, while additional packages are introduced
only when they provide functionality directly relevant to the statistical analysis.
This project is intentionally limited in scope.
It is not intended to be a general introduction to statistics or a comprehensive treatment of experimental design. It focuses on a connected set of statistical models that are particularly useful for understanding method-comparison data in analytical chemistry.
The central question throughout the project is not simply whether a statistical test produces a significant p-value, but whether the statistical model adequately represents the experimental structure and supports the scientific conclusion being drawn from it.