Skip to content

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ANOVA and Linear Models for Analytical Chemistry

When ANOVA stops being simple

A reproducible journey through ANOVA, linear models and experimental design using a realistic analytical-chemistry problem.

We have four analytical methods and three environmental matrices.

Our response is relative bias.

At first, the experiment is perfectly balanced and the observations are independent.

Then we deliberately make the problem harder.

Why?

Because real statistical problems rarely arrive labelled "two-way ANOVA" or "ANCOVA". They emerge when the scientific question becomes more precise.

The scientific question

Which analytical method should we use?

At first this seems to be a comparison of four means.

But progressively we discover that:

  1. not all differences are scientifically relevant;
  2. the matrix matters;
  3. the effect of the method depends on the matrix;
  4. matrix composition is a continuous variable;
  5. the design can be unbalanced;
  6. model assumptions can fail;
  7. observations can become dependent.

Each chapter introduces a statistical model because the previous model has become insufficient to answer the next scientific question.

01 Four methods ↓ One-way ANOVA

02 Which differences matter? ↓ Contrasts

03 Does matrix matter? ↓ Two-way ANOVA

04 Does method performance depend on matrix? ↓ Interaction

05 What happens when the design is unbalanced? ↓ Type I / II / III

06 Is matrix composition the real driver? ↓ ANCOVA

07 Can we trust the model? ↓ Diagnostics

08 What if assumptions fail? ↓ Robustness

09 What if we have several responses? ↓ MANOVA / MANCOVA

10 What if observations are not independent? ↓ Mixed models

The sequence is deliberately cumulative: later models build on questions and concepts introduced earlier.

Data

The repository contains simulated analytical chemistry data designed to reproduce common experimental structures encountered when comparing analytical methods.

The main dataset contains measurements obtained by applying four analytical methods to samples belonging to three matrices. Each sample is analysed by all four methods, reating a repeated-measures structure that becomes important in the later episodes.

The data are synthetic and are intended for educational purposes. They should not be interpreted as results from a real analytical method comparison.

Reproducibility

The project is developed as a reproducible Quarto project.

R package dependencies are managed with renv, with the project environment recorded in renv.lock.

To reproduce the analysis:

  1. Clone the repository.
  2. Open the project in RStudio or another R environment.
  3. Restore the project library with:
  4. renv::restore()
  5. Render the desired Quarto document, or render the complete project.

Using renv ensures that the package versions used for the analysis can be restored rather than relying on whatever versions happen to be installed on the user's system.

Project structure

.
├── data/              # Input datasets
├── analyses/          # Quarto documents for the individual episodes
├── posts/             # Quarto documents for LinkedIn posts
├── renv/              # Project-specific renv infrastructure
├── renv.lock          # Locked R package versions
└── ...

The exact directory structure may evolve as the project develops.

Software

The analyses are written in R and documented with Quarto.

The code favours transparent base R and lightweight packages where practical. data.table is used for data manipulation, while additional packages are introduced only when they provide functionality directly relevant to the statistical analysis.

Scope

This project is intentionally limited in scope.

It is not intended to be a general introduction to statistics or a comprehensive treatment of experimental design. It focuses on a connected set of statistical models that are particularly useful for understanding method-comparison data in analytical chemistry.

The central question throughout the project is not simply whether a statistical test produces a significant p-value, but whether the statistical model adequately represents the experimental structure and supports the scientific conclusion being drawn from it.

About

A reproducible journey through ANOVA, linear models and experimental design using a realistic analytical-chemistry problem.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages