Skip to content

Repository files navigation

Experiment Runner

An adaptive experiment runner for behavioral research. Define branching, looping, and randomized participant flows as typed graph configs; render them as structured multi-screen web studies.


The Problem

Researchers running online behavioral studies — psychology experiments, UX tests, clinical screeners — face a difficult tradeoff:

Commercial survey tools (Qualtrics, SurveyMonkey, Gorilla) offer drag-and-drop authoring but fall short on complex adaptive designs: multi-level loops, probability-weighted random assignment, composite branching conditions across many responses, or nested path structures with per-path progress indicators.

Custom code offers full flexibility but requires developer involvement for every new study, making iteration slow and error-prone.

This project occupies the space between them: a typed, declarative experiment schema expressive enough to handle complex adaptive designs, paired with a React participant UI and a static validator that catches misconfigured flows before anyone runs them.


Core Ideas

Experiments as graphs

An experiment is a directed graph of nodes connected by edges. The researcher defines the graph; the engine traverses it as a state machine.

start → screen[welcome] → branch[age-check] → screen[adult-path]
                                            ↘ screen[minor-path]

Nodes represent steps; edges define how the engine moves between them.

Node types

Node Purpose
start Entry point. Supports multiple named entry points via URL parameters.
screen Displays a screen to the participant. Each screen is a list of components.
branch Evaluates conditions against collected data and routes to the first matching arm (or a default).
fork Random assignment. Each arm has a weight; the engine selects one probabilistically.
path A sequential group of child nodes, optionally randomized. Supports a progress stepper.
loop Repeats a template node over a list of values — either static or drawn from collected data.
checkpoint Triggers a data submission at a specific point in the flow (useful for long studies).

Screen components

Each screen is composed of typed components across four families:

  • Content — rich-text, image, video, audio. Display only, no data collected.
  • Response — slider, radio, checkboxes, likert-scale, text-input, text-area, numeric-input, date-input, time-input, single-checkbox, dropdown. Each stores its value under a typed dataKey.
  • Layout — button (advances the screen), group (wraps child components).
  • Control — conditional (shows a component only when a condition is met), for-each (renders a component template once per item in a list).

Data references

Collected values are referenced using a prefix notation:

Prefix Scope Example
$$ Experiment-wide collected data $$demographics.age
@ Current loop item (keyed by loop node ID) @loop-sports.value, @loop-sports.index
$ Current screen's live form values $hasChildren
# Current for-each item (keyed by for-each component ID) #foreach-sport.value, #foreach-sport.index
% Shared option sets defined in ExperimentFlow.options %agreement-scale

These references are used in branch conditions, answer piping (string interpolation inside labels and content), and conditional rendering.

Answer piping

Any string prop — labels, rich-text content, image URLs and alt text — can interpolate collected values using {{ }} template syntax:

"How are you feeling about {{$$welcome.name}}'s results?"
"Rate your experience with {{@value}}"

Conditions

Branch conditions and conditional components use a composable condition structure:

// simple comparison
{ type: "simple", operator: "gte", dataKey: "$$screening.age", value: 18 }

// compound — and / or / not
{ type: "and", conditions: [
  { type: "simple", operator: "eq", dataKey: "$$consent.agreed", value: true },
  { type: "simple", operator: "gte", dataKey: "$$screening.age", value: 18 },
]}
{ type: "not", condition: { type: "simple", operator: "eq", dataKey: "$skip", value: true } }

Available operators: eq, neq, lt, lte, gt, gte, contains, length-eq, length-neq, length-lt, length-lte, length-gt, length-gte.

Static validation

validateExperiment(flow) walks the graph before the participant ever sees it and returns typed errors for:

  • Duplicate node IDs, missing start node
  • Edges referencing nodes that don't exist
  • Branch nodes missing a default arm
  • Screen nodes with no matching screen definition
  • $$ references used before the corresponding data is collected in the flow
  • @ references used outside a loop context

This makes misconfigured experiments a build-time problem rather than a runtime surprise.


Architecture

packages/engine/        # Pure, framework-agnostic engine (@experiment-hub/engine)
  flow/                 # State machine: traverse, enterStep, traverseInPath/Loop
  types.ts              # ExperimentFlow, FlowStep, State, Context
  nodes.ts, edges.ts    # Node and edge type definitions
  conditions.ts         # Condition evaluation
  resolve.ts            # Data key resolution and string interpolation
  field-schema.ts, screen-schema.ts  # Per-screen form validation
  experiment-validation/# Static experiment graph validator
  components/           # Component type definitions (content, response, layout, control)
  specs/                # Unit tests for the flow engine

apps/frontend/          # Next.js React application (@experiment-hub/frontend)
  app/                  # App Router — experiments routed at /experiments/[slug]
  src/
    Experiment.tsx      # Top-level experiment runner component
    Screen.tsx          # Screen renderer with react-hook-form integration
    data/
      experiments/      # EXPERIMENTS record — one file per experiment
      store.ts          # Zustand store: step, start(), next()
      send.ts           # Checkpoint POST to the backend
    components/         # RenderComponent + content/response/layout/control
    specs/              # Unit tests for React components
  e2e/                  # Playwright tests

apps/backend/           # Effect 4 API (@experiment-hub/backend)
  src/                  # HttpApi contract, Checkpoints service, SQLite layers

infra/nginx/            # Single-origin reverse proxy conf
infra/cloudflared/      # Locally-managed tunnel ingress config
docs/                   # Domain reference documentation

The flow engine in packages/engine/ has no React dependency and is fully unit-tested. The React layer drives it by calling startExperiment() once and traverse(step, formData) on each screen submission.


Current State

This project is an early-stage working prototype. The flow engine and component library are functional; the infrastructure around them is scaffolding.

What works

  • Full flow traversal: branch, fork, path, loop, checkpoint
  • All 12 response component types with Zod-based form validation
  • Answer piping in labels and rich-text content (rich-text, image, labels, placeholders, option labels, button text)
  • Conditional rendering within screens (conditional, for-each)
  • Checkpoint persistence: checkpoint nodes and run completion POST full-context snapshots to apps/backend (Effect + SQLite); runs are server-issued (POST /api/runs → run id + signed token); researcher export GET /api/experiments/:slug/export (NDJSON, bearer-gated)
  • Static experiment validator with 36 error codes
  • Unit test suite for the flow engine

What is missing or incomplete

Session persistence

Data submission hardening — anyone can still register a run for a public study; the run token stops forged run ids and slug swaps, not a determined fabricator (see #108).

Visual flow builder — Experiments are currently defined as TypeScript object literals in apps/frontend/src/data/experiments/. A drag-and-drop canvas editor using @xyflow/react is planned but not started.

Score variables — There is no way to compute derived values (e.g. sum of five Likert items) and branch on them.

Back navigation — Participants cannot go back to a previous screen.

Debug artifacts — The current UI renders raw JSON debug panels (flow state, form values) directly on screen. These exist for development only and must be removed before any participant-facing deployment.


Running locally

pnpm install
pnpm dev            # frontend on :3000
pnpm dev:backend    # backend on :3100 (needed for checkpoint persistence)

Open http://localhost:3000. Experiments are defined in apps/frontend/src/data/experiments/ and routed by slug (/experiments/ocean).

In dev, Next rewrites /api/* to the backend (BACKEND_URL env var overrides the default http://localhost:3100).

Deploying

docker-compose.yml runs the whole stack on a single host behind a Cloudflare Tunnel — backend (Effect + SQLite on a volume), frontend (Next standalone), nginx (single origin: / → frontend, /api/* → backend), cloudflared (only public ingress).

The tunnel is locally managed: ingress rules live in infra/cloudflared/config.yml and the connector authenticates with ./.cloudflared/credentials.json — a gitignored copy of the ~/.cloudflared/<tunnel-id>.json that cloudflared tunnel create produces (after cloudflared tunnel login). No dashboard config needed.

Provisioning a new deployment (the committed config.yml references this project's tunnel id and hostname — substitute your own):

cloudflared tunnel login                            # browser auth, writes ~/.cloudflared/cert.pem
cloudflared tunnel create <name>                    # writes ~/.cloudflared/<tunnel-id>.json
cloudflared tunnel route dns <name> <your-domain>   # creates the CNAME for the apex/hostname

Then update tunnel: and hostname: in infra/cloudflared/config.yml, and copy ~/.cloudflared/<tunnel-id>.json to .cloudflared/credentials.json.

Copy .env.example to .env, set EXPORT_TOKEN and RUN_TOKEN_SECRET (optionally ALLOWED_EXPERIMENTS), place credentials.json in .cloudflared/, then:

docker compose up -d --build

Tests

pnpm test

Unit tests live in packages/engine/specs/, apps/frontend/src/specs/, and apps/backend/src/. The flow engine tests cover branch, fork, path, loop, integration, and validation scenarios.


Defining an experiment

An experiment config is an ExperimentFlow object:

import { ExperimentFlow } from '@experiment-hub/engine/types';

export const experiment: ExperimentFlow = {
  nodes: [
    { id: 'start', type: 'start' },
    { id: 'screen-welcome', type: 'screen', props: { slug: 'welcome' } },
    {
      id: 'branch-age',
      type: 'branch',
      props: {
        name: 'Age gate',
        branches: [
          {
            id: 'adult',
            name: 'Adult',
            config: {
              type: 'simple',
              operator: 'gte',
              dataKey: '$$welcome.age',
              value: 18,
            },
          },
        ],
      },
    },
    { id: 'screen-adult', type: 'screen', props: { slug: 'adult-content' } },
    { id: 'screen-ineligible', type: 'screen', props: { slug: 'ineligible' } },
  ],
  edges: [
    { type: 'sequential', from: 'start', to: 'screen-welcome' },
    { type: 'sequential', from: 'screen-welcome', to: 'branch-age' },
    { type: 'branch-condition', from: 'branch-age.adult', to: 'screen-adult' },
    { type: 'branch-default', from: 'branch-age', to: 'screen-ineligible' },
  ],
  screens: [
    {
      slug: 'welcome',
      components: [
        {
          componentFamily: 'content',
          template: 'rich-text',
          props: { content: '## Welcome\nHow old are you?' },
        },
        {
          componentFamily: 'response',
          template: 'numeric-input',
          props: { label: 'Age', dataKey: 'age', min: 0, max: 120 },
        },
        {
          componentFamily: 'layout',
          template: 'button',
          props: { text: 'Continue' },
        },
      ],
    },
    // ...
  ],
};

Before using a config in production, run:

import { validateExperiment } from '@experiment-hub/engine/experiment-validation';

const errors = validateExperiment(experiment);
if (errors.length > 0) {
  console.error(errors);
}

About

Adaptive experiment runner for behavioral research — define branching, looping, and randomized participant flows as typed graph configs, render them as structured multi-screen web studies.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages